Parse documents into structured content for agent ingestion with Dedoc
Extract document text, tables, logical structure, and metadata into normalized output before RAG, review, or workflow automation.
npx skills add agentskillexchange/skills --skill parse-documents-into-structured-content-for-agent-ingestion-with-dedoc
docker run -p 1231:1231 --rm dedocproject/dedoc python3 /dedoc_root/dedoc/main.py, or install the Python package from the upstream project and call Dedoc from the parsing workflow.Use Dedoc when an agent needs to turn PDFs, scans, DOCX, HTML, spreadsheets, archives, or text files into structured content for downstream review, RAG ingestion, data extraction, or compliance workflows. The operator runs Dedoc as a Docker service or Python library, submits documents, captures the logical tree, tables, metadata, and extracted text, then routes the normalized output into the next workflow step. Invoke this instead of using the product normally when the job is repeatable document parsing with structured evidence, not manual document viewing or general OCR exploration. The boundary is document-to-structured-output extraction and normalization; it is not a generic Python library listing or a full document management platform.
What this skill actually does
Inputs and prerequisites: Dedoc Docker service or Python library, source documents.
Setup notes: Run the published Docker image with docker run -p 1231:1231 –rm dedocproject/dedoc python3 /dedoc_root/dedoc/main.py , or install the Python package from the upstream project and call Dedoc from the parsing workflow.
Source and verification boundary: use https://dedoc.readthedocs.io/en/latest/ as the canonical reference before running the workflow; keep commands, API calls, CLI usage, and generated outputs reviewable against that upstream source.
Framework fit: publish this as a Multi-Framework workflow only when the operator can invoke the documented toolchain directly, rather than treating the upstream project as a generic product listing.