Skill Detail

Parse documents into structured content for agent ingestion with Dedoc

Extract document text, tables, logical structure, and metadata into normalized output before RAG, review, or workflow automation.

Data Extraction & TransformationMulti-Framework
Data Extraction & Transformation Multi-Framework Security Reviewed
⭐ 712 GitHub stars
COPY SKILL INSTRUCTIONS (OPTIONAL)
npx skills add agentskillexchange/skills --skill parse-documents-into-structured-content-for-agent-ingestion-with-dedoc Copy
Uses the third-party skills CLI, not an ASE-owned installer. Check your agent’s compatibility. This copies instructions; complete the upstream tool setup below separately.
At a glance
Tools required
Dedoc Docker service or Python library, source documents
Install & setup
Run the published Docker image with docker run -p 1231:1231 --rm dedocproject/dedoc python3 /dedoc_root/dedoc/main.py, or install the Python package from the upstream project and call Dedoc from the parsing workflow.
Author
ISP RAS
Publisher
Organization
Last updated
Jul 6, 2026
Quick brief

Use Dedoc when an agent needs to turn PDFs, scans, DOCX, HTML, spreadsheets, archives, or text files into structured content for downstream review, RAG ingestion, data extraction, or compliance workflows. The operator runs Dedoc as a Docker service or Python library, submits documents, captures the logical tree, tables, metadata, and extracted text, then routes the normalized output into the next workflow step. Invoke this instead of using the product normally when the job is repeatable document parsing with structured evidence, not manual document viewing or general OCR exploration. The boundary is document-to-structured-output extraction and normalization; it is not a generic Python library listing or a full document management platform.

How it works

What this skill actually does

Inputs and prerequisites: Dedoc Docker service or Python library, source documents.

Setup notes: Run the published Docker image with docker run -p 1231:1231 –rm dedocproject/dedoc python3 /dedoc_root/dedoc/main.py , or install the Python package from the upstream project and call Dedoc from the parsing workflow.

Source and verification boundary: use https://dedoc.readthedocs.io/en/latest/ as the canonical reference before running the workflow; keep commands, API calls, CLI usage, and generated outputs reviewable against that upstream source.

Framework fit: publish this as a Multi-Framework workflow only when the operator can invoke the documented toolchain directly, rather than treating the upstream project as a generic product listing.