Skill Detail

Convert complex PDFs and document images into agent-ready Markdown with OCRFlux

Use OCRFlux when agents need local PDF or image parsing into Markdown/JSONL with layout-aware OCR, table handling, and cross-page paragraph or table merging.

Data Extraction & TransformationMulti-Framework
Data Extraction & Transformation Multi-Framework Security Reviewed
⭐ 2.5k GitHub stars
COPY SKILL INSTRUCTIONS (OPTIONAL)
npx skills add agentskillexchange/skills --skill convert-complex-pdfs-and-document-images-into-agent-ready-markdown-with-ocrflux Copy
Uses the third-party skills CLI, not an ASE-owned installer. Check your agent’s compatibility. This copies instructions; complete the upstream tool setup below separately.
At a glance
Tools required
Python 3.11, conda, NVIDIA GPU, vLLM, poppler-utils, OCRFlux-3B model
Install & setup
Install system dependencies such as poppler-utils, poppler-data, Microsoft core fonts, Croscore fonts, gsfonts, and lcdf-typetools. Create a clean environment with conda create -n ocrflux python=3.11, activate it, clone https://github.com/chatdoc-com/OCRFlux.git, then run pip install -e . --find-links https://flashinfer.ai/whl/cu124/torch2.5/flashinfer/. Download or mount the OCRFlux-3B model before running the pipeline.
Author
ChatDOC
Publisher
Open Source
Last updated
Jun 21, 2026
Quick brief

OCRFlux is a local multimodal document parsing toolkit for turning scanned PDFs, layout-heavy PDFs, and document images into clean Markdown and structured JSONL that agents can inspect, cite, summarize, or load into retrieval workflows.

How it works

What this skill actually does

Use this skill when an operator needs to prepare difficult documents for downstream agent work: long reports, papers, forms, or tables where ordinary OCR loses reading order, headers, footers, equations, or cross-page table continuity. The agent workflow is to stage the source documents, run the OCRFlux pipeline against a bounded workspace, review fallback pages and generated JSONL, then convert the results into final Markdown artifacts for ingestion or human review.

Invoke OCRFlux instead of using a normal PDF viewer, manual copy/paste, or generic OCR when the document needs repeatable conversion into LLM-ready text and when local GPU inference is acceptable. It is not a general document management system, a hosted extraction product, or a generic OCR library listing; the boundary is document-to-Markdown preprocessing for agent ingestion and review.

Typical workflow:
1. Confirm the machine has a supported NVIDIA GPU, enough disk space, poppler-utils, fonts, and the OCRFlux model available.
2. Create a clean Python 3.11 environment and install OCRFlux from the upstream repository.
3. Run `python -m ocrflux.pipeline ./localworkspace –data –model /model_dir/OCRFlux-3B` for a file or directory.
4. Inspect `./localworkspace/results` JSONL for `document_text`, `page_texts`, and `fallback_pages`.
5. Run `python -m ocrflux.jsonl_to_markdown ./localworkspace` and hand the Markdown directory to the summarization, extraction, RAG, or evidence-review workflow.

Keep the scope narrow: OCRFlux is appropriate for high-fidelity OCR and PDF/image-to-Markdown conversion, especially cross-page table and paragraph merging. Use lighter tools when documents are already text-native, when no local GPU is available, or when a hosted compliance-reviewed document extraction service is required.