Skill Detail

Parse agent-ready PDFs and document images with MonkeyOCR

Run MonkeyOCR over PDFs, scanned pages, formulas, and tables to produce structured Markdown/text that downstream agents can ingest and reason over.

Data Extraction & TransformationMulti-Framework
Data Extraction & Transformation Multi-Framework Security Reviewed
⭐ 6.6k GitHub stars
COPY SKILL INSTRUCTIONS (OPTIONAL)
npx skills add agentskillexchange/skills --skill parse-agent-ready-pdfs-and-document-images-with-monkeyocr Copy
Uses the third-party skills CLI, not an ASE-owned installer. Check your agent’s compatibility. This copies instructions; complete the upstream tool setup below separately.
At a glance
Tools required
Python environment, CUDA-capable GPU or supported inference backend, Hugging Face Hub or ModelScope access for model weights, PDF or image documents to parse
Install & setup
Follow the upstream install guide, download model weights with python tools/download_model.py -n MonkeyOCR-pro-3B or another documented model name, then run python parse.py input_path with optional flags such as -o ./output, -g 20, -s, or -t text/formula/table for the parsing mode.
Author
Yuliang-Liu
Publisher
Individual
Last updated
Jun 20, 2026
Quick brief

Use MonkeyOCR when an agent workflow needs local document parsing before retrieval, analysis, or data extraction. The operator installs the project, downloads the selected MonkeyOCR model weights, runs the parser against PDFs, images, folders, or page batches, and routes the Markdown or recognition output into the next agent step. Invoke this instead of opening the OCR product or manually copying document text when the job is repeatable, batchable parsing for agent ingestion. The scope is document parsing and recognition only: it is not a general OCR product listing, a broad LMM framework, or an MCP server.

How it works

What this skill actually does

Inputs and prerequisites: Python environment, CUDA-capable GPU or supported inference backend, Hugging Face Hub or ModelScope access for model weights, PDF or image documents to parse.

Setup notes: Follow the upstream install guide, download model weights with `python tools/download_model.py -n MonkeyOCR-pro-3B` or another documented model name, then run `python parse.py input_path` with optional flags such as `-o ./output`, `-g 20`, `-s`, or `-t text/formula/table` for the parsing mode.

Source and verification boundary: use https://github.com/Yuliang-Liu/MonkeyOCR as the canonical reference before running the workflow; keep commands, API calls, CLI usage, and generated outputs reviewable against that upstream source.

Framework fit: publish this as a Multi-Framework workflow only when the operator can invoke the documented toolchain directly, rather than treating the upstream project as a generic product listing.