Skill Detail

Deploy document-to-JSON extraction APIs and ETL pipelines with Unstract

Use Unstract when an operator needs agents or automation pipelines to turn recurring PDFs, scans, and document batches into structured JSON through prompt-defined extraction workflows.

Data Extraction & TransformationMulti-Framework
Data Extraction & Transformation Multi-Framework Security Reviewed
⭐ 6.7k GitHub stars
COPY SKILL INSTRUCTIONS (OPTIONAL)
npx skills add agentskillexchange/skills --skill deploy-document-to-json-extraction-apis-and-etl-pipelines-with-unstract Copy
Uses the third-party skills CLI, not an ASE-owned installer. Check your agent’s compatibility. This copies instructions; complete the upstream tool setup below separately.
At a glance
Tools required
Unstract platform; Docker and Docker Compose; Git; LLM provider credentials; target document source; optional API, ETL, MCP, or n8n integration
Install & setup
Clone https://github.com/Zipstack/unstract, run ./run-platform.sh, open the local Unstract frontend, configure provider credentials, define the extraction prompt/schema in Prompt Studio, then deploy as a REST API or ETL pipeline using the official Unstract docs.
Author
Zipstack
Publisher
Open Source
Last updated
Jun 16, 2026
Quick brief

Use Unstract when an operator needs a repeatable document extraction workflow: define the fields to extract in Prompt Studio, connect an LLM provider and document source, then expose the extraction as a REST API, ETL pipeline, MCP server, or automation node that returns structured JSON for downstream systems.

How it works

What this skill actually does

This is skill-shaped because the agent/operator job is bounded: build, test, and run a document-to-JSON extraction workflow for known document families, then hand off validated structured output to an API consumer, warehouse, or workflow runner. Invoke it when documents need repeatable schema extraction and deployment, not when the user only wants to browse the Unstract product or use a generic LLM SDK.

Inputs and prerequisites: Unstract platform; Docker and Docker Compose for local runs; Git; access to the target documents; configured LLM provider credentials; optional API deployment, ETL, MCP, or n8n integration.

Setup notes: Clone https://github.com/Zipstack/unstract, run ./run-platform.sh for a local platform, open the local frontend, configure provider credentials, define the extraction prompt/schema in Prompt Studio, and deploy the workflow as an API or ETL pipeline according to the upstream documentation.

Source and verification boundary: use https://github.com/Zipstack/unstract and https://docs.unstract.com as the canonical references before running the workflow; keep extraction prompts, schema changes, provider credentials, and produced JSON reviewable against those upstream docs.

Framework fit: publish this as a Multi-Framework workflow because upstream intentionally exposes the same bounded extraction workflow through REST APIs, ETL pipelines, MCP, and automation integrations rather than requiring one agent runtime.