Skill Detail

Run local document RAG with citations over MCP using Haiku.RAG

Index local or self-hosted documents, search them with hybrid and multimodal retrieval, and answer agent questions through an MCP server with citations.

Data Extraction & TransformationMCP
Data Extraction & Transformation MCP Security Reviewed
โญ 581 GitHub stars
COPY SKILL INSTRUCTIONS (OPTIONAL)
npx skills add agentskillexchange/skills --skill run-local-document-rag-with-citations-over-mcp-using-haiku-rag Copy
Uses the third-party skills CLI, not an ASE-owned installer. Check your agent’s compatibility. This copies instructions; complete the upstream tool setup below separately.
At a glance
Tools required
Python 3.12+, haiku.rag or haiku.rag-slim, an embedding provider such as Ollama/OpenAI/VoyageAI/Cohere/LM Studio/vLLM, and an MCP-compatible client
Install & setup
Install with pip install haiku.rag or uv pip install haiku.rag, index documents with commands such as haiku-rag add-src paper.pdf, then expose the knowledge base to an MCP client with haiku-rag mcp --stdio. Use haiku-rag --read-only mcp --stdio when the agent should only search and ask questions.
Author
Yiorgis Gozadinos
Publisher
Individual
Last updated
Aug 18, 2026
Quick brief

Use Haiku.RAG when an agent needs a local-first retrieval layer over private documents rather than a hosted knowledge product. The operator installs the Python package, indexes files or URLs into an embedded LanceDB store, starts the MCP server, and lets an MCP-compatible assistant search documents, ask cited questions, analyze document collections, and optionally ingest new sources through bounded tools. This is strongest for evidence-backed PDF, document, and image-grounded research workflows where citations, local storage, read-only mode, and repeatable retrieval behavior matter. The scope boundary is document retrieval and cited QA over a configured knowledge base; it is not a generic RAG framework listing, vector database listing, or product card.

How it works

What this skill actually does

Inputs and prerequisites: Python 3.12+, haiku.rag or haiku.rag-slim, an embedding provider such as Ollama/OpenAI/VoyageAI/Cohere/LM Studio/vLLM, and an MCP-compatible client.

Setup notes: Install with pip install haiku.rag or uv pip install haiku.rag , index documents with commands such as haiku-rag add-src paper.pdf , then expose the knowledge base to an MCP client with haiku-rag mcp –stdio . Use haiku-rag –read-only mcp –stdio when the agent should only search and ask questions.

Source and verification boundary: use https://ggozad.github.io/haiku.rag/ as the canonical reference before running the workflow; keep commands, API calls, CLI usage, and generated outputs reviewable against that upstream source.

Framework fit: publish this as a MCP workflow only when the operator can invoke the documented toolchain directly, rather than treating the upstream project as a generic product listing.