Run self-hosted media research and RAG workflows with tldw_server
Use tldw_server when an agent workflow needs to ingest video, audio, documents, and web pages into a local research layer with RAG, evals, OpenAI-compatible APIs, and MCP access.
npx skills add agentskillexchange/skills --skill run-self-hosted-media-research-and-rag-workflows-with-tldw-server
Use tldw_server when an operator needs a self-hosted research workspace that turns media and documents into searchable agent context. The workflow is to deploy the FastAPI service, ingest approved video, audio, PDFs, documents, websites, or browser-clipped pages, run transcription, OCR, chunking, embeddings, hybrid search, and RAG, then let trusted agents query the resulting workspace through documented APIs or the MCP server. Invoke this instead of using a hosted note-taking or video-summary product when privacy, local model routing, repeatable ingestion, reviewable citations, or a consistent OpenAI-compatible API surface matters.
What this skill actually does
The scope boundary is the research and context layer: collect approved source material, process it into a local knowledge base, tune retrieval or evaluation recipes, and expose bounded search/chat tools to an agent. It is not a generic media app listing, not a replacement for checking source rights, and not a reason to give agents unrestricted filesystem or network access. Keep deployments single-user or properly hardened unless the upstream multi-user profile and auth controls have been configured.
Inputs and prerequisites: Python 3.10+, ffmpeg, Docker or local Python setup, storage for media and embeddings, optional local or hosted LLM/STT/TTS providers, and an MCP-capable client when using the MCP surface.
Setup notes: Follow the upstream getting-started profile, run the prerequisite check, choose Docker single-user or local single-user setup, configure auth secrets for any shared deployment, verify the health endpoint, and test ingestion on a small approved document or media file before connecting agent clients.
Source and verification boundary: use https://github.com/rmusser01/tldw_server and https://tldwproject.com as the canonical references before running the workflow; keep API calls, MCP tools, ingestion jobs, and generated outputs reviewable against that upstream source.
Framework fit: publish this as an MCP workflow because upstream documents an MCP server and MCP governance surfaces, while the same research layer can still be operated through its API and web UI.