Benchmark enterprise RAG agents with EnterpriseRAG-Bench
Use EnterpriseRAG-Bench to evaluate an enterprise RAG or knowledge-agent system against a realistic synthetic company corpus with answer, recall, and comparative scoring.
npx skills add agentskillexchange/skills --skill benchmark-enterprise-rag-agents-with-enterpriserag-bench
EnterpriseRAG-Bench provides a repeatable benchmark for retrieval and answer quality on synthetic internal-company data. Use it when an agent team needs to test whether a RAG system can answer questions over Slack, Gmail, Drive, GitHub, Jira, Confluence, CRM, and meeting-style documents before relying on it in production. The workflow is to download the benchmark corpus and questions, run the user’s RAG or knowledge agent over the questions, write answers and retrieved document IDs as JSONL, then run the provided metrics or comparative evaluation scripts. This is scoped to evaluation of enterprise RAG behavior and dataset generation; it is not an Onyx product card, a generic RAG framework listing, or an open-ended retrieval tutorial.
What this skill actually does
Inputs and prerequisites: Python 3.10+, EnterpriseRAG-Bench dataset and questions, a RAG or knowledge-agent system under test, OpenAI or Anthropic compatible LLM credentials for evaluation.
Setup notes: Clone https://github.com/onyx-dot-app/EnterpriseRAG-Bench, install Python dependencies with pip install -r requirements.txt, set LLM_PROVIDER and LLM_API_KEY, download the benchmark dataset from the latest GitHub release or Hugging Face, write system outputs to answer_evaluation/answers.jsonl, then run python -m src.scripts.answer_evaluation.metrics_based_eval –answers-file answer_evaluation/answers.jsonl.
Source and verification boundary: use https://github.com/onyx-dot-app/EnterpriseRAG-Bench/blob/main/quickstart.md as the canonical reference before running the workflow; keep commands, API calls, CLI usage, and generated outputs reviewable against that upstream source.
Framework fit: publish this as a Multi-Framework workflow only when the operator can invoke the documented toolchain directly, rather than treating the upstream project as a generic product listing.