Skill Detail

Benchmark enterprise RAG agents with EnterpriseRAG-Bench

Use EnterpriseRAG-Bench to evaluate an enterprise RAG or knowledge-agent system against a realistic synthetic company corpus with answer, recall, and comparative scoring.

Security & VerificationMulti-Framework
Security & Verification Multi-Framework Security Reviewed
⭐ 562 GitHub stars
COPY SKILL INSTRUCTIONS (OPTIONAL)
npx skills add agentskillexchange/skills --skill benchmark-enterprise-rag-agents-with-enterpriserag-bench Copy
Uses the third-party skills CLI, not an ASE-owned installer. Check your agent’s compatibility. This copies instructions; complete the upstream tool setup below separately.
At a glance
Tools required
Python 3.10+, EnterpriseRAG-Bench dataset and questions, a RAG or knowledge-agent system under test, OpenAI or Anthropic compatible LLM credentials for evaluation
Install & setup
Clone https://github.com/onyx-dot-app/EnterpriseRAG-Bench, install Python dependencies with pip install -r requirements.txt, set LLM_PROVIDER and LLM_API_KEY, download the benchmark dataset from the latest GitHub release or Hugging Face, write system outputs to answer_evaluation/answers.jsonl, then run python -m src.scripts.answer_evaluation.metrics_based_eval –answers-file answer_evaluation/answers.jsonl.
Author
Onyx
Publisher
Organization
Last updated
Sep 20, 2026
Quick brief

EnterpriseRAG-Bench provides a repeatable benchmark for retrieval and answer quality on synthetic internal-company data. Use it when an agent team needs to test whether a RAG system can answer questions over Slack, Gmail, Drive, GitHub, Jira, Confluence, CRM, and meeting-style documents before relying on it in production. The workflow is to download the benchmark corpus and questions, run the user’s RAG or knowledge agent over the questions, write answers and retrieved document IDs as JSONL, then run the provided metrics or comparative evaluation scripts. This is scoped to evaluation of enterprise RAG behavior and dataset generation; it is not an Onyx product card, a generic RAG framework listing, or an open-ended retrieval tutorial.

How it works

What this skill actually does

Inputs and prerequisites: Python 3.10+, EnterpriseRAG-Bench dataset and questions, a RAG or knowledge-agent system under test, OpenAI or Anthropic compatible LLM credentials for evaluation.

Setup notes: Clone https://github.com/onyx-dot-app/EnterpriseRAG-Bench, install Python dependencies with pip install -r requirements.txt, set LLM_PROVIDER and LLM_API_KEY, download the benchmark dataset from the latest GitHub release or Hugging Face, write system outputs to answer_evaluation/answers.jsonl, then run python -m src.scripts.answer_evaluation.metrics_based_eval –answers-file answer_evaluation/answers.jsonl.

Source and verification boundary: use https://github.com/onyx-dot-app/EnterpriseRAG-Bench/blob/main/quickstart.md as the canonical reference before running the workflow; keep commands, API calls, CLI usage, and generated outputs reviewable against that upstream source.

Framework fit: publish this as a Multi-Framework workflow only when the operator can invoke the documented toolchain directly, rather than treating the upstream project as a generic product listing.