How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
VAKRA (eValuating API and Knowledge Retrieval Agents using multi-hop, multi-source dialogues) is a tool-grounded, executable benchmark designed to evaluate how well AI agents reason end-to-end in enterprise-like settings.
The main features of ibm/vakra are: Benchmarks and Evaluation.
Open-source alternatives to ibm/vakra include: graphrag-bench/graphrag-benchmark — 🧩Task Examples. petergpt/bullshit-benchmark — BullshitBench v2.