Benchmarking whether LLM agents admit tool failure instead of hallucinating success — 100 tasks, 1,800 responses, deterministic simulator, reproducible harness.

ai-agents ai-evaluation benchmark guardrails hallucination llm reproducibility tool-use
2 Open Issues Need Help Last updated: Sep 11, 2026

Open Issues Need Help

View All on GitHub

Benchmarking whether LLM agents admit tool failure instead of hallucinating success — 100 tasks, 1,800 responses, deterministic simulator, reproducible harness.

Python
#ai-agents#ai-evaluation#benchmark#guardrails#hallucination#llm#reproducibility#tool-use
help wanted model-result replication

Benchmarking whether LLM agents admit tool failure instead of hallucinating success — 100 tasks, 1,800 responses, deterministic simulator, reproducible harness.

Python
#ai-agents#ai-evaluation#benchmark#guardrails#hallucination#llm#reproducibility#tool-use