Open Issues Need Help
View All on GitHub help wanted human-validation
Benchmarking whether LLM agents admit tool failure instead of hallucinating success — 100 tasks, 1,800 responses, deterministic simulator, reproducible harness.
Python
#ai-agents#ai-evaluation#benchmark#guardrails#hallucination#llm#reproducibility#tool-use
help wanted model-result replication
Benchmarking whether LLM agents admit tool failure instead of hallucinating success — 100 tasks, 1,800 responses, deterministic simulator, reproducible harness.
Python
#ai-agents#ai-evaluation#benchmark#guardrails#hallucination#llm#reproducibility#tool-use