An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

agent-skills ai-agents ai-coding benchmark claude-code developer-tools evaluation llm prompt-engineering python reproducible-research sycophancy
22 Open Issues Need Help Last updated: Sep 2, 2026

Open Issues Need Help

View All on GitHub
good first issue adapter

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
good first issue adapter

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
good first issue adapter

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
good first issue adapter

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
good first issue adapter

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
good first issue adapter

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
help wanted result

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
help wanted result

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
help wanted result

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
help wanted result

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
help wanted result

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
help wanted result

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
help wanted result

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
help wanted result

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
help wanted result

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
help wanted result

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
help wanted result

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy
help wanted result

An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.

Python
#agent-skills#ai-agents#ai-coding#benchmark#claude-code#developer-tools#evaluation#llm#prompt-engineering#python#reproducible-research#sycophancy