Open Issues Need Help
View All on GitHubAn evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.
An evidence gate for AI coding agents, and a benchmark that grades with a test exit code instead of another model. 424 runs, 6 predictions filed before the data, 2 of them lost. It makes agents paste real command output ~30x more often — and does not make them any less wrong. Both numbers are on the front page.