Agent-12 — the local agent leaderboard. Filesystem-judged agent tasks on consumer hardware, with reference-validated judges.

agent-benchmark apple-silicon benchmark coding-agent leaderboard llm-evaluation local-llm mlx ollama
1 Open Issue Need Help Last updated: Sep 19, 2026

Open Issues Need Help

View All on GitHub

Agent-12 — the local agent leaderboard. Filesystem-judged agent tasks on consumer hardware, with reference-validated judges.

Python
#agent-benchmark#apple-silicon#benchmark#coding-agent#leaderboard#llm-evaluation#local-llm#mlx#ollama