Open, local-first behavioral evidence and evaluation infrastructure for AI systems, benchmarks, and agents.

1 stars 1 forks 1 watchers TypeScript MIT License
ai ai-agents ai-evaluation behavioral-evaluation benchmarking evidence llm local-first observability open-source reproducibility research-infrastructure
3 Open Issues Need Help Last updated: Sep 1, 2026

Open Issues Need Help

View All on GitHub

Open, local-first behavioral evidence and evaluation infrastructure for AI systems, benchmarks, and agents.

TypeScript
#ai#ai-agents#ai-evaluation#behavioral-evaluation#benchmarking#evidence#llm#local-first#observability#open-source#reproducibility#research-infrastructure

Open, local-first behavioral evidence and evaluation infrastructure for AI systems, benchmarks, and agents.

TypeScript
#ai#ai-agents#ai-evaluation#behavioral-evaluation#benchmarking#evidence#llm#local-first#observability#open-source#reproducibility#research-infrastructure

Open, local-first behavioral evidence and evaluation infrastructure for AI systems, benchmarks, and agents.

TypeScript
#ai#ai-agents#ai-evaluation#behavioral-evaluation#benchmarking#evidence#llm#local-first#observability#open-source#reproducibility#research-infrastructure