Evidence-backed verification for action-taking AI agents. Compile the task into a contract once, then verify every run against real state and trajectory evidence with zero model calls.

0 stars 0 forks 0 watchers TypeScript MIT License
agent-evaluation agentic-ai ai-agents ai-safety anthropic benchmark claude deterministic evals llm llm-evaluation typescript verification
2 Open Issues Need Help Last updated: Sep 8, 2026

Open Issues Need Help

View All on GitHub
enhancement good first issue product

Evidence-backed verification for action-taking AI agents. Compile the task into a contract once, then verify every run against real state and trajectory evidence with zero model calls.

TypeScript
#agent-evaluation#agentic-ai#ai-agents#ai-safety#anthropic#benchmark#claude#deterministic#evals#llm#llm-evaluation#typescript#verification

Evidence-backed verification for action-taking AI agents. Compile the task into a contract once, then verify every run against real state and trajectory evidence with zero model calls.

TypeScript
#agent-evaluation#agentic-ai#ai-agents#ai-safety#anthropic#benchmark#claude#deterministic#evals#llm#llm-evaluation#typescript#verification