Open Issues Need Help
View All on GitHub enhancement good first issue product
Evidence-backed verification for action-taking AI agents. Compile the task into a contract once, then verify every run against real state and trajectory evidence with zero model calls.
TypeScript
#agent-evaluation#agentic-ai#ai-agents#ai-safety#anthropic#benchmark#claude#deterministic#evals#llm#llm-evaluation#typescript#verification
bug good first issue
Evidence-backed verification for action-taking AI agents. Compile the task into a contract once, then verify every run against real state and trajectory evidence with zero model calls.
TypeScript
#agent-evaluation#agentic-ai#ai-agents#ai-safety#anthropic#benchmark#claude#deterministic#evals#llm#llm-evaluation#typescript#verification