Open Issues Need Help
View All on GitHub Help design the first fraud-prevention benchmark scenarios about 8 hours ago
help wanted expert review
Measuring harmful refusals in AI models
Python
#ai-benchmark#cybersecurity#false-refusal#llm-evaluation#responsible-ai
Test the first-run guide on a clean macOS or Linux environment about 8 hours ago
documentation good first issue
Measuring harmful refusals in AI models
Python
#ai-benchmark#cybersecurity#false-refusal#llm-evaluation#responsible-ai