Diagnose whether a low-kappa LLM judge panel fails from item ambiguity or rubric underspecification.

ai-evaluation bias-detection cohen-kappa evaluation-metrics fleiss-kappa inter-rater-reliability llm-as-a-judge llm-evaluation python reliability
1 Open Issue Need Help Last updated: Aug 5, 2026

Open Issues Need Help

View All on GitHub

Diagnose whether a low-kappa LLM judge panel fails from item ambiguity or rubric underspecification.

Python
#ai-evaluation#bias-detection#cohen-kappa#evaluation-metrics#fleiss-kappa#inter-rater-reliability#llm-as-a-judge#llm-evaluation#python#reliability