An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

42 Open Issues Need Help Last updated: Jul 8, 2026

Open Issues Need Help

View All on GitHub
help wanted adapter priority: medium

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted adapter priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted adapter priority: medium

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted statistics priority: medium

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted statistics priority: high

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted adapter priority: medium

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted adapter priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted adapter priority: medium

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted adapter priority: medium

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted adapter priority: medium

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted adapter priority: medium

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted statistics priority: medium

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted priority: medium

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted priority: high

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
documentation help wanted good first issue priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted good first issue priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted good first issue priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted good first issue priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
documentation help wanted good first issue priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
documentation help wanted good first issue priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted good first issue priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted good first issue priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted good first issue priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted good first issue priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted good first issue adapter priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted priority: low

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted statistics priority: medium

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted adapter priority: medium

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted adapter

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted adapter

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted statistics

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted statistics

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted statistics

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement help wanted

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement good first issue

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
enhancement good first issue

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted good first issue

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
help wanted good first issue

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python
good first issue

An auditor for LLM evaluations — tells you whether you can trust your eval results, not just what the score is.

Python