Open Issues Need Help
View All on GitHub A suite can clear --min-score with the prompt-injection case failing about 1 month ago
bug good first issue
Evaluation harness for Amazon Bedrock: run a fixed case suite across models and prompts, and compare scores, cost and latency between runs.
Python
#amazon-bedrock#aws#eval-harness#genai#llm-evaluation#llmops
score exits 0 even when every case fails, so it cannot gate CI about 1 month ago
bug good first issue
Evaluation harness for Amazon Bedrock: run a fixed case suite across models and prompts, and compare scores, cost and latency between runs.
Python
#amazon-bedrock#aws#eval-harness#genai#llm-evaluation#llmops