Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

9 stars 2 forks 9 watchers Python Apache License 2.0
cli llm metrics rag rag-evaluation
37 Open Issues Need Help Last updated: Jul 13, 2026

Open Issues Need Help

View All on GitHub
enhancement good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
enhancement good first issue in-progress

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
documentation good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
documentation good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
documentation enhancement good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
documentation enhancement good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
documentation enhancement good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
documentation enhancement good first issue help wanted

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
enhancement good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
enhancement good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
enhancement good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
enhancement good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
enhancement good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
documentation good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
enhancement good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
enhancement good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation
documentation good first issue

Open-source CLI framework for evaluating RAG systems and AI agents with standardized metrics, experiment tracking, and reproducible reports.

Python
#cli#llm#metrics#rag#rag-evaluation