Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

ai ai-evaluation artificial-intelligence benchmarking developer-tools evaluation fastapi g-eval llm llm-as-a-judge llm-evaluation open-source postgresql python rag rag-evaluation react redis typescript
27 Open Issues Need Help Last updated: Jul 28, 2026

Open Issues Need Help

View All on GitHub
help wanted good first issue frontend

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted good first issue testing

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted frontend testing

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted backend

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted backend

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted frontend performance

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted backend

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted good first issue

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted good first issue frontend

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted good first issue frontend

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted good first issue backend

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted good first issue backend

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
documentation good first issue

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
good first issue backend testing

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted good first issue backend testing

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
help wanted good first issue backend

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript
good first issue frontend

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript

Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.

Python
#ai#ai-evaluation#artificial-intelligence#benchmarking#developer-tools#evaluation#fastapi#g-eval#llm#llm-as-a-judge#llm-evaluation#open-source#postgresql#python#rag#rag-evaluation#react#redis#typescript