The robust European language model benchmark.

european evaluation-framework leaderboard llm nlp
55 Open Issues Need Help Last updated: Jul 27, 2026

Open Issues Need Help

View All on GitHub

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted good first issue leaderboards

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted leaderboards

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted good first issue leaderboards

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
benchmark dataset request help wanted

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
benchmark dataset request help wanted

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
benchmark dataset request help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
benchmark dataset request help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
documentation good first issue minor

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted good first issue minor

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp

AI Summary: The user is encountering an error when running `euroeval` with a GPT-2 model, stating that the `generative` extra is missing, despite explicitly including `euroeval[generative]` in their `uv` script's dependencies. This suggests an issue where the `generative` extra is not being correctly recognized or installed, potentially due to a packaging problem within `euroeval` or a flaw in its internal dependency check.

Complexity: 3/5
documentation help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted new language

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
benchmark dataset request help wanted

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp

AI Summary: This GitHub issue is a request to add the 'Exam-et' dataset, an Estonian multiple-choice exam dataset available on Hugging Face, to be used as an Estonian knowledge benchmark. The request provides a direct link and a brief description of the dataset's content and purpose.

Complexity: 1/5
benchmark dataset request help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
Support Polish 11 months ago
benchmark dataset request help wanted

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
benchmark dataset request help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
Support Lithuanian 11 months ago
benchmark dataset request help wanted

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp

AI Summary: This GitHub issue requests the addition of the `copa-lv` dataset, a Latvian translation and post-edited version of the English COPA common-sense reasoning dataset. The request specifies using machine-translated train and validation splits, and post-edited test splits, all available at the provided GitHub link.

Complexity: 2/5
benchmark dataset request help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
documentation help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp

AI Summary: This GitHub issue reports several grammatical flaws and awkward direct translations in Icelandic prompts used across various datasets, including 'Hotter and Colder', 'MÍM-GOLD-NER', and 'ScaLA-is'. The author provides specific suggestions for improving word choice, verb conjugation, and overall naturalness of the prompts to make them more idiomatic Icelandic.

Complexity: 1/5
help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
Support Latvian 12 months ago
benchmark dataset request help wanted

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
Support Estonian 12 months ago
benchmark dataset request help wanted

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp
help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp

AI Summary: The task is to improve the EuroEval benchmark by replacing the current manual evaluation method for encoder models on multiple-choice tasks with the Hugging Face `AutoModelForMultipleChoice` class. This involves modifying the existing code to utilize this class and comparing the results to the previous method to assess the impact of the change.

Complexity: 4/5
help wanted

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp

AI Summary: Implement a new separator (#) for specifying generation arguments in the EuroEval benchmark, alongside the existing @ separator for LiteLLM models. The new separator should work for all model types, while maintaining backward compatibility with the @ separator (with a deprecation warning). The implementation must handle any order of @ and # separators.

Complexity: 4/5
help wanted good first issue

The robust European language model benchmark.

Python
#european#evaluation-framework#leaderboard#llm#nlp