eval_framework.benchmarks package¶
Submodules¶
eval_framework.benchmarks.arc module¶
ARC (AI2 Reasoning Challenge): https://huggingface.co/datasets/allenai/ai2_arc
Grade-school science multiple-choice questions, split into the ARC-Easy and ARC-Challenge subjects. The base task scores the full answer text (cloze); the OLMES variant shows the options as space-prefixed lettered choices and scores the letters; the IDK variant lets the model abstain with “I do not know”.
- final class eval_framework.benchmarks.arc.ArcReader[source]¶
Bases:
ChoiceReaderReads an ARC item: the question and its answer options, with the correct one at
answerKey— a letter (A–E) or a 1-based number, both normalised to a 0-based index.
- eval_framework.benchmarks.arc.arc(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.arc.arc_idk(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.arc.arc_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.arc_de module¶
ARC German (ARC-DE).
- final class eval_framework.benchmarks.arc_de.ArcDeReader[source]¶
Bases:
ChoiceReaderReads an ARC-DE row into choice fields: the German question, its answer texts, and the correct index.
answerKeyarrives as either a 1-based number or a letter;answer_key_to_indexnormalises both, which also frees the reader from caring how many answers a given row offers.
- eval_framework.benchmarks.arc_de.arc_de(dataset=None)[source]¶
ARC-DE as cloze/ranked classification.
https://huggingface.co/datasets/LeoLM/ArcChallenge_de
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.arc_ellamind module¶
German ARC (EllaMind): https://huggingface.co/datasets/ellamind/arc-multilingual
Grade-school science questions in German. Every slice loads the single German config (deu); the
ARC-Easy and ARC-Challenge subjects are the rows of that config tagged by the arc_config column. Cloze
scores the full answer text, MC the letter labels, BPB only the ground-truth answer’s bits-per-byte.
- final class eval_framework.benchmarks.arc_ellamind.ArcEllamindReader[source]¶
Bases:
ChoiceReaderReads a German ARC item: the question and its answer options, with the correct one at
answer_key(a letter A–E or a 1-based number).
- eval_framework.benchmarks.arc_ellamind.arc_ellamind_bpb_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.arc_ellamind.arc_ellamind_cloze_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.arc_ellamind.arc_ellamind_mc_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.bigcodebench module¶
BigCodeBench: https://huggingface.co/datasets/bigcode/bigcodebench
The model completes a self-contained function; the CodeExecutionPassAtOne metric merges the generated
snippet with the problem’s unittest harness (via functions carried, serialized, in the sample’s context) and
runs it. Only the OLMES 3-shot variant is registered.
- eval_framework.benchmarks.bigcodebench.bigcodebench_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.copa module¶
COPA (Choice of Plausible Alternatives): https://huggingface.co/datasets/aps/super_glue
Causal-reasoning items: a premise, a cause/effect cue, and two alternatives. The registered variant is OLMES-style: the premise is recast as a sentence stem (its final period replaced by the causal connector), the two alternatives are shown as space-prefixed lettered options (” A. …”), and the model is scored over the letter labels. The prompt carries no assistant cue — the options directly continue the stem.
- final class eval_framework.benchmarks.copa.CopaReader[source]¶
Bases:
ChoiceReaderReads a COPA item: the premise recast as a stem (final period → connector) and its two alternatives, each lower-cased to continue the stem.
- eval_framework.benchmarks.copa.copa_mc_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.cot module¶
Shared chain-of-thought scaffolding for multiple-choice benchmarks.
A CoT variant asks the model to reason freely and conclude with a stated answer letter, which is pulled back
out at scoring time (the injected ExtractFromCompletion). The prompt surface — an optional preamble, the body,
and any inert scored candidates — is injected per benchmark; the CoT contract is fixed here: one free-form
sample, no assistant cue, the bare answer letter as ground truth, scored by accuracy.
- final class eval_framework.benchmarks.cot.Cot(reader, *, build_prompt, preamble=None, candidates=None)[source]¶
Bases:
EvalKindMultiple-choice chain-of-thought (see the module docstring).
build_promptrenders the body,preamblean optional subject-templated top line, andcandidatesany inert scored letters kept for faithful parity with a loglikelihood baseline (free-form scoring ignores them).- Parameters:
reader (ChoiceReader)
build_prompt (Callable[[str, list[str]], str])
preamble (Callable[[str], str] | None)
candidates (Callable[[list[str]], list[str]] | None)
- messages(body, *, fewshot, subject_label)[source]¶
The full message list for one sample — typically
assemble_messageswith the kind’s own preamble and/or system prompt.- Return type:
list[Message]- Parameters:
body (SampleBody)
fewshot (list[FewshotExample])
subject_label (str)
- metrics()[source]¶
The metrics this kind is scored with.
- Return type:
list[type[BaseMetric]]
- samples(item)[source]¶
The scored sample(s) for one eval item — one for most kinds, more when a kind fans out.
- Return type:
list[SampleBody]- Parameters:
item (dict[str, Any])
- eval_framework.benchmarks.cot.tulu3_cot_prompt(raw_question, choices)[source]¶
The parenthesised-option CoT body shared by MMLU-Pro and GPQA.
Reasoning prompt from Figure 44 of the Tülu 3 paper: https://arxiv.org/pdf/2411.15124
- Return type:
str- Parameters:
raw_question (str)
choices (list[str])
- eval_framework.benchmarks.cot.tulu_answer()[source]¶
Extracts the parenthesised letter that
tulu3_cot_promptasks the model to conclude with —"Therefore, the answer is (X)". Kept here beside the prompt because both encode the same(X)format. The accepted letters are the fixed A–J the prompt’s"(A), (B), ..., (E), etc."implies.- Return type:
- eval_framework.benchmarks.cot.tulu_answer_v2(n_options)[source]¶
Lenient variant of
tulu_answer: no required"Therefore,", optional colon and parentheses, case-insensitive, taking the last match. Only the accepted letter range is benchmark-specific, so it is built fromn_options.- Return type:
- Parameters:
n_options (int)
eval_framework.benchmarks.csqa module¶
CommonsenseQA (English): https://huggingface.co/datasets/tau/commonsense_qa
Five-way multiple-choice commonsense questions. The registered variant is OLMES-style: the options are shown as space-prefixed lettered choices (” A. …”) and the model is scored over the letter labels.
- final class eval_framework.benchmarks.csqa.CommonsenseQaReader[source]¶
Bases:
ChoiceReaderReads a CommonsenseQA item: the question and its labelled options, with the correct one at
answerKey.
- eval_framework.benchmarks.csqa.csqa_mc_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.csqa_ellamind module¶
German CommonsenseQA (EllaMind) tasks.
https://huggingface.co/datasets/ellamind/csqa-multilingual
CSQA supplies separate easy and hard distractors.
- final class eval_framework.benchmarks.csqa_ellamind.CsqaReader(distractor_level)[source]¶
Bases:
ChoiceReaderReads a CSQA item: easy/hard distractor lists per level, shuffled in with the correct answer.
- Parameters:
distractor_level (Literal['easy', 'hard'])
- eval_framework.benchmarks.csqa_ellamind.csqa_ellamind_bpb_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.csqa_ellamind.csqa_ellamind_cloze_easy_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.csqa_ellamind.csqa_ellamind_cloze_hard_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.csqa_ellamind.csqa_ellamind_mc_easy_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.csqa_ellamind.csqa_ellamind_mc_hard_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.drop module¶
DROP (Discrete Reasoning Over Paragraphs): https://huggingface.co/datasets/EleutherAI/drop
A passage and a question over it. DropCompletion_OLMES generates a free-form answer scored by DROP F1 /
exact match (EleutherAI/drop); DropMC_OLMES scores labelled candidate answers by loglikelihood
(allenai/drop-gen2mc).
- eval_framework.benchmarks.drop.drop_completion_olmes(dataset=None)[source]¶
OLMES: a reading-comprehension preamble, few-shot from the train split, and a 100-token answer budget.
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.drop.drop_mc_olmes(dataset=None)[source]¶
OLMES lays out the options with a leading space (” A. …”).
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.global_mmlu module¶
Global-MMLU: https://huggingface.co/datasets/CohereLabs/Global-MMLU
MMLU translated into many languages; we evaluate French, German, Spanish, Italian, Portuguese and Arabic.
- eval_framework.benchmarks.global_mmlu.global_mmlu(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.global_mmlu.global_mmlu_german(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.goldenswag module¶
GoldenSwag: https://huggingface.co/datasets/PleIAs/GoldenSwag
A curated subset of HellaSwag, read identically (sentence completion, scored as cloze — see
HellaswagReader). GoldenSwag_IDK additionally lets the model abstain with an “I do not know” answer.
- eval_framework.benchmarks.goldenswag.goldenswag(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.goldenswag.goldenswag_idk(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.gpqa module¶
GPQA (Graduate-level Google-Proof Q&A): https://huggingface.co/datasets/Idavidrein/gpqa
Gated, expert-written multiple-choice science questions. Each item has one Correct Answer and three
Incorrect Answer N distractors; the reader shuffles them together, seeded from the option texts so an
item’s option order is stable across runs. The registered variants:
GPQA_OLMES: OLMES-style loglikelihood over the fullgpqa_extendedconfig.GPQA_DIAMOND_COT/_V2: chain-of-thought completion over the hardergpqa_diamondconfig; they share one prompt and differ only in how the concluding answer letter is extracted.
One question carries a raw sequence far too long for the prompt budget and is dropped from every config.
- final class eval_framework.benchmarks.gpqa.GpqaReader[source]¶
Bases:
ChoiceReaderReads a GPQA item: three
Incorrect Answer Ndistractors and theCorrect Answer, shuffled together. The shuffle is seeded from the (preprocessed) option texts, so the same item always yields the same option order regardless of the order in which items are evaluated.
- eval_framework.benchmarks.gpqa.gpqa_diamond_cot(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.gpqa.gpqa_diamond_cot_v2(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.gpqa.gpqa_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.gpqa_ellamind module¶
German GPQA (Graduate-level Professional QA, EllaMind) tasks.
https://huggingface.co/datasets/ellamind/gpqa-multilingual
GPQA uses a single distractor set (incorrect_answers). The diamond variants restrict evaluation to
the diamond subset — the 198 hardest questions (is_diamond) from the original GPQA-Diamond benchmark.
The COT variant has the model reason in German and conclude with the answer letter, which is
leniently regex-extracted from the generation (free-form completion, 0-shot).
- final class eval_framework.benchmarks.gpqa_ellamind.GpqaReader[source]¶
Bases:
ChoiceReaderReads a GPQA item: a single
incorrect_answersdistractor set, shuffled in with the correct answer.
- eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_bpb_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_cloze_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_diamond_bpb_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_diamond_cloze_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_diamond_cot_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_diamond_mc_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_mc_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.gpqa_ellamind.tulu3_cot_prompt_de(raw_question, choices)[source]¶
German translation of
tulu3_cot_prompt(Figure 44 of the Tülu 3 paper, https://arxiv.org/pdf/2411.15124): the model reasons briefly and concludes with “Daher ist die Antwort (X)”. The answer format is stated twice — once before the question and once as a reminder after it.- Return type:
str- Parameters:
raw_question (str)
choices (list[str])
- eval_framework.benchmarks.gpqa_ellamind.tulu_answer_de()[source]¶
Extracts the letter that
tulu3_cot_prompt_deasks the model to conclude with, as leniently astulu_answer_v2: the last match wins, the parentheses are optional, and there is no stop sequence to cut the generation short. It also accepts the English “answer is X”, because a model prompted in German often still concludes in English. The match is anchored on the answer phrase, so a bare “Antwort D” in the reasoning does not count. GPQA always has four options, so only A–D are accepted.- Return type:
eval_framework.benchmarks.gsm8k module¶
GSM8K: https://huggingface.co/datasets/openai/gsm8k
Grade-school math word problems. Both registered variants use eight fixed, hand-written exemplars (not sampled from the dataset) and OLMES answer normalisation:
GSM8K_OLMES: free-form completion, scored on the final number of the generation.GSM8KBPB: bits-per-byte of the single normalised gold solution (one forward pass).
- eval_framework.benchmarks.gsm8k.clean_short_answer(continuation)[source]¶
Reduce a solution to its final number, commas removed — the OLMES short-answer form.
- Return type:
str- Parameters:
continuation (str)
- eval_framework.benchmarks.gsm8k.gsm8k_bpb(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.gsm8k.gsm8k_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.gsm8k_ellamind module¶
German GSM8K (EllaMind): https://huggingface.co/datasets/ellamind/gsm8k-platinum-multilingual
The German counterpart of GSM8K (Frage: / Antwort:), with few-shot demonstrations sampled from the
test split. Each item carries a worked solution and a final_answer; a demonstration ends with the
German final-answer line "Daher ist die Antwort N.". Two registered variants:
GSM8K_Ellamind_DE_Platinum: free-form completion, scored on the final integer of the generation.GSM8K_Ellamind_DE_BPB_Platinum: bits-per-byte of the single gold solution.
- eval_framework.benchmarks.gsm8k_ellamind.gsm8k_ellamind_de_bpb_platinum(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.gsm8k_ellamind.gsm8k_ellamind_de_platinum(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.hellaswag module¶
HellaSwag: https://huggingface.co/datasets/Rowan/hellaswag
Sentence completion: each item gives an activity label and a context, and the model scores which ending
best continues it. Scored as cloze — the prompt is "{activity}: {context}" and the candidates are the
full endings. HellaSwag evaluates on the validation split; HellaSwag_OLMES on the (larger) train
split, matching OLMES.
- final class eval_framework.benchmarks.hellaswag.HellaswagReader[source]¶
Bases:
ChoiceReaderReads a HellaSwag item: the shown text is
"{activity}: {context}"(context isctx_aplus a capitalisedctx_b), and each choice is a candidate ending. Markup is stripped from all shown text.
- eval_framework.benchmarks.hellaswag.hellaswag(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.hellaswag.hellaswag_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.hellaswag_ellamind module¶
German HellaSwag (EllaMind) tasks.
https://huggingface.co/datasets/ellamind/hellaswag-multilingual
HellaSwag is a sentence-completion task: the prompt is a partial sentence ("{activity}: {context}")
and the model scores full-sentence endings. There is no natural MC variant. HellaSwag supplies separate
easy and hard distractors.
- final class eval_framework.benchmarks.hellaswag_ellamind.HellaswagReader(distractor_level)[source]¶
Bases:
ChoiceReaderReads a HellaSwag item: the partial sentence
"{activity}: {context}", with the easy/hard full-sentence endings for the level shuffled in with the correct ending.- Parameters:
distractor_level (Literal['easy', 'hard'])
- eval_framework.benchmarks.hellaswag_ellamind.hellaswag_ellamind_bpb_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.hellaswag_ellamind.hellaswag_ellamind_easy_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.hellaswag_ellamind.hellaswag_ellamind_hard_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.hendrycks_math_ellamind module¶
German Hendrycks Math (EllaMind), minerva-style — composed.
https://huggingface.co/datasets/ellamind/hendrycks-math-multilingual
The German counterpart of the Minerva-OLMES MATH tasks: Aufgabe: / Lösung: prompt markers, few-shot
demonstrations sampled from the dataset (single-paragraph solutions only, so the \n\n stop separates
blocks), German final-answer lines, minerva_de extraction and the German Minerva metrics.
MATHMinervaDE_OLMES/MATHMinervaDE_OLMES_NONL: free-form, differ only in stop sequences.MATHMinervaDE_BPB_OLMES: bits-per-byte of the single gold solution (same prompt + few-shot).
- eval_framework.benchmarks.hendrycks_math_ellamind.mathminerva_de_bpb_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.hendrycks_math_ellamind.mathminerva_de_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.hendrycks_math_ellamind.mathminerva_de_olmes_nonl(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.hle_ellamind module¶
German HLE (Humanity’s Last Exam, EllaMind) tasks.
https://huggingface.co/datasets/ellamind/hle-multilingual
HLE uses a single distractor set (incorrect_answers). The NATIVE variants restrict evaluation to the
items that are natively multiple-choice in the original benchmark (answer_type == "multipleChoice").
- final class eval_framework.benchmarks.hle_ellamind.HleReader[source]¶
Bases:
ChoiceReaderReads an HLE item: a single
incorrect_answersdistractor set, shuffled in with the correct answer.
- eval_framework.benchmarks.hle_ellamind.hle_ellamind_bpb_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.hle_ellamind.hle_ellamind_cloze_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.hle_ellamind.hle_ellamind_cloze_native_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.hle_ellamind.hle_ellamind_mc_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.hle_ellamind.hle_ellamind_mc_native_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.humaneval module¶
HumanEval code generation: https://huggingface.co/datasets/openai/openai_humaneval
The model completes a Python function stub; a sandboxed metric runs the result against the problem’s tests.
The _OLMES variants generate the body and are scored by execution; the BPB variants instead score the
loglikelihood of the gold solution as a single candidate. The builders and prompt shapes here are exposed for
the sibling humaneval_plus and humaneval_ellamind modules to reuse (retargeted to their datasets /
language), mirroring how their BaseTask ancestors subclassed this task.
- class eval_framework.benchmarks.humaneval.HumanEvalMetricContext(**data)[source]¶
Bases:
BaseMetricContext- Parameters:
test (str)
entry_point (str)
prompt (str)
extra_data (Any)
- entry_point: str¶
- model_config: ClassVar[ConfigDict] = {'extra': 'allow'}¶
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- prompt: str¶
- test: str¶
- class eval_framework.benchmarks.humaneval.SingleGoldReader(question, gold)[source]¶
Bases:
ChoiceReaderChoice reader for the BPB (loglikelihood) variants: the sole candidate is the gold solution, always at index 0.
questionrenders the code prompt andgoldthe reference solution from the raw item, so the same reader serves both the scored sample and its few-shot demonstrations.- Parameters:
question (Callable[[dict[str, Any]], str])
gold (Callable[[dict[str, Any]], str])
- gold: Callable[[dict[str, Any]], str]¶
- question: Callable[[dict[str, Any]], str]¶
- eval_framework.benchmarks.humaneval.bpb(id, *, dataset_path, styler, question, gold, subjects=None, language=Language.ENG, dataset)[source]¶
A BPB (loglikelihood-of-the-gold-solution) benchmark: one candidate, scored by
styler.- Return type:
- Parameters:
id (str)
dataset_path (str)
styler (BPBStyle)
question (Callable[[dict[str, Any]], str])
gold (Callable[[dict[str, Any]], str])
subjects (SubjectsSelector | None)
language (Language | dict[str, Language] | dict[str, tuple[Language, Language]] | None)
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.humaneval.execution(id, *, dataset_path, metrics, build_prompt, fewshot_target, answer=None, subjects=None, language=Language.ENG, dataset)[source]¶
A code-generation-scored-by-execution benchmark.
answerdefaults to the standard HumanEval reconstruction (truncate at a stop sequence, splice into the test harness); an instruct variant can inject its own (e.g. extracting a markdown code block).- Return type:
- Parameters:
id (str)
dataset_path (str)
metrics (list[type[BaseMetric]])
build_prompt (Callable[[dict[str, Any]], str])
fewshot_target (Callable[[dict[str, Any]], str])
answer (AnswerPolicy | None)
subjects (SubjectsSelector | None)
language (Language | dict[str, Language] | dict[str, tuple[Language, Language]] | None)
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.humaneval.humaneval_bpb(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.humaneval.humaneval_bpb_v2(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.humaneval.humaneval_context(item)[source]¶
- Return type:
- Parameters:
item (dict[str, Any])
- eval_framework.benchmarks.humaneval.humaneval_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.humaneval.humaneval_olmes_v2(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.humaneval.olmes_target(item)[source]¶
- Return type:
str- Parameters:
item (dict[str, Any])
- eval_framework.benchmarks.humaneval.v2_bpb(id, *, dataset_path, subjects=None, language=Language.ENG, dataset)[source]¶
The
BPB_V2loglikelihood shape (fenced gold, no leading space), retargetable to another dataset / language — the building blockhumaneval_plusandhumaneval_ellamindreuse.- Return type:
- Parameters:
id (str)
dataset_path (str)
subjects (SubjectsSelector | None)
language (Language | dict[str, Language] | dict[str, tuple[Language, Language]] | None)
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.humaneval.v2_execution(id, *, dataset_path, metrics, subjects=None, language=Language.ENG, dataset)[source]¶
The
_OLMES_V2execution shape (rstrip + fenced), retargetable to another dataset / metric / language — the building blockhumaneval_plusandhumaneval_ellamindreuse.- Return type:
- Parameters:
id (str)
dataset_path (str)
metrics (list[type[BaseMetric]])
subjects (SubjectsSelector | None)
language (Language | dict[str, Language] | dict[str, tuple[Language, Language]] | None)
dataset (DatasetPolicy | None)
eval_framework.benchmarks.humaneval_ellamind module¶
German HumanEval (EllaMind): https://huggingface.co/datasets/ellamind/humaneval-multilingual
The German counterpart of HumanEval — the dataset mirrors the original, so these reuse the builders and prompt
shapes from the humaneval module, retargeted to the German deu subject. HumanEvalDEInstruct differs:
it asks (in German) for the function in a markdown block and extracts the code from the free-form response.
- eval_framework.benchmarks.humaneval_ellamind.humaneval_de_bpb_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.humaneval_ellamind.humaneval_de_bpb_olmes_v2(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.humaneval_ellamind.humaneval_de_instruct(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.humaneval_ellamind.humaneval_de_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.humaneval_ellamind.humaneval_de_olmes_v2(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.humaneval_plus module¶
HumanEvalPlus: HumanEval on the EvalPlus dataset (https://huggingface.co/datasets/evalplus/humanevalplus),
which adds far more test cases per problem. The prompt shape is identical to HumanEval’s _V2 variants, so
these reuse the builders exposed by the humaneval module; only the dataset (and the execution metric,
which needs numpy in the sandbox) differ.
- eval_framework.benchmarks.humaneval_plus.humaneval_plus(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.humaneval_plus.humaneval_plus_bpb(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.ifeval module¶
IFEval: Instruction Following Eval (https://arxiv.org/pdf/2311.07911).
The model follows a natural-language prompt carrying verifiable constraints (word counts, formats, casing, …).
The instruction checks run from a per-sample IFEvalMetricContext, so the task has no gold answer and is
0-shot only.
- eval_framework.benchmarks.ifeval.ifeval(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.ifeval.ifeval_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.math_reasoning module¶
Math reasoning benchmarks.
- eval_framework.benchmarks.math_reasoning.aime(id, *, dataset_policy, ground_truth, sample_split, build_prompt=<function _aime_default_prompt>, subjects=<eval_framework.subjects.NoSubject object>, language=Language.ENG)[source]¶
Build an AIME-style boxed-answer benchmark: 0-shot generative, boxed extraction with the preserved MATH normalisation (its bug is inert on integer answers), scored by
MathReasoningCompletion. Callers vary the prompt (build_prompt, defaulting to the English NeMo-Skills template), dataset, subjects and language — e.g. localized AIME variants in the companion package reuse this.- Return type:
- Parameters:
id (str)
dataset_policy (DatasetPolicy)
ground_truth (Callable[[dict[str, Any]], str])
sample_split (str)
build_prompt (Callable[[dict[str, Any]], str])
subjects (SubjectsSelector)
language (Language)
- eval_framework.benchmarks.math_reasoning.aime2024(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.math_reasoning.aime2025(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.math_reasoning.aime2026(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.math_reasoning.extract_hash_answer(text)[source]¶
The GSM8K gold-answer form: the number after
####with commas removed, or"[invalid]".- Return type:
str- Parameters:
text (str)
- eval_framework.benchmarks.math_reasoning.gsm8k_reasoning(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.math_reasoning.math500_v2(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.math_reasoning.math500_with_bug(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.math_reasoning.mathminerva_bpb(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.math_reasoning.mathminerva_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.math_reasoning.mathminerva_olmes_nonl(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.mbpp module¶
MBPP: https://huggingface.co/datasets/google-research-datasets/mbpp
The model writes a Python function; a sandboxed metric appends the gold assert tests and runs it. The
_OLMES / _EvalPlus variants generate the body (scored by execution) from a fixed 3-shot block; the
BPB variants score the loglikelihood of the gold solution as a single candidate. The builders, reconstruct
functions, and shared pieces here are exposed for the sibling mbpp_ellamind module to reuse (retargeted to
the German dataset), mirroring how its BaseTask ancestors subclassed this task.
- class eval_framework.benchmarks.mbpp.MBPPMetricContext(**data)[source]¶
Bases:
BaseMetricContext- Parameters:
tests_code (str)
extra_data (Any)
- model_config: ClassVar[ConfigDict] = {'extra': 'allow'}¶
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- tests_code: str¶
- class eval_framework.benchmarks.mbpp.SingleGoldReader(question, gold)[source]¶
Bases:
ChoiceReaderChoice reader for the BPB (loglikelihood) variants: the sole candidate is the gold solution, always at index 0.
questionrenders the code prompt andgoldthe reference solution from the raw item.- Parameters:
question (Callable[[dict[str, Any]], str])
gold (Callable[[dict[str, Any]], str])
- gold: Callable[[dict[str, Any]], str]¶
- question: Callable[[dict[str, Any]], str]¶
- eval_framework.benchmarks.mbpp.bpb(id, *, dataset_path, reader, styler, fewshot, subjects=None, language=Language.ENG, dataset)[source]¶
A BPB (loglikelihood-of-the-gold-solution) benchmark: one candidate, scored by
styler. The few-shot policy is passed whole because its source (predefined block vs sampled split) and renderer vary.- Return type:
- Parameters:
id (str)
dataset_path (str)
reader (ChoiceReader)
styler (BPBStyle)
fewshot (FewShotPolicy)
subjects (SubjectsSelector | None)
language (Language | dict[str, Language] | dict[str, tuple[Language, Language]] | None)
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mbpp.evalplus_reconstruct(completion_text, *, context, ground_truth, messages)[source]¶
- Return type:
str- Parameters:
completion_text (str)
context (BaseMetricContext | list[BaseMetricContext] | None)
ground_truth (str | list[str] | None)
messages (list[Message])
- eval_framework.benchmarks.mbpp.execution(id, *, dataset_path, instruction, cue, fewshot_target, answer, fewshot_source, subjects=None, language=Language.ENG, dataset, metrics=None)[source]¶
A code-generation-scored-by-execution benchmark: the gold answers are the
asserttests (appended byanswer’s reconstruction), and demonstrations are drawn fromfewshot_source.- Return type:
- Parameters:
id (str)
dataset_path (str)
instruction (Callable[[dict[str, Any]], str])
cue (str)
fewshot_target (Callable[[dict[str, Any]], str])
answer (AnswerPolicy)
fewshot_source (FewShotSource)
subjects (SubjectsSelector | None)
language (Language | dict[str, Language] | dict[str, tuple[Language, Language]] | None)
dataset (DatasetPolicy | None)
metrics (list[type[BaseMetric]] | None)
- eval_framework.benchmarks.mbpp.instruct_reconstruct(completion_text, *, context, ground_truth, messages)[source]¶
- Return type:
str- Parameters:
completion_text (str)
context (BaseMetricContext | list[BaseMetricContext] | None)
ground_truth (str | list[str] | None)
messages (list[Message])
- eval_framework.benchmarks.mbpp.mbpp_bpb(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mbpp.mbpp_bpb_evalplus(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mbpp.mbpp_context(item)[source]¶
- Return type:
- Parameters:
item (dict[str, Any])
- eval_framework.benchmarks.mbpp.mbpp_evalplus(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mbpp.mbpp_ground_truth(item)[source]¶
- Return type:
str- Parameters:
item (dict[str, Any])
- eval_framework.benchmarks.mbpp.mbpp_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.mbpp_ellamind module¶
German MBPP (EllaMind): https://huggingface.co/datasets/ellamind/mbpp-multilingual
The German counterpart of MBPP — the dataset mirrors the text / code / test_list schema, so these
reuse the builders and reconstruct functions from the mbpp module with a German instruction wrapper. Unlike
the English variants (fixed English 3-shot block), the German ones sample demonstrations from the (German) eval
split. MBPPDEEvalPlusInstruct asks for the function in a markdown block and extracts it from the response.
- eval_framework.benchmarks.mbpp_ellamind.mbpp_de_bpb_evalplus(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mbpp_ellamind.mbpp_de_bpb_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mbpp_ellamind.mbpp_de_evalplus(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mbpp_ellamind.mbpp_de_evalplus_instruct(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mbpp_ellamind.mbpp_de_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.mmlu module¶
MMLU: https://huggingface.co/datasets/cais/mmlu
Multiple-choice knowledge questions across 57 subjects (each an HF config, loaded per subject); every prompt is prefaced by a subject-templated preamble. The composed variants:
MMLU/MMLU_OLMES— score the letter labels; OLMES adds a space before each option label.Full Text MMLU— shows the options as a bulleted list and scores the full answer text.MMLU_IDK— lets the model abstain with"?"and reports confidence-aware metrics.MMLU_COT— the model reasons freely and concludes with the answer, which is regex-extracted from the generation (free-form completion, 0-shot).
- final class eval_framework.benchmarks.mmlu.MmluReader[source]¶
Bases:
ChoiceReaderReads an MMLU item: the question and its four answer choices, with the correct one at
answer.
- eval_framework.benchmarks.mmlu.mmlu(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mmlu.mmlu_cot(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mmlu.mmlu_full_text(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mmlu.mmlu_idk(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mmlu.mmlu_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.mmlu_pro module¶
MMLU-Pro: https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro
Harder, ten-option multiple-choice questions across 14 categories. All questions live in one config and are
split into subjects by the category column. Every prompt is prefaced by a subject-templated preamble.
The composed variants:
Note: MMLU-Pro questions carry a variable number of options (6–10), but every loglikelihood variant scores
a fixed ten letters A–J regardless (see _MmluProMCStyle) — A bug faithfully preserved from the original
implementation, in order to not change the meaning of the score silently.
- final class eval_framework.benchmarks.mmlu_pro.MmluProReader[source]¶
Bases:
ChoiceReaderReads an MMLU-Pro item: the question and its (6–10) options, with the correct one at
answer_index.
- eval_framework.benchmarks.mmlu_pro.mmlu_pro(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mmlu_pro.mmlu_pro_cot(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mmlu_pro.mmlu_pro_cot_v2(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mmlu_pro.mmlu_pro_idk(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.mmlu_pro.mmlu_pro_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.multipl_e module¶
MultiPL-E: translations of HumanEval and MBPP into 6 programming languages (nuprl/MultiPL-E).
Corresponds to the OLMES suites multipl_e_{humaneval,mbpp}:{cpp,java,js,php,rs,sh}::olmo3:n32:v2. Each of
the 12 variants loads one language config (<humaneval|mbpp>-<lang>), is 0-shot only (there are no gold
examples to draw from), and prompts with the target-language function stub verbatim. Grading is entirely
test-based via MultiPLECodeAssertion; the generation is trimmed at language-specific stop tokens.
Recommended run settings for OLMES parity: 0-shot, temperature=0.6, top_p=0.6, repeats=32, max_tokens=1024. Paper: https://ieeexplore.ieee.org/abstract/document/10103177
eval_framework.benchmarks.naturalqs_open module¶
Natural Questions (open): https://huggingface.co/datasets/google-research-datasets/nq_open
Short factoid questions, each with one or more equally-correct gold answers. NaturalQsOpen generates a
free-form answer scored by DROP F1 / exact match (against every gold answer, carried as a per-sample
DropMetricContext); NaturalQsOpenMC_OLMES scores labelled candidate answers by loglikelihood
(allenai/nq-gen2mc).
- eval_framework.benchmarks.naturalqs_open.natural_qs_open(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.naturalqs_open.natural_qs_open_mc_olmes(dataset=None)[source]¶
OLMES lays out the options with a leading space (” A. …”).
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.piqa module¶
PIQA (physical commonsense QA).
https://huggingface.co/datasets/ybisk/piqa
Each item is a goal with two candidate solutions (sol1, sol2); label selects the correct one.
The base task scores the full solution text (cloze); the OLMES variant shows the solutions as lettered
options (multiple choice); the IDK variant lets the model abstain with an “I do not know” answer.
- final class eval_framework.benchmarks.piqa.PiqaReader[source]¶
Bases:
ChoiceReaderReads a PIQA item: the goal and its two candidate solutions (
sol1,sol2) in fixed order.
- eval_framework.benchmarks.piqa.piqa(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.piqa.piqa_idk(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.piqa.piqa_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.piqa_ellamind module¶
German PIQA (EllaMind) tasks.
https://huggingface.co/datasets/ellamind/piqa-multilingual
PIQA supplies separate easy and hard distractors.
- final class eval_framework.benchmarks.piqa_ellamind.PiqaReader(distractor_level)[source]¶
Bases:
ChoiceReaderReads a PIQA item: a single easy/hard distractor per level, shuffled in with the correct solution.
- Parameters:
distractor_level (Literal['easy', 'hard'])
- eval_framework.benchmarks.piqa_ellamind.piqa_ellamind_bpb_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.piqa_ellamind.piqa_ellamind_cloze_easy_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.piqa_ellamind.piqa_ellamind_cloze_hard_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.piqa_ellamind.piqa_ellamind_mc_easy_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.piqa_ellamind.piqa_ellamind_mc_hard_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.sciq module¶
SciQ (English): https://huggingface.co/datasets/allenai/sciq
Science-exam questions, each with three distractors and one correct answer. The four options are shuffled deterministically per item (seeded by the question + answer). The registered variant is OLMES-style: the options are shown as space-prefixed lettered choices (” A. …”) and the model is scored over the letters.
- final class eval_framework.benchmarks.sciq.SciqReader[source]¶
Bases:
ChoiceReaderReads a SciQ item: the question and its options — the correct answer shuffled among its three distractors, deterministically per item (seeded by question + answer).
- eval_framework.benchmarks.sciq.sciq_mc_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.simpleqa_ellamind module¶
German SimpleQA (verified, EllaMind) tasks.
https://huggingface.co/datasets/ellamind/simpleqa-verified-multilingual
SimpleQA supplies separate easy and hard distractors. The answer_aliases field is unused.
- final class eval_framework.benchmarks.simpleqa_ellamind.SimpleqaReader(distractor_level)[source]¶
Bases:
ChoiceReaderReads a SimpleQA item: easy/hard distractor lists per level, shuffled in with the correct answer.
- Parameters:
distractor_level (Literal['easy', 'hard'])
- eval_framework.benchmarks.simpleqa_ellamind.simpleqa_ellamind_bpb_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.simpleqa_ellamind.simpleqa_ellamind_cloze_easy_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.simpleqa_ellamind.simpleqa_ellamind_cloze_hard_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.simpleqa_ellamind.simpleqa_ellamind_mc_easy_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.simpleqa_ellamind.simpleqa_ellamind_mc_hard_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.siqa_ellamind module¶
German Social IQa (EllaMind) tasks.
https://huggingface.co/datasets/ellamind/siqa-multilingual
SIQA supplies separate easy and hard distractors.
- final class eval_framework.benchmarks.siqa_ellamind.SiqaReader(distractor_level)[source]¶
Bases:
ChoiceReaderReads a Social IQa item: the shown question is the context followed by the question, and the easy/hard distractor list for the level is shuffled in with the correct answer.
- Parameters:
distractor_level (Literal['easy', 'hard'])
- eval_framework.benchmarks.siqa_ellamind.siqa_ellamind_bpb_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.siqa_ellamind.siqa_ellamind_cloze_easy_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.siqa_ellamind.siqa_ellamind_cloze_hard_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.siqa_ellamind.siqa_ellamind_mc_easy_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.siqa_ellamind.siqa_ellamind_mc_hard_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.squad module¶
SQuAD reading comprehension: a passage and a question whose answer is a span of the passage (or, in v2, unanswerable). Answers come as several equally-correct annotator spans, scored by (SQuAD-normalised) F1.
SQuAD_OLMES: v1 (rajpurkar/squad), OLMES Title/Background/Question layout, F1 on the raw generation.SQuAD2_MA/SQuAD2_MA_NO_SYSPROMPT: v2 (rajpurkar/squad_v2), the MA-training prompt; the model is told to begin with “Final answer:”, which is stripped back off before scoring. The two differ only in whether the MA system prompt is present.
- eval_framework.benchmarks.squad.squad2_ma(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.squad.squad2_ma_no_sysprompt(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.squad.squad_olmes(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.winogrande module¶
Winogrande: https://huggingface.co/datasets/allenai/winogrande
Pronoun-resolution sentences with a blank _ filled by option1 or option2; answer selects
the correct one. The registered task uses partial evaluation: each item becomes two samples that score the
shared sentence suffix under each option-augmented prefix — p(suffix | prefix + option). WinograndeReader
and PartialEval live here and are reused by the multilingual EllaMind variants.
- final class eval_framework.benchmarks.winogrande.PartialEval[source]¶
Bases:
EvalKindWinogrande partial evaluation: one item becomes two samples, each scoring the shared sentence suffix under one option —
p(suffix | prefix + option).PartialEvalAccuracypairs the two (consecutive ids) and picks the option under which the suffix is likelier.- messages(body, *, fewshot, subject_label)[source]¶
The full message list for one sample — typically
assemble_messageswith the kind’s own preamble and/or system prompt.- Return type:
list[Message]- Parameters:
body (SampleBody)
fewshot (list[FewshotExample])
subject_label (str)
- metrics()[source]¶
The metrics this kind is scored with.
- Return type:
list[type[BaseMetric]]
- samples(item)[source]¶
The scored sample(s) for one eval item — one for most kinds, more when a kind fans out.
- Return type:
list[SampleBody]- Parameters:
item (dict[str, Any])
- final class eval_framework.benchmarks.winogrande.WinograndeReader[source]¶
Bases:
ChoiceReaderReads a Winogrande item: the shown question is the sentence prefix (before the blank
_); each choice is an option completed by the shared suffix (after the blank).
- eval_framework.benchmarks.winogrande.winogrande_cloze(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
eval_framework.benchmarks.winogrande_ellamind module¶
German Winogrande (EllaMind) tasks.
https://huggingface.co/datasets/ellamind/winogrande-multilingual
The sentence contains a blank _ filled by option1 or option2; answer selects the correct
one. Cloze and MC score the two full “option + suffix” completions; partial evaluation instead scores the
shared suffix under each option-augmented prefix.
- eval_framework.benchmarks.winogrande_ellamind.winogrande_ellamind_cloze_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.winogrande_ellamind.winogrande_ellamind_mc_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)
- eval_framework.benchmarks.winogrande_ellamind.winogrande_ellamind_partial_eval_de(dataset=None)[source]¶
- Return type:
- Parameters:
dataset (DatasetPolicy | None)