eval_framework.benchmarks package

Submodules

eval_framework.benchmarks.arc module

ARC (AI2 Reasoning Challenge): https://huggingface.co/datasets/allenai/ai2_arc

Grade-school science multiple-choice questions, split into the ARC-Easy and ARC-Challenge subjects. The base task scores the full answer text (cloze); the OLMES variant shows the options as space-prefixed lettered choices and scores the letters; the IDK variant lets the model abstain with “I do not know”.

final class eval_framework.benchmarks.arc.ArcReader[source]

Bases: ChoiceReader

Reads an ARC item: the question and its answer options, with the correct one at answerKey — a letter (A–E) or a 1-based number, both normalised to a 0-based index.

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.arc.arc(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.arc.arc_idk(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.arc.arc_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.arc_de module

ARC German (ARC-DE).

final class eval_framework.benchmarks.arc_de.ArcDeReader[source]

Bases: ChoiceReader

Reads an ARC-DE row into choice fields: the German question, its answer texts, and the correct index.

answerKey arrives as either a 1-based number or a letter; answer_key_to_index normalises both, which also frees the reader from caring how many answers a given row offers.

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.arc_de.arc_de(dataset=None)[source]

ARC-DE as cloze/ranked classification.

https://huggingface.co/datasets/LeoLM/ArcChallenge_de

Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.arc_ellamind module

German ARC (EllaMind): https://huggingface.co/datasets/ellamind/arc-multilingual

Grade-school science questions in German. Every slice loads the single German config (deu); the ARC-Easy and ARC-Challenge subjects are the rows of that config tagged by the arc_config column. Cloze scores the full answer text, MC the letter labels, BPB only the ground-truth answer’s bits-per-byte.

final class eval_framework.benchmarks.arc_ellamind.ArcEllamindReader[source]

Bases: ChoiceReader

Reads a German ARC item: the question and its answer options, with the correct one at answer_key (a letter A–E or a 1-based number).

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.arc_ellamind.arc_ellamind_bpb_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.arc_ellamind.arc_ellamind_cloze_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.arc_ellamind.arc_ellamind_mc_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.bigcodebench module

BigCodeBench: https://huggingface.co/datasets/bigcode/bigcodebench

The model completes a self-contained function; the CodeExecutionPassAtOne metric merges the generated snippet with the problem’s unittest harness (via functions carried, serialized, in the sample’s context) and runs it. Only the OLMES 3-shot variant is registered.

eval_framework.benchmarks.bigcodebench.bigcodebench_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.copa module

COPA (Choice of Plausible Alternatives): https://huggingface.co/datasets/aps/super_glue

Causal-reasoning items: a premise, a cause/effect cue, and two alternatives. The registered variant is OLMES-style: the premise is recast as a sentence stem (its final period replaced by the causal connector), the two alternatives are shown as space-prefixed lettered options (” A. …”), and the model is scored over the letter labels. The prompt carries no assistant cue — the options directly continue the stem.

final class eval_framework.benchmarks.copa.CopaReader[source]

Bases: ChoiceReader

Reads a COPA item: the premise recast as a stem (final period → connector) and its two alternatives, each lower-cased to continue the stem.

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.copa.copa_mc_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.cot module

Shared chain-of-thought scaffolding for multiple-choice benchmarks.

A CoT variant asks the model to reason freely and conclude with a stated answer letter, which is pulled back out at scoring time (the injected ExtractFromCompletion). The prompt surface — an optional preamble, the body, and any inert scored candidates — is injected per benchmark; the CoT contract is fixed here: one free-form sample, no assistant cue, the bare answer letter as ground truth, scored by accuracy.

final class eval_framework.benchmarks.cot.Cot(reader, *, build_prompt, preamble=None, candidates=None)[source]

Bases: EvalKind

Multiple-choice chain-of-thought (see the module docstring). build_prompt renders the body, preamble an optional subject-templated top line, and candidates any inert scored letters kept for faithful parity with a loglikelihood baseline (free-form scoring ignores them).

Parameters:
  • reader (ChoiceReader)

  • build_prompt (Callable[[str, list[str]], str])

  • preamble (Callable[[str], str] | None)

  • candidates (Callable[[list[str]], list[str]] | None)

messages(body, *, fewshot, subject_label)[source]

The full message list for one sample — typically assemble_messages with the kind’s own preamble and/or system prompt.

Return type:

list[Message]

Parameters:
metrics()[source]

The metrics this kind is scored with.

Return type:

list[type[BaseMetric]]

samples(item)[source]

The scored sample(s) for one eval item — one for most kinds, more when a kind fans out.

Return type:

list[SampleBody]

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.cot.tulu3_cot_prompt(raw_question, choices)[source]

The parenthesised-option CoT body shared by MMLU-Pro and GPQA.

Reasoning prompt from Figure 44 of the Tülu 3 paper: https://arxiv.org/pdf/2411.15124

Return type:

str

Parameters:
  • raw_question (str)

  • choices (list[str])

eval_framework.benchmarks.cot.tulu_answer()[source]

Extracts the parenthesised letter that tulu3_cot_prompt asks the model to conclude with — "Therefore, the answer is (X)". Kept here beside the prompt because both encode the same (X) format. The accepted letters are the fixed A–J the prompt’s "(A), (B), ..., (E), etc." implies.

Return type:

ExtractFromCompletion

eval_framework.benchmarks.cot.tulu_answer_v2(n_options)[source]

Lenient variant of tulu_answer: no required "Therefore,", optional colon and parentheses, case-insensitive, taking the last match. Only the accepted letter range is benchmark-specific, so it is built from n_options.

Return type:

ExtractFromCompletion

Parameters:

n_options (int)

eval_framework.benchmarks.csqa module

CommonsenseQA (English): https://huggingface.co/datasets/tau/commonsense_qa

Five-way multiple-choice commonsense questions. The registered variant is OLMES-style: the options are shown as space-prefixed lettered choices (” A. …”) and the model is scored over the letter labels.

final class eval_framework.benchmarks.csqa.CommonsenseQaReader[source]

Bases: ChoiceReader

Reads a CommonsenseQA item: the question and its labelled options, with the correct one at answerKey.

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.csqa.csqa_mc_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.csqa_ellamind module

German CommonsenseQA (EllaMind) tasks.

https://huggingface.co/datasets/ellamind/csqa-multilingual

CSQA supplies separate easy and hard distractors.

final class eval_framework.benchmarks.csqa_ellamind.CsqaReader(distractor_level)[source]

Bases: ChoiceReader

Reads a CSQA item: easy/hard distractor lists per level, shuffled in with the correct answer.

Parameters:

distractor_level (Literal['easy', 'hard'])

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.csqa_ellamind.csqa_ellamind_bpb_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.csqa_ellamind.csqa_ellamind_cloze_easy_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.csqa_ellamind.csqa_ellamind_cloze_hard_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.csqa_ellamind.csqa_ellamind_mc_easy_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.csqa_ellamind.csqa_ellamind_mc_hard_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.drop module

DROP (Discrete Reasoning Over Paragraphs): https://huggingface.co/datasets/EleutherAI/drop

A passage and a question over it. DropCompletion_OLMES generates a free-form answer scored by DROP F1 / exact match (EleutherAI/drop); DropMC_OLMES scores labelled candidate answers by loglikelihood (allenai/drop-gen2mc).

eval_framework.benchmarks.drop.drop_completion_olmes(dataset=None)[source]

OLMES: a reading-comprehension preamble, few-shot from the train split, and a 100-token answer budget.

Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.drop.drop_mc_olmes(dataset=None)[source]

OLMES lays out the options with a leading space (” A. …”).

Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.global_mmlu module

Global-MMLU: https://huggingface.co/datasets/CohereLabs/Global-MMLU

MMLU translated into many languages; we evaluate French, German, Spanish, Italian, Portuguese and Arabic.

eval_framework.benchmarks.global_mmlu.global_mmlu(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.global_mmlu.global_mmlu_german(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.goldenswag module

GoldenSwag: https://huggingface.co/datasets/PleIAs/GoldenSwag

A curated subset of HellaSwag, read identically (sentence completion, scored as cloze — see HellaswagReader). GoldenSwag_IDK additionally lets the model abstain with an “I do not know” answer.

eval_framework.benchmarks.goldenswag.goldenswag(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.goldenswag.goldenswag_idk(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gpqa module

GPQA (Graduate-level Google-Proof Q&A): https://huggingface.co/datasets/Idavidrein/gpqa

Gated, expert-written multiple-choice science questions. Each item has one Correct Answer and three Incorrect Answer N distractors; the reader shuffles them together, seeded from the option texts so an item’s option order is stable across runs. The registered variants:

  • GPQA_OLMES: OLMES-style loglikelihood over the full gpqa_extended config.

  • GPQA_DIAMOND_COT / _V2: chain-of-thought completion over the harder gpqa_diamond config; they share one prompt and differ only in how the concluding answer letter is extracted.

One question carries a raw sequence far too long for the prompt budget and is dropped from every config.

final class eval_framework.benchmarks.gpqa.GpqaReader[source]

Bases: ChoiceReader

Reads a GPQA item: three Incorrect Answer N distractors and the Correct Answer, shuffled together. The shuffle is seeded from the (preprocessed) option texts, so the same item always yields the same option order regardless of the order in which items are evaluated.

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.gpqa.gpqa_diamond_cot(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gpqa.gpqa_diamond_cot_v2(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gpqa.gpqa_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gpqa_ellamind module

German GPQA (Graduate-level Professional QA, EllaMind) tasks.

https://huggingface.co/datasets/ellamind/gpqa-multilingual

GPQA uses a single distractor set (incorrect_answers). The diamond variants restrict evaluation to the diamond subset — the 198 hardest questions (is_diamond) from the original GPQA-Diamond benchmark. The COT variant has the model reason in German and conclude with the answer letter, which is leniently regex-extracted from the generation (free-form completion, 0-shot).

final class eval_framework.benchmarks.gpqa_ellamind.GpqaReader[source]

Bases: ChoiceReader

Reads a GPQA item: a single incorrect_answers distractor set, shuffled in with the correct answer.

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_bpb_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_cloze_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_diamond_bpb_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_diamond_cloze_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_diamond_cot_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_diamond_mc_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gpqa_ellamind.gpqa_ellamind_mc_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gpqa_ellamind.tulu3_cot_prompt_de(raw_question, choices)[source]

German translation of tulu3_cot_prompt (Figure 44 of the Tülu 3 paper, https://arxiv.org/pdf/2411.15124): the model reasons briefly and concludes with “Daher ist die Antwort (X)”. The answer format is stated twice — once before the question and once as a reminder after it.

Return type:

str

Parameters:
  • raw_question (str)

  • choices (list[str])

eval_framework.benchmarks.gpqa_ellamind.tulu_answer_de()[source]

Extracts the letter that tulu3_cot_prompt_de asks the model to conclude with, as leniently as tulu_answer_v2: the last match wins, the parentheses are optional, and there is no stop sequence to cut the generation short. It also accepts the English “answer is X”, because a model prompted in German often still concludes in English. The match is anchored on the answer phrase, so a bare “Antwort D” in the reasoning does not count. GPQA always has four options, so only A–D are accepted.

Return type:

ExtractFromCompletion

eval_framework.benchmarks.gsm8k module

GSM8K: https://huggingface.co/datasets/openai/gsm8k

Grade-school math word problems. Both registered variants use eight fixed, hand-written exemplars (not sampled from the dataset) and OLMES answer normalisation:

  • GSM8K_OLMES: free-form completion, scored on the final number of the generation.

  • GSM8KBPB: bits-per-byte of the single normalised gold solution (one forward pass).

eval_framework.benchmarks.gsm8k.clean_short_answer(continuation)[source]

Reduce a solution to its final number, commas removed — the OLMES short-answer form.

Return type:

str

Parameters:

continuation (str)

eval_framework.benchmarks.gsm8k.gsm8k_bpb(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gsm8k.gsm8k_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gsm8k_ellamind module

German GSM8K (EllaMind): https://huggingface.co/datasets/ellamind/gsm8k-platinum-multilingual

The German counterpart of GSM8K (Frage: / Antwort:), with few-shot demonstrations sampled from the test split. Each item carries a worked solution and a final_answer; a demonstration ends with the German final-answer line "Daher ist die Antwort N.". Two registered variants:

  • GSM8K_Ellamind_DE_Platinum: free-form completion, scored on the final integer of the generation.

  • GSM8K_Ellamind_DE_BPB_Platinum: bits-per-byte of the single gold solution.

eval_framework.benchmarks.gsm8k_ellamind.gsm8k_ellamind_de_bpb_platinum(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.gsm8k_ellamind.gsm8k_ellamind_de_platinum(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hellaswag module

HellaSwag: https://huggingface.co/datasets/Rowan/hellaswag

Sentence completion: each item gives an activity label and a context, and the model scores which ending best continues it. Scored as cloze — the prompt is "{activity}: {context}" and the candidates are the full endings. HellaSwag evaluates on the validation split; HellaSwag_OLMES on the (larger) train split, matching OLMES.

final class eval_framework.benchmarks.hellaswag.HellaswagReader[source]

Bases: ChoiceReader

Reads a HellaSwag item: the shown text is "{activity}: {context}" (context is ctx_a plus a capitalised ctx_b), and each choice is a candidate ending. Markup is stripped from all shown text.

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.hellaswag.hellaswag(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hellaswag.hellaswag_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hellaswag_ellamind module

German HellaSwag (EllaMind) tasks.

https://huggingface.co/datasets/ellamind/hellaswag-multilingual

HellaSwag is a sentence-completion task: the prompt is a partial sentence ("{activity}: {context}") and the model scores full-sentence endings. There is no natural MC variant. HellaSwag supplies separate easy and hard distractors.

final class eval_framework.benchmarks.hellaswag_ellamind.HellaswagReader(distractor_level)[source]

Bases: ChoiceReader

Reads a HellaSwag item: the partial sentence "{activity}: {context}", with the easy/hard full-sentence endings for the level shuffled in with the correct ending.

Parameters:

distractor_level (Literal['easy', 'hard'])

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.hellaswag_ellamind.hellaswag_ellamind_bpb_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hellaswag_ellamind.hellaswag_ellamind_easy_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hellaswag_ellamind.hellaswag_ellamind_hard_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hendrycks_math_ellamind module

German Hendrycks Math (EllaMind), minerva-style — composed.

https://huggingface.co/datasets/ellamind/hendrycks-math-multilingual

The German counterpart of the Minerva-OLMES MATH tasks: Aufgabe: / Lösung: prompt markers, few-shot demonstrations sampled from the dataset (single-paragraph solutions only, so the \n\n stop separates blocks), German final-answer lines, minerva_de extraction and the German Minerva metrics.

  • MATHMinervaDE_OLMES / MATHMinervaDE_OLMES_NONL: free-form, differ only in stop sequences.

  • MATHMinervaDE_BPB_OLMES: bits-per-byte of the single gold solution (same prompt + few-shot).

eval_framework.benchmarks.hendrycks_math_ellamind.mathminerva_de_bpb_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hendrycks_math_ellamind.mathminerva_de_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hendrycks_math_ellamind.mathminerva_de_olmes_nonl(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hle_ellamind module

German HLE (Humanity’s Last Exam, EllaMind) tasks.

https://huggingface.co/datasets/ellamind/hle-multilingual

HLE uses a single distractor set (incorrect_answers). The NATIVE variants restrict evaluation to the items that are natively multiple-choice in the original benchmark (answer_type == "multipleChoice").

final class eval_framework.benchmarks.hle_ellamind.HleReader[source]

Bases: ChoiceReader

Reads an HLE item: a single incorrect_answers distractor set, shuffled in with the correct answer.

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.hle_ellamind.hle_ellamind_bpb_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hle_ellamind.hle_ellamind_cloze_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hle_ellamind.hle_ellamind_cloze_native_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hle_ellamind.hle_ellamind_mc_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.hle_ellamind.hle_ellamind_mc_native_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.humaneval module

HumanEval code generation: https://huggingface.co/datasets/openai/openai_humaneval

The model completes a Python function stub; a sandboxed metric runs the result against the problem’s tests. The _OLMES variants generate the body and are scored by execution; the BPB variants instead score the loglikelihood of the gold solution as a single candidate. The builders and prompt shapes here are exposed for the sibling humaneval_plus and humaneval_ellamind modules to reuse (retargeted to their datasets / language), mirroring how their BaseTask ancestors subclassed this task.

class eval_framework.benchmarks.humaneval.HumanEvalMetricContext(**data)[source]

Bases: BaseMetricContext

Parameters:
  • test (str)

  • entry_point (str)

  • prompt (str)

  • extra_data (Any)

entry_point: str
model_config: ClassVar[ConfigDict] = {'extra': 'allow'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

prompt: str
test: str
class eval_framework.benchmarks.humaneval.SingleGoldReader(question, gold)[source]

Bases: ChoiceReader

Choice reader for the BPB (loglikelihood) variants: the sole candidate is the gold solution, always at index 0. question renders the code prompt and gold the reference solution from the raw item, so the same reader serves both the scored sample and its few-shot demonstrations.

Parameters:
  • question (Callable[[dict[str, Any]], str])

  • gold (Callable[[dict[str, Any]], str])

gold: Callable[[dict[str, Any]], str]
question: Callable[[dict[str, Any]], str]
read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.humaneval.bpb(id, *, dataset_path, styler, question, gold, subjects=None, language=Language.ENG, dataset)[source]

A BPB (loglikelihood-of-the-gold-solution) benchmark: one candidate, scored by styler.

Return type:

Benchmark

Parameters:
eval_framework.benchmarks.humaneval.execution(id, *, dataset_path, metrics, build_prompt, fewshot_target, answer=None, subjects=None, language=Language.ENG, dataset)[source]

A code-generation-scored-by-execution benchmark. answer defaults to the standard HumanEval reconstruction (truncate at a stop sequence, splice into the test harness); an instruct variant can inject its own (e.g. extracting a markdown code block).

Return type:

Benchmark

Parameters:
eval_framework.benchmarks.humaneval.humaneval_bpb(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.humaneval.humaneval_bpb_v2(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.humaneval.humaneval_context(item)[source]
Return type:

HumanEvalMetricContext

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.humaneval.humaneval_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.humaneval.humaneval_olmes_v2(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.humaneval.olmes_target(item)[source]
Return type:

str

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.humaneval.v2_bpb(id, *, dataset_path, subjects=None, language=Language.ENG, dataset)[source]

The BPB_V2 loglikelihood shape (fenced gold, no leading space), retargetable to another dataset / language — the building block humaneval_plus and humaneval_ellamind reuse.

Return type:

Benchmark

Parameters:
eval_framework.benchmarks.humaneval.v2_execution(id, *, dataset_path, metrics, subjects=None, language=Language.ENG, dataset)[source]

The _OLMES_V2 execution shape (rstrip + fenced), retargetable to another dataset / metric / language — the building block humaneval_plus and humaneval_ellamind reuse.

Return type:

Benchmark

Parameters:
eval_framework.benchmarks.humaneval.v2_prompt(item)[source]
Return type:

str

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.humaneval_ellamind module

German HumanEval (EllaMind): https://huggingface.co/datasets/ellamind/humaneval-multilingual

The German counterpart of HumanEval — the dataset mirrors the original, so these reuse the builders and prompt shapes from the humaneval module, retargeted to the German deu subject. HumanEvalDEInstruct differs: it asks (in German) for the function in a markdown block and extracts the code from the free-form response.

eval_framework.benchmarks.humaneval_ellamind.humaneval_de_bpb_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.humaneval_ellamind.humaneval_de_bpb_olmes_v2(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.humaneval_ellamind.humaneval_de_instruct(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.humaneval_ellamind.humaneval_de_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.humaneval_ellamind.humaneval_de_olmes_v2(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.humaneval_plus module

HumanEvalPlus: HumanEval on the EvalPlus dataset (https://huggingface.co/datasets/evalplus/humanevalplus), which adds far more test cases per problem. The prompt shape is identical to HumanEval’s _V2 variants, so these reuse the builders exposed by the humaneval module; only the dataset (and the execution metric, which needs numpy in the sandbox) differ.

eval_framework.benchmarks.humaneval_plus.humaneval_plus(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.humaneval_plus.humaneval_plus_bpb(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.ifeval module

IFEval: Instruction Following Eval (https://arxiv.org/pdf/2311.07911).

The model follows a natural-language prompt carrying verifiable constraints (word counts, formats, casing, …). The instruction checks run from a per-sample IFEvalMetricContext, so the task has no gold answer and is 0-shot only.

eval_framework.benchmarks.ifeval.ifeval(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.ifeval.ifeval_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.math_reasoning module

Math reasoning benchmarks.

eval_framework.benchmarks.math_reasoning.aime(id, *, dataset_policy, ground_truth, sample_split, build_prompt=<function _aime_default_prompt>, subjects=<eval_framework.subjects.NoSubject object>, language=Language.ENG)[source]

Build an AIME-style boxed-answer benchmark: 0-shot generative, boxed extraction with the preserved MATH normalisation (its bug is inert on integer answers), scored by MathReasoningCompletion. Callers vary the prompt (build_prompt, defaulting to the English NeMo-Skills template), dataset, subjects and language — e.g. localized AIME variants in the companion package reuse this.

Return type:

Benchmark

Parameters:
  • id (str)

  • dataset_policy (DatasetPolicy)

  • ground_truth (Callable[[dict[str, Any]], str])

  • sample_split (str)

  • build_prompt (Callable[[dict[str, Any]], str])

  • subjects (SubjectsSelector)

  • language (Language)

eval_framework.benchmarks.math_reasoning.aime2024(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.math_reasoning.aime2025(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.math_reasoning.aime2026(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.math_reasoning.extract_hash_answer(text)[source]

The GSM8K gold-answer form: the number after #### with commas removed, or "[invalid]".

Return type:

str

Parameters:

text (str)

eval_framework.benchmarks.math_reasoning.gsm8k_reasoning(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.math_reasoning.math500_v2(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.math_reasoning.math500_with_bug(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.math_reasoning.mathminerva_bpb(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.math_reasoning.mathminerva_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.math_reasoning.mathminerva_olmes_nonl(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mbpp module

MBPP: https://huggingface.co/datasets/google-research-datasets/mbpp

The model writes a Python function; a sandboxed metric appends the gold assert tests and runs it. The _OLMES / _EvalPlus variants generate the body (scored by execution) from a fixed 3-shot block; the BPB variants score the loglikelihood of the gold solution as a single candidate. The builders, reconstruct functions, and shared pieces here are exposed for the sibling mbpp_ellamind module to reuse (retargeted to the German dataset), mirroring how its BaseTask ancestors subclassed this task.

class eval_framework.benchmarks.mbpp.MBPPMetricContext(**data)[source]

Bases: BaseMetricContext

Parameters:
  • tests_code (str)

  • extra_data (Any)

model_config: ClassVar[ConfigDict] = {'extra': 'allow'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

tests_code: str
class eval_framework.benchmarks.mbpp.SingleGoldReader(question, gold)[source]

Bases: ChoiceReader

Choice reader for the BPB (loglikelihood) variants: the sole candidate is the gold solution, always at index 0. question renders the code prompt and gold the reference solution from the raw item.

Parameters:
  • question (Callable[[dict[str, Any]], str])

  • gold (Callable[[dict[str, Any]], str])

gold: Callable[[dict[str, Any]], str]
question: Callable[[dict[str, Any]], str]
read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.mbpp.bpb(id, *, dataset_path, reader, styler, fewshot, subjects=None, language=Language.ENG, dataset)[source]

A BPB (loglikelihood-of-the-gold-solution) benchmark: one candidate, scored by styler. The few-shot policy is passed whole because its source (predefined block vs sampled split) and renderer vary.

Return type:

Benchmark

Parameters:
eval_framework.benchmarks.mbpp.evalplus_reconstruct(completion_text, *, context, ground_truth, messages)[source]
Return type:

str

Parameters:
  • completion_text (str)

  • context (BaseMetricContext | list[BaseMetricContext] | None)

  • ground_truth (str | list[str] | None)

  • messages (list[Message])

eval_framework.benchmarks.mbpp.execution(id, *, dataset_path, instruction, cue, fewshot_target, answer, fewshot_source, subjects=None, language=Language.ENG, dataset, metrics=None)[source]

A code-generation-scored-by-execution benchmark: the gold answers are the assert tests (appended by answer’s reconstruction), and demonstrations are drawn from fewshot_source.

Return type:

Benchmark

Parameters:
eval_framework.benchmarks.mbpp.instruct_reconstruct(completion_text, *, context, ground_truth, messages)[source]
Return type:

str

Parameters:
  • completion_text (str)

  • context (BaseMetricContext | list[BaseMetricContext] | None)

  • ground_truth (str | list[str] | None)

  • messages (list[Message])

eval_framework.benchmarks.mbpp.mbpp_bpb(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mbpp.mbpp_bpb_evalplus(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mbpp.mbpp_context(item)[source]
Return type:

MBPPMetricContext

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.mbpp.mbpp_evalplus(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mbpp.mbpp_ground_truth(item)[source]
Return type:

str

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.mbpp.mbpp_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mbpp.olmes_reconstruct(completion_text, *, context, ground_truth, messages)[source]
Return type:

str

Parameters:
  • completion_text (str)

  • context (BaseMetricContext | list[BaseMetricContext] | None)

  • ground_truth (str | list[str] | None)

  • messages (list[Message])

eval_framework.benchmarks.mbpp_ellamind module

German MBPP (EllaMind): https://huggingface.co/datasets/ellamind/mbpp-multilingual

The German counterpart of MBPP — the dataset mirrors the text / code / test_list schema, so these reuse the builders and reconstruct functions from the mbpp module with a German instruction wrapper. Unlike the English variants (fixed English 3-shot block), the German ones sample demonstrations from the (German) eval split. MBPPDEEvalPlusInstruct asks for the function in a markdown block and extracts it from the response.

eval_framework.benchmarks.mbpp_ellamind.mbpp_de_bpb_evalplus(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mbpp_ellamind.mbpp_de_bpb_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mbpp_ellamind.mbpp_de_evalplus(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mbpp_ellamind.mbpp_de_evalplus_instruct(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mbpp_ellamind.mbpp_de_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mmlu module

MMLU: https://huggingface.co/datasets/cais/mmlu

Multiple-choice knowledge questions across 57 subjects (each an HF config, loaded per subject); every prompt is prefaced by a subject-templated preamble. The composed variants:

  • MMLU / MMLU_OLMES — score the letter labels; OLMES adds a space before each option label.

  • Full Text MMLU — shows the options as a bulleted list and scores the full answer text.

  • MMLU_IDK — lets the model abstain with "?" and reports confidence-aware metrics.

  • MMLU_COT — the model reasons freely and concludes with the answer, which is regex-extracted from the generation (free-form completion, 0-shot).

final class eval_framework.benchmarks.mmlu.MmluReader[source]

Bases: ChoiceReader

Reads an MMLU item: the question and its four answer choices, with the correct one at answer.

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.mmlu.mmlu(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mmlu.mmlu_cot(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mmlu.mmlu_full_text(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mmlu.mmlu_idk(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mmlu.mmlu_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mmlu_pro module

MMLU-Pro: https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro

Harder, ten-option multiple-choice questions across 14 categories. All questions live in one config and are split into subjects by the category column. Every prompt is prefaced by a subject-templated preamble. The composed variants:

Note: MMLU-Pro questions carry a variable number of options (6–10), but every loglikelihood variant scores a fixed ten letters A–J regardless (see _MmluProMCStyle) — A bug faithfully preserved from the original implementation, in order to not change the meaning of the score silently.

final class eval_framework.benchmarks.mmlu_pro.MmluProReader[source]

Bases: ChoiceReader

Reads an MMLU-Pro item: the question and its (6–10) options, with the correct one at answer_index.

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.mmlu_pro.mmlu_pro(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mmlu_pro.mmlu_pro_cot(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mmlu_pro.mmlu_pro_cot_v2(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mmlu_pro.mmlu_pro_idk(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.mmlu_pro.mmlu_pro_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.multipl_e module

MultiPL-E: translations of HumanEval and MBPP into 6 programming languages (nuprl/MultiPL-E).

Corresponds to the OLMES suites multipl_e_{humaneval,mbpp}:{cpp,java,js,php,rs,sh}::olmo3:n32:v2. Each of the 12 variants loads one language config (<humaneval|mbpp>-<lang>), is 0-shot only (there are no gold examples to draw from), and prompts with the target-language function stub verbatim. Grading is entirely test-based via MultiPLECodeAssertion; the generation is trimmed at language-specific stop tokens.

Recommended run settings for OLMES parity: 0-shot, temperature=0.6, top_p=0.6, repeats=32, max_tokens=1024. Paper: https://ieeexplore.ieee.org/abstract/document/10103177

eval_framework.benchmarks.naturalqs_open module

Natural Questions (open): https://huggingface.co/datasets/google-research-datasets/nq_open

Short factoid questions, each with one or more equally-correct gold answers. NaturalQsOpen generates a free-form answer scored by DROP F1 / exact match (against every gold answer, carried as a per-sample DropMetricContext); NaturalQsOpenMC_OLMES scores labelled candidate answers by loglikelihood (allenai/nq-gen2mc).

eval_framework.benchmarks.naturalqs_open.natural_qs_open(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.naturalqs_open.natural_qs_open_mc_olmes(dataset=None)[source]

OLMES lays out the options with a leading space (” A. …”).

Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.piqa module

PIQA (physical commonsense QA).

https://huggingface.co/datasets/ybisk/piqa

Each item is a goal with two candidate solutions (sol1, sol2); label selects the correct one. The base task scores the full solution text (cloze); the OLMES variant shows the solutions as lettered options (multiple choice); the IDK variant lets the model abstain with an “I do not know” answer.

final class eval_framework.benchmarks.piqa.PiqaReader[source]

Bases: ChoiceReader

Reads a PIQA item: the goal and its two candidate solutions (sol1, sol2) in fixed order.

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.piqa.piqa(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.piqa.piqa_idk(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.piqa.piqa_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.piqa_ellamind module

German PIQA (EllaMind) tasks.

https://huggingface.co/datasets/ellamind/piqa-multilingual

PIQA supplies separate easy and hard distractors.

final class eval_framework.benchmarks.piqa_ellamind.PiqaReader(distractor_level)[source]

Bases: ChoiceReader

Reads a PIQA item: a single easy/hard distractor per level, shuffled in with the correct solution.

Parameters:

distractor_level (Literal['easy', 'hard'])

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.piqa_ellamind.piqa_ellamind_bpb_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.piqa_ellamind.piqa_ellamind_cloze_easy_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.piqa_ellamind.piqa_ellamind_cloze_hard_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.piqa_ellamind.piqa_ellamind_mc_easy_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.piqa_ellamind.piqa_ellamind_mc_hard_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.sciq module

SciQ (English): https://huggingface.co/datasets/allenai/sciq

Science-exam questions, each with three distractors and one correct answer. The four options are shuffled deterministically per item (seeded by the question + answer). The registered variant is OLMES-style: the options are shown as space-prefixed lettered choices (” A. …”) and the model is scored over the letters.

final class eval_framework.benchmarks.sciq.SciqReader[source]

Bases: ChoiceReader

Reads a SciQ item: the question and its options — the correct answer shuffled among its three distractors, deterministically per item (seeded by question + answer).

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.sciq.sciq_mc_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.simpleqa_ellamind module

German SimpleQA (verified, EllaMind) tasks.

https://huggingface.co/datasets/ellamind/simpleqa-verified-multilingual

SimpleQA supplies separate easy and hard distractors. The answer_aliases field is unused.

final class eval_framework.benchmarks.simpleqa_ellamind.SimpleqaReader(distractor_level)[source]

Bases: ChoiceReader

Reads a SimpleQA item: easy/hard distractor lists per level, shuffled in with the correct answer.

Parameters:

distractor_level (Literal['easy', 'hard'])

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.simpleqa_ellamind.simpleqa_ellamind_bpb_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.simpleqa_ellamind.simpleqa_ellamind_cloze_easy_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.simpleqa_ellamind.simpleqa_ellamind_cloze_hard_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.simpleqa_ellamind.simpleqa_ellamind_mc_easy_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.simpleqa_ellamind.simpleqa_ellamind_mc_hard_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.siqa_ellamind module

German Social IQa (EllaMind) tasks.

https://huggingface.co/datasets/ellamind/siqa-multilingual

SIQA supplies separate easy and hard distractors.

final class eval_framework.benchmarks.siqa_ellamind.SiqaReader(distractor_level)[source]

Bases: ChoiceReader

Reads a Social IQa item: the shown question is the context followed by the question, and the easy/hard distractor list for the level is shuffled in with the correct answer.

Parameters:

distractor_level (Literal['easy', 'hard'])

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.siqa_ellamind.siqa_ellamind_bpb_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.siqa_ellamind.siqa_ellamind_cloze_easy_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.siqa_ellamind.siqa_ellamind_cloze_hard_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.siqa_ellamind.siqa_ellamind_mc_easy_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.siqa_ellamind.siqa_ellamind_mc_hard_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.social_iqa module

Social IQa (English): https://huggingface.co/datasets/allenai/social_i_qa

Commonsense reasoning about social situations: a context and question with three answer options (answerA/answerB/answerC) and a 1-indexed label marking the correct one. The registered variant is OLMES-style: options are shown as space-prefixed lettered choices (” A. …”) and the model is scored over the letter labels.

allenai/social_i_qa ships only a (no-longer-supported) loading script on its main branch, so it cannot be loaded by a pinned commit under datasets >= 4. We load Hugging Face’s auto-generated parquet branch instead.

final class eval_framework.benchmarks.social_iqa.SocialIqaReader[source]

Bases: ChoiceReader

Reads a Social IQa item: the shown question is the context followed by the question, with the correct option at the 1-indexed label.

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.social_iqa.social_iqa_mc_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.squad module

SQuAD reading comprehension: a passage and a question whose answer is a span of the passage (or, in v2, unanswerable). Answers come as several equally-correct annotator spans, scored by (SQuAD-normalised) F1.

  • SQuAD_OLMES: v1 (rajpurkar/squad), OLMES Title/Background/Question layout, F1 on the raw generation.

  • SQuAD2_MA / SQuAD2_MA_NO_SYSPROMPT: v2 (rajpurkar/squad_v2), the MA-training prompt; the model is told to begin with “Final answer:”, which is stripped back off before scoring. The two differ only in whether the MA system prompt is present.

eval_framework.benchmarks.squad.squad2_ma(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.squad.squad2_ma_no_sysprompt(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.squad.squad_olmes(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.winogrande module

Winogrande: https://huggingface.co/datasets/allenai/winogrande

Pronoun-resolution sentences with a blank _ filled by option1 or option2; answer selects the correct one. The registered task uses partial evaluation: each item becomes two samples that score the shared sentence suffix under each option-augmented prefix — p(suffix | prefix + option). WinograndeReader and PartialEval live here and are reused by the multilingual EllaMind variants.

final class eval_framework.benchmarks.winogrande.PartialEval[source]

Bases: EvalKind

Winogrande partial evaluation: one item becomes two samples, each scoring the shared sentence suffix under one option — p(suffix | prefix + option). PartialEvalAccuracy pairs the two (consecutive ids) and picks the option under which the suffix is likelier.

messages(body, *, fewshot, subject_label)[source]

The full message list for one sample — typically assemble_messages with the kind’s own preamble and/or system prompt.

Return type:

list[Message]

Parameters:
metrics()[source]

The metrics this kind is scored with.

Return type:

list[type[BaseMetric]]

samples(item)[source]

The scored sample(s) for one eval item — one for most kinds, more when a kind fans out.

Return type:

list[SampleBody]

Parameters:

item (dict[str, Any])

final class eval_framework.benchmarks.winogrande.WinograndeReader[source]

Bases: ChoiceReader

Reads a Winogrande item: the shown question is the sentence prefix (before the blank _); each choice is an option completed by the shared suffix (after the blank).

read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.benchmarks.winogrande.winogrande_cloze(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.winogrande_ellamind module

German Winogrande (EllaMind) tasks.

https://huggingface.co/datasets/ellamind/winogrande-multilingual

The sentence contains a blank _ filled by option1 or option2; answer selects the correct one. Cloze and MC score the two full “option + suffix” completions; partial evaluation instead scores the shared suffix under each option-augmented prefix.

eval_framework.benchmarks.winogrande_ellamind.winogrande_ellamind_cloze_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.winogrande_ellamind.winogrande_ellamind_mc_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

eval_framework.benchmarks.winogrande_ellamind.winogrande_ellamind_partial_eval_de(dataset=None)[source]
Return type:

Benchmark

Parameters:

dataset (DatasetPolicy | None)

Module contents