eval_framework package

Subpackages

Submodules

eval_framework.answer module

How a benchmark obtains the model’s scored answer.

An AnswerPolicy owns the “answer side” of a composed benchmark — the mode the model answers in (loglikelihood over candidates vs. free-form completion), the bounds on any generation (stop sequences, token limit), and how the raw generation is distilled into the answer that metrics score. It is injected into compose next to the eval kind, so the kind stays purely about the prompt and candidates.

class eval_framework.answer.AnswerPolicy[source]

Bases: ABC

The answer side of a kind: the response type, the generation bounds, and answer extraction.

abstractmethod extract_answer(completion_text, *, context, ground_truth, messages)[source]

The answer to score, distilled from the raw generation.

Return type:

str

Parameters:
  • completion_text (str)

  • context (BaseMetricContext | list[BaseMetricContext] | None)

  • ground_truth (str | list[str] | None)

  • messages (list[Message])

abstractmethod max_tokens()[source]

Token limit for completion generation, or None for no limit.

Return type:

int | None

abstractmethod metrics()[source]

The bookkeeping metrics this answer mode always reports (efficiency / token counts), added to the kind’s scoring metrics.

Return type:

list[type[BaseMetric]]

abstractmethod response_type()[source]

Whether the model is scored by loglikelihood over candidates or by free-form completion.

Return type:

ResponseType

abstractmethod stop_sequences()[source]

Stop sequences for completion generation (empty when nothing is generated).

Return type:

list[str]

class eval_framework.answer.CodeReconstructor(*args, **kwargs)[source]

Bases: Protocol

Assembles the runnable program a code-execution metric runs, out of the raw generation plus the sample’s scoring material — the test harness / code prompt carried in context and the gold asserts in ground_truth. This is the code-generation counterpart of ExtractFromCompletion’s extractor, but it needs more than the generation text: the snippet only becomes runnable once spliced together with the problem’s tests.

final class eval_framework.answer.ExtractFromCompletion(extract, stop_sequences=None, *, max_tokens=None)[source]

Bases: AnswerPolicy

Free-form completion: the model generates (bounded by stop_sequences / max_tokens) and the scored answer is produced by extract applied to the generation. Regex extractors are available as first_match / last_match.

Parameters:
  • extract (Callable[[str], str])

  • stop_sequences (list[str] | None)

  • max_tokens (int | None)

extract_answer(completion_text, *, context, ground_truth, messages)[source]

The answer to score, distilled from the raw generation.

Return type:

str

Parameters:
  • completion_text (str)

  • context (BaseMetricContext | list[BaseMetricContext] | None)

  • ground_truth (str | list[str] | None)

  • messages (list[Message])

max_tokens()[source]

Token limit for completion generation, or None for no limit.

Return type:

int | None

metrics()[source]

The bookkeeping metrics this answer mode always reports (efficiency / token counts), added to the kind’s scoring metrics.

Return type:

list[type[BaseMetric]]

response_type()[source]

Whether the model is scored by loglikelihood over candidates or by free-form completion.

Return type:

ResponseType

stop_sequences()[source]

Stop sequences for completion generation (empty when nothing is generated).

Return type:

list[str]

final class eval_framework.answer.PickFromCandidates[source]

Bases: AnswerPolicy

Loglikelihood scoring: the model is scored over fixed candidate completions and the answer is the best-scoring candidate, taken verbatim — nothing is generated, so nothing is bounded or extracted.

extract_answer(completion_text, *, context, ground_truth, messages)[source]

The answer to score, distilled from the raw generation.

Return type:

str

Parameters:
  • completion_text (str)

  • context (BaseMetricContext | list[BaseMetricContext] | None)

  • ground_truth (str | list[str] | None)

  • messages (list[Message])

max_tokens()[source]

Token limit for completion generation, or None for no limit.

Return type:

int | None

metrics()[source]

The bookkeeping metrics this answer mode always reports (efficiency / token counts), added to the kind’s scoring metrics.

Return type:

list[type[BaseMetric]]

response_type()[source]

Whether the model is scored by loglikelihood over candidates or by free-form completion.

Return type:

ResponseType

stop_sequences()[source]

Stop sequences for completion generation (empty when nothing is generated).

Return type:

list[str]

final class eval_framework.answer.ReconstructProgram(reconstruct, *, stop_sequences=None, max_tokens=None)[source]

Bases: AnswerPolicy

Free-form code generation scored by execution: the model generates a solution (bounded by stop_sequences / max_tokens) and reconstruct turns it into the runnable program the metric executes — typically the generated snippet spliced into the prompt and test harness from the sample’s context. The reconstructed program is the scored answer, so the executing metric runs it verbatim.

Parameters:
  • reconstruct (CodeReconstructor)

  • stop_sequences (list[str] | None)

  • max_tokens (int | None)

extract_answer(completion_text, *, context, ground_truth, messages)[source]

The answer to score, distilled from the raw generation.

Return type:

str

Parameters:
  • completion_text (str)

  • context (BaseMetricContext | list[BaseMetricContext] | None)

  • ground_truth (str | list[str] | None)

  • messages (list[Message])

max_tokens()[source]

Token limit for completion generation, or None for no limit.

Return type:

int | None

metrics()[source]

The bookkeeping metrics this answer mode always reports (efficiency / token counts), added to the kind’s scoring metrics.

Return type:

list[type[BaseMetric]]

response_type()[source]

Whether the model is scored by loglikelihood over candidates or by free-form completion.

Return type:

ResponseType

stop_sequences()[source]

Stop sequences for completion generation (empty when nothing is generated).

Return type:

list[str]

eval_framework.answer.first_match(answer_re)[source]

Extractor: group 1 of the first regex match, returned as-is, or "[invalid]".

Return type:

Callable[[str], str]

Parameters:

answer_re (Pattern[str])

eval_framework.answer.last_match(answer_re)[source]

Extractor: the last regex match, upper-cased (for lenient case-insensitive patterns), or "[invalid]".

Return type:

Callable[[str], str]

Parameters:

answer_re (Pattern[str])

eval_framework.base_config module

class eval_framework.base_config.BaseConfig(**data)[source]

Bases: BaseModel

as_dict()[source]
Return type:

dict[str, Any]

classmethod from_yaml(yml_filename)[source]
Return type:

BaseConfig

Parameters:

yml_filename (str | Path)

model_config: ClassVar[ConfigDict] = {'extra': 'forbid', 'frozen': True, 'protected_namespaces': ()}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

save(out_file)[source]
Return type:

None

Parameters:

out_file (Path)

eval_framework.choices module

Choice readers: extracting the fields a choice-based styler needs out of a raw dataset item, so neither the eval nor the styler has to know a benchmark’s item schema.

class eval_framework.choices.ChoiceFields(raw_question, choices, correct_index)[source]

Bases: object

The fields a choice-based styler needs out of a single dataset item.

choices and correct_index are produced together (a benchmark may shuffle the correct answer in among distractors), so a reader yields them in one read rather than via separate calls that would each have to re-derive the same shuffle.

Parameters:
  • raw_question (str)

  • choices (list[str])

  • correct_index (int)

choices: list[str]
correct_index: int
raw_question: str
class eval_framework.choices.ChoiceReader[source]

Bases: ABC

Reads the fields a styler needs out of a raw dataset item.

Isolates dataset-schema knowledge here, so neither the eval nor the styler has to know the shape of a benchmark’s items.

abstractmethod read(item)[source]
Return type:

ChoiceFields

Parameters:

item (dict[str, Any])

eval_framework.composed module

final class eval_framework.composed.ComposedBenchmark(*, id, display_name, subjects, kind, answer, sample_split, fewshot, dataset_policy, language)[source]

Bases: Benchmark

A Benchmark that builds a ComposedEval from an injected eval kind and dataset policy.

Parameters:
classmethod choice(*, id, reader, styler, sample_split, fewshot_split, subjects=None, dataset_policy, language, display_name=None)[source]

Build a choice-based benchmark. The same reader + styler drive both the scored Choice and its matching few-shot demonstrations, so they are given once. A choice is always scored by loglikelihood over its candidates, so the answer is fixed to PickFromCandidates. The few-shot source is picked from whether fewshot_split is the eval split (leak-safe) or a separate one.

Return type:

Self

Parameters:
classmethod compose(*, id, kind, answer, sample_split, fewshot, subjects=None, dataset_policy, language, display_name=None)[source]

Build a ComposedBenchmark from its inputs; subjects defaults to NoSubject (a single unnamed slice) and display_name to id.

Return type:

Self

Parameters:
create(num_fewshot, custom_subjects, custom_hf_revision, seed=None)[source]
Return type:

Eval

Parameters:
  • num_fewshot (int)

  • custom_subjects (list[str] | None)

  • custom_hf_revision (str | None)

  • seed (int | None)

display_name()[source]

The benchmark’s human-readable display name.

Return type:

str

id()[source]

Uniquely identifies the benchmark

Return type:

str

markdown_doc(formatters)[source]

Render the benchmarks’s documentation as markdown.

Return type:

str

Parameters:

formatters (Sequence[BaseFormatter])

metrics()[source]

The benchmark’s scoring metrics (from the kind) plus the answer’s bookkeeping metrics.

Return type:

list[type[BaseMetric]]

response_type()[source]

The benchmark’s response type

Return type:

ResponseType

subjects()[source]

The benchmark’s subjects

Return type:

list[Any]

final class eval_framework.composed.ComposedEval(*, display_name, kind, answer, loader, sample_split, fewshot, subjects, language, rnd)[source]

Bases: Eval

Parameters:
display_name()[source]

Human-readable display name. Is allowed to have special characters and whitespaces.

Return type:

str

generate_completions(llm, samples, stop_sequences=None, max_tokens=None, fail_on_error=True)[source]

Generates completions for the sample. :param sample: sample to generate completions for :type stop_sequences: list[str] | None :param stop_sequences: stop sequences to use in completion generation :type max_tokens: int | None :param max_tokens: maximum tokens to use in completion generation :type fail_on_error: bool :param fail_on_error: if True, re-raise the original exception instead of capturing it

into a per-sample Error completion

Return type:

list[Completion]

Returns:

completion

Parameters:
  • llm (BaseLLM)

  • samples (list[Sample])

  • stop_sequences (list[str] | None)

  • max_tokens (int | None)

  • fail_on_error (bool)

get_max_tokens()[source]

Token limit the eval requests for completion generation, or None for no limit.

Return type:

int | None

get_metadata()[source]

Descriptive metadata about the eval for result reporting.

Return type:

dict[str, str | list[str]]

get_response_type()[source]
Return type:

ResponseType

get_stop_sequences()[source]

Stop sequences the eval requests for completion generation.

Return type:

list[str]

iterate_samples(num_samples=None)[source]

Yield the eval’s samples across all subjects. num_samples caps how many are yielded PER SUBJECT (None = no cap), so a benchmark with S subjects yields up to S * num_samples samples. A sample’s id is its index within its subject, restarting at 0 for each subject.

Return type:

Iterable[Sample]

Parameters:

num_samples (int | None)

eval_framework.contract module

class eval_framework.contract.Benchmark[source]

Bases: ABC

A benchmark is used to provide means of measuring model performance in a domain.

Benchmark act as factories for Eval. They bind all the a prior known information and enrich it with the arguments provided at runtime to create concrete instances of Eval which are used to provide measurements of the models performance.

abstractmethod create(num_fewshot, custom_subjects, custom_hf_revision, seed=None)[source]
Return type:

Eval

Parameters:
  • num_fewshot (int)

  • custom_subjects (list[str] | None)

  • custom_hf_revision (str | None)

  • seed (int | None)

abstractmethod display_name()[source]

Human-readable display name. Is allowed to have special characters and whitespaces.

Return type:

str

abstractmethod id()[source]

Uniquely identifies the benchmark

Return type:

str

abstractmethod markdown_doc(formatters)[source]

Render the benchmarks’s documentation as markdown.

Return type:

str

Parameters:

formatters (Sequence[BaseFormatter])

abstractmethod metrics()[source]

The benchmark’s metrics

Return type:

list[type[BaseMetric]]

abstractmethod response_type()[source]

The benchmark’s response type

Return type:

ResponseType

abstractmethod subjects()[source]

Subjects of the benchmark

Return type:

list[Any]

class eval_framework.contract.Eval[source]

Bases: ABC

The contract a caller relies on to run an evaluation

abstractmethod display_name()[source]

Human-readable display name. Is allowed to have special characters and whitespaces.

Return type:

str

abstractmethod generate_completions(llm, samples, stop_sequences=None, max_tokens=None, fail_on_error=True)[source]

Run llm over samples and return their completions.

Return type:

list[Completion]

Parameters:
  • llm (BaseLLM)

  • samples (list[Sample])

  • stop_sequences (list[str] | None)

  • max_tokens (int | None)

  • fail_on_error (bool)

abstractmethod get_max_tokens()[source]

Token limit the eval requests for completion generation, or None for no limit.

Return type:

int | None

abstractmethod get_metadata()[source]

Descriptive metadata about the eval for result reporting.

Return type:

dict[str, str | list[str]]

abstractmethod get_response_type()[source]
Return type:

ResponseType

abstractmethod get_stop_sequences()[source]

Stop sequences the eval requests for completion generation.

Return type:

list[str]

abstractmethod iterate_samples(num_samples=None)[source]

Yield the eval’s samples across all subjects. num_samples caps how many are yielded PER SUBJECT (None = no cap), so a benchmark with S subjects yields up to S * num_samples samples. A sample’s id is its index within its subject, restarting at 0 for each subject.

Return type:

Iterable[Sample]

Parameters:

num_samples (int | None)

class eval_framework.contract.ResponseType(*values)[source]

Bases: Enum

COMPLETION = 'completion'
LOGLIKELIHOODS = 'loglikelihoods'
class eval_framework.contract.Sample(**data)[source]

Bases: BaseModel

Parameters:
  • id (int)

  • subject (str)

  • messages (list[Message])

  • ground_truth (str | list[str] | None)

  • possible_completions (list[str] | None)

  • context (BaseMetricContext | list[BaseMetricContext] | None)

context: BaseMetricContext | list[BaseMetricContext] | None
ground_truth: str | list[str] | None
id: int
messages: list[Message]
model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

possible_completions: list[str] | None
subject: str

eval_framework.eval_kind module

final class eval_framework.eval_kind.Choice(reader, styler)[source]

Bases: EvalKind

Choice-based eval kind: wraps a reader (item -> ChoiceFields) and a styler (multiple-choice / cloze / BPB), producing exactly one scored sample per item.

Parameters:
messages(body, *, fewshot, subject_label)[source]

The full message list for one sample — typically assemble_messages with the kind’s own preamble and/or system prompt.

Return type:

list[Message]

Parameters:
metadata()[source]

Kind-specific metadata merged into the eval’s get_metadata (e.g. the task style).

Return type:

dict[str, str]

metrics()[source]

The metrics this kind is scored with.

Return type:

list[type[BaseMetric]]

samples(item)[source]

The scored sample(s) for one eval item — one for most kinds, more when a kind fans out.

Return type:

list[SampleBody]

Parameters:

item (dict[str, Any])

class eval_framework.eval_kind.EvalKind[source]

Bases: ABC

The prompt side of a task: the messages to put in front of the model and what its candidates/ground truth are.

E.g. Multiple choice vs Free Form answers. A kind owns the full prompt for an item — the few-shot demonstrations, its own USER turn and (optional) ASSISTANT cue, and any preamble — assembled by messages; ComposedEval only supplies the drawn few-shot examples and wraps the result. The answer side (response type, generation bounds, extraction) is an injected AnswerPolicy.

abstractmethod messages(body, *, fewshot, subject_label)[source]

The full message list for one sample — typically assemble_messages with the kind’s own preamble and/or system prompt.

Return type:

list[Message]

Parameters:
metadata()[source]

Kind-specific metadata merged into the eval’s get_metadata (e.g. the task style).

Return type:

dict[str, str]

abstractmethod metrics()[source]

The metrics this kind is scored with.

Return type:

list[type[BaseMetric]]

abstractmethod samples(item)[source]

The scored sample(s) for one eval item — one for most kinds, more when a kind fans out.

Return type:

list[SampleBody]

Parameters:

item (dict[str, Any])

final class eval_framework.eval_kind.Generative(*, build_prompt, cue, ground_truth, metrics, context=None, initial_prompt=None, system_prompt=None)[source]

Bases: EvalKind

Free-form question -> answer kind: one sample per item, no scored candidates (the answer is extracted from the generation by the injected AnswerPolicy). build_prompt frames the question, cue primes the answer turn ("" for none), ground_truth derives the gold answer, and metrics are the scoring metrics. context derives the per-sample scoring material a metric needs beyond the gold string (see SampleBody.context); initial_prompt is a preamble prepended once above the first (few-shot) turn, and system_prompt derives a leading SYSTEM turn per item (None for no system turn — e.g. an instruction-following task that carries its constraints in the system prompt).

Parameters:
  • build_prompt (Callable[[dict[str, Any]], str])

  • cue (str)

  • ground_truth (Callable[[dict[str, Any]], str | list[str] | None])

  • metrics (list[type[BaseMetric]])

  • context (Callable[[dict[str, Any]], BaseMetricContext | list[BaseMetricContext] | None] | None)

  • initial_prompt (str | None)

  • system_prompt (Callable[[dict[str, Any]], str] | None)

messages(body, *, fewshot, subject_label)[source]

The full message list for one sample — typically assemble_messages with the kind’s own preamble and/or system prompt.

Return type:

list[Message]

Parameters:
metrics()[source]

The metrics this kind is scored with.

Return type:

list[type[BaseMetric]]

samples(item)[source]

The scored sample(s) for one eval item — one for most kinds, more when a kind fans out.

Return type:

list[SampleBody]

Parameters:

item (dict[str, Any])

eval_framework.eval_kind.NoContext(item)[source]

The default ItemContext: the sample carries no metric context.

Return type:

None

Parameters:

item (dict[str, Any])

class eval_framework.eval_kind.SampleBody(prompt, cue, possible_completions, ground_truth, context=None, system_prompt=None)[source]

Bases: object

Parameters:
  • prompt (str)

  • cue (str)

  • possible_completions (list[str])

  • ground_truth (str | list[str] | None)

  • context (BaseMetricContext | list[BaseMetricContext] | None)

  • system_prompt (str | None)

context: BaseMetricContext | list[BaseMetricContext] | None = None
cue: str
ground_truth: str | list[str] | None
possible_completions: list[str]
prompt: str
system_prompt: str | None = None
eval_framework.eval_kind.assemble_messages(fewshot, body, *, initial_prompt=None)[source]

The standard prompt: an optional SYSTEM turn (from body.system_prompt), the few-shot demonstrations as USER / ASSISTANT pairs, then the item’s USER turn and (optional) ASSISTANT cue — with initial_prompt folded once into the first turn (above the demonstrations). The single assembler every kind’s messages delegates to.

Return type:

list[Message]

Parameters:

eval_framework.evaluation_generator module

class eval_framework.evaluation_generator.EvaluationGenerator(config, result_processor)[source]

Bases: object

Parameters:
run_eval()[source]

Runs evaluation using saved completions.

Return type:

list[Result]

eval_framework.exceptions module

exception eval_framework.exceptions.LogicError[source]

Bases: Exception

eval_framework.fewshot module

Few-shot policies: what demonstrations a composed eval shows before each item, and how they render.

final class eval_framework.fewshot.ChoiceRenderer(reader, styler)[source]

Bases: FewShotRenderer

Renders through the same choice reader + styler that score the task, so the shots look exactly like the scored prompt (used by every choice / loglikelihood benchmark).

Parameters:
render(item)[source]

Render one dataset row into a demonstration (prompt + shown answer).

Return type:

FewshotExample

Parameters:

item (dict[str, Any])

final class eval_framework.fewshot.FewShot(source, renderer)[source]

Bases: FewShotPolicy

Draw demonstrations from a source and render each with a renderer — the two orthogonal axes of few-shot (which rows, and how they look).

Parameters:
bind(num_fewshot)[source]

Resolve the shot count — failing fast on an unsupported request, or pinning it for a fixed source — and produce the per-run generator that holds it. Called at eval creation, before any load.

Return type:

FewShotGenerator

Parameters:

num_fewshot (int)

documentation(sample_split)[source]

The demonstration split and example shot count for the rendered task docs.

Return type:

FewShotDoc

Parameters:

sample_split (str)

class eval_framework.fewshot.FewShotDoc(split, example_shots)[source]

Bases: object

What markdown_doc needs to describe a source without running it: the demonstration split (None for a fixed block or no few-shot) and how many demonstrations to show in the rendered example.

Parameters:
  • split (str | None)

  • example_shots (int)

example_shots: int
split: str | None
class eval_framework.fewshot.FewShotGenerator[source]

Bases: ABC

A per-run few-shot worker, bound to a shot count: it remembers its demonstration pool once the data is loaded, then renders the demonstrations to show before each eval item.

abstractmethod for_item(item, rnd)[source]

The rendered demonstrations to show before item — leak-safe against item itself.

Return type:

list[FewshotExample]

Parameters:
  • item (dict[str, Any])

  • rnd (Random)

abstractmethod metadata(sample_split)[source]

Few-shot metadata merged into the eval’s get_metadata (e.g. the source split).

Return type:

dict[str, str]

Parameters:

sample_split (str)

abstractmethod prepare(dataset, *, sample_split, sample_rows)[source]

Materialise the demonstration pool from the already-loaded dataset (called once per subject).

Return type:

None

Parameters:
  • dataset (Mapping[str, Any])

  • sample_split (str)

  • sample_rows (list[dict[str, Any]])

class eval_framework.fewshot.FewShotPolicy[source]

Bases: ABC

The immutable few-shot spec a benchmark holds; binds a run’s shot count into a generator (mirrors DatasetPolicy → DatasetLoader).

abstractmethod bind(num_fewshot)[source]

Resolve the shot count — failing fast on an unsupported request, or pinning it for a fixed source — and produce the per-run generator that holds it. Called at eval creation, before any load.

Return type:

FewShotGenerator

Parameters:

num_fewshot (int)

abstractmethod documentation(sample_split)[source]

The demonstration split and example shot count for the rendered task docs.

Return type:

FewShotDoc

Parameters:

sample_split (str)

class eval_framework.fewshot.FewShotRenderer[source]

Bases: ABC

Turns one drawn dataset row into a solved demonstration — the shown prompt and its correct answer.

The rendering half of few-shot, orthogonal to which rows a source provides.

abstractmethod render(item)[source]

Render one dataset row into a demonstration (prompt + shown answer).

Return type:

FewshotExample

Parameters:

item (dict[str, Any])

class eval_framework.fewshot.FewShotSource[source]

Bases: ABC

Where a few-shot policy draws its demonstration rows, and how many. An immutable spec — the per-run pool and count live on the generator, so a source instance is safe to share across evals.

build_pool (once, when data is loaded) and draw (per eval item) are split so a candidate pool is filtered/materialised once rather than per item.

abstractmethod build_pool(dataset, *, sample_split, sample_rows)[source]

The candidate demonstration rows, from the already-loaded data (called once per subject).

Return type:

list[dict[str, Any]]

Parameters:
  • dataset (Mapping[str, Any])

  • sample_split (str)

  • sample_rows (list[dict[str, Any]])

abstractmethod documentation(sample_split)[source]

The demonstration split and example shot count for the rendered task docs.

Return type:

FewShotDoc

Parameters:

sample_split (str)

abstractmethod draw(pool, *, item, count, rnd)[source]

Select count rows from pool to show before item.

Return type:

list[dict[str, Any]]

Parameters:
  • pool (list[dict[str, Any]])

  • item (dict[str, Any])

  • count (int)

  • rnd (Random)

abstractmethod metadata(sample_split)[source]

Source metadata merged into the eval’s get_metadata (e.g. the demonstration split).

Return type:

dict[str, str]

Parameters:

sample_split (str)

abstractmethod resolve_count(num_fewshot)[source]

The effective shot count — usually num_fewshot unchanged; a fixed source pins it.

Return type:

int

Parameters:

num_fewshot (int)

final class eval_framework.fewshot.FewShotSplit(split, *, keep=None)[source]

Bases: FewShotSource

Draws demonstrations from a dedicated split separate from the eval split — no leak check, since the pool never overlaps the eval items. Optionally restricted to rows passing keep.

Parameters:
  • split (str)

  • keep (Callable[[dict[str, Any]], bool] | None)

build_pool(dataset, *, sample_split, sample_rows)[source]

The candidate demonstration rows, from the already-loaded data (called once per subject).

Return type:

list[dict[str, Any]]

Parameters:
  • dataset (Mapping[str, Any])

  • sample_split (str)

  • sample_rows (list[dict[str, Any]])

documentation(sample_split)[source]

The demonstration split and example shot count for the rendered task docs.

Return type:

FewShotDoc

Parameters:

sample_split (str)

draw(pool, *, item, count, rnd)[source]

Select count rows from pool to show before item.

Return type:

list[dict[str, Any]]

Parameters:
  • pool (list[dict[str, Any]])

  • item (dict[str, Any])

  • count (int)

  • rnd (Random)

metadata(sample_split)[source]

Source metadata merged into the eval’s get_metadata (e.g. the demonstration split).

Return type:

dict[str, str]

Parameters:

sample_split (str)

resolve_count(num_fewshot)[source]

The effective shot count — usually num_fewshot unchanged; a fixed source pins it.

Return type:

int

Parameters:

num_fewshot (int)

class eval_framework.fewshot.FewshotExample(prompt, answer)[source]

Bases: object

Parameters:
  • prompt (str)

  • answer (str)

answer: str
prompt: str
final class eval_framework.fewshot.FunctionRenderer(render)[source]

Bases: FewShotRenderer

Renders via a benchmark-supplied item -> FewshotExample function — for generative tasks that build the demonstration directly rather than through a choice styler.

Parameters:

render (Callable[[dict[str, Any]], FewshotExample])

render(item)[source]

Render one dataset row into a demonstration (prompt + shown answer).

Return type:

FewshotExample

Parameters:

item (dict[str, Any])

final class eval_framework.fewshot.NoFewShot[source]

Bases: FewShotPolicy

A benchmark that only runs 0-shot: it rejects any few-shot request and shows no demonstrations.

bind(num_fewshot)[source]

Resolve the shot count — failing fast on an unsupported request, or pinning it for a fixed source — and produce the per-run generator that holds it. Called at eval creation, before any load.

Return type:

FewShotGenerator

Parameters:

num_fewshot (int)

documentation(sample_split)[source]

The demonstration split and example shot count for the rendered task docs.

Return type:

FewShotDoc

Parameters:

sample_split (str)

final class eval_framework.fewshot.Predefined(items, *, count, label)[source]

Bases: FewShotSource

A fixed, hand-written block of demonstrations (not drawn from the dataset). The count is pinned to count — a benchmark whose prompt uses a canonical fixed few-shot block — warning (rather than showing a different number) if a different count is requested.

Parameters:
  • items (list[dict[str, Any]])

  • count (int)

  • label (str)

build_pool(dataset, *, sample_split, sample_rows)[source]

The candidate demonstration rows, from the already-loaded data (called once per subject).

Return type:

list[dict[str, Any]]

Parameters:
  • dataset (Mapping[str, Any])

  • sample_split (str)

  • sample_rows (list[dict[str, Any]])

documentation(sample_split)[source]

The demonstration split and example shot count for the rendered task docs.

Return type:

FewShotDoc

Parameters:

sample_split (str)

draw(pool, *, item, count, rnd)[source]

Select count rows from pool to show before item.

Return type:

list[dict[str, Any]]

Parameters:
  • pool (list[dict[str, Any]])

  • item (dict[str, Any])

  • count (int)

  • rnd (Random)

metadata(sample_split)[source]

Source metadata merged into the eval’s get_metadata (e.g. the demonstration split).

Return type:

dict[str, str]

Parameters:

sample_split (str)

resolve_count(num_fewshot)[source]

The effective shot count — usually num_fewshot unchanged; a fixed source pins it.

Return type:

int

Parameters:

num_fewshot (int)

final class eval_framework.fewshot.SampleSplit(*, keep=None)[source]

Bases: FewShotSource

Draws demonstrations from the eval (sample) split itself — leak-safe: the current item is excluded so its own answer never appears in its prompt. Optionally restricted to rows passing keep.

Parameters:

keep (Callable[[dict[str, Any]], bool] | None)

build_pool(dataset, *, sample_split, sample_rows)[source]

The candidate demonstration rows, from the already-loaded data (called once per subject).

Return type:

list[dict[str, Any]]

Parameters:
  • dataset (Mapping[str, Any])

  • sample_split (str)

  • sample_rows (list[dict[str, Any]])

documentation(sample_split)[source]

The demonstration split and example shot count for the rendered task docs.

Return type:

FewShotDoc

Parameters:

sample_split (str)

draw(pool, *, item, count, rnd)[source]

Select count rows from pool to show before item.

Return type:

list[dict[str, Any]]

Parameters:
  • pool (list[dict[str, Any]])

  • item (dict[str, Any])

  • count (int)

  • rnd (Random)

metadata(sample_split)[source]

Source metadata merged into the eval’s get_metadata (e.g. the demonstration split).

Return type:

dict[str, str]

Parameters:

sample_split (str)

resolve_count(num_fewshot)[source]

The effective shot count — usually num_fewshot unchanged; a fixed source pins it.

Return type:

int

Parameters:

num_fewshot (int)

eval_framework.logger module

eval_framework.main module

eval_framework.main.main(llm, config, should_preempt_callable=None, trial_id=None, *args, resource_cleanup=False, verbosity=1)[source]

Runs the entire evaluation process: responses generation and evaluation.

Return type:

list[Result]

Parameters:
  • llm (BaseLLM)

  • config (EvalConfig)

  • should_preempt_callable (Callable[[], bool] | None)

  • trial_id (int | None)

  • args (Any)

  • resource_cleanup (bool)

  • verbosity (int)

eval_framework.response_generator module

class eval_framework.response_generator.ResponseGenerator(llm, config, result_processor, *, benchmark_registry=None)[source]

Bases: object

Parameters:
generate(should_preempt_callable)[source]

Generates responses and saves them along with metadata. :type should_preempt_callable: Callable[[], bool] :param should_preempt_callable: function to check if preempt is called

Return type:

tuple[list[Completion | Loglikelihood], bool]

Returns:

list of responses, preempted: whether the process was preempted or not

Parameters:

should_preempt_callable (Callable[[], bool])

eval_framework.response_generator.map_language_to_value(language)[source]
Return type:

str | dict[str, str] | dict[str, tuple[str, str]] | None

Parameters:

language (Language | dict[str, Language] | dict[str, tuple[Language, Language]] | None)

eval_framework.response_generator.repeat_samples(samples, repeats)[source]

Flatten repeats into a single stream of samples.

After expansion original sample indices do not point to the same sample anymore. They Original sample can be recovered by original_index = expanded_index // repeats.

Return type:

Iterable[Sample]

Parameters:
  • samples (Iterable[Sample])

  • repeats (int)

eval_framework.run module

eval_framework.run.parse_args()[source]
Return type:

Namespace

eval_framework.run.run()[source]
Return type:

None

eval_framework.run.run_with_kwargs(kwargs)[source]
Return type:

None

Parameters:

kwargs (dict)

eval_framework.run_direct module

eval_framework.subjects module

Subjects: the slices a benchmark’s evaluation partitions into — which dataset config each loads, and how each is labelled in samples, metadata, and result aggregation.

A SubjectsSelector turns the --task-subjects selector tokens a run requests into the concrete Subjects to evaluate (an empty token list means “all”).

final class eval_framework.subjects.ListOfSubjects(names)[source]

Bases: SubjectsSelector

A task whose subjects are named dataset configs. Each name is both the config to load and the slice’s label; a selector picks names exactly, or "*" picks all.

Parameters:

names (list[str])

select(tokens)[source]
Return type:

Sequence[Subject]

Parameters:

tokens (list[str])

final class eval_framework.subjects.NoSubject[source]

Bases: SubjectsSelector

A task with no subjects: a single unnamed slice. Any selector other than "*" is an error.

select(tokens)[source]
Return type:

Sequence[Subject]

Parameters:

tokens (list[str])

class eval_framework.subjects.Subject(load_key, label)[source]

Bases: object

One evaluation slice.

load_key selects the dataset config to load (None means the dataset’s single config); label identifies the slice in samples, metadata, and result aggregation.

Parameters:
  • load_key (str | None)

  • label (str)

label: str
load_key: str | None
class eval_framework.subjects.SubjectsSelector[source]

Bases: ABC

Selects which slices a run evaluates from its --task-subjects tokens; [] selects all.

abstractmethod select(tokens)[source]
Return type:

Sequence[Subject]

Parameters:

tokens (list[str])

eval_framework.suite module

class eval_framework.suite.MetricSource(**data)[source]

Bases: BaseModel

A single (child, metric) pair used as an input to a SuiteAggregate. See the examples folder for how these are used.

Parameters:
  • child (str)

  • metric (str)

child: str
metric: str
model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

class eval_framework.suite.SuiteAggregate(**data)[source]

Bases: BaseModel

Model to aggregate results from a suite of tasks.

Parameters:
  • name (str)

  • sources (list[MetricSource])

  • method (str | Callable[[list[float]], float])

method: str | Callable[[list[float]], float]
model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

name: str
sources: list[MetricSource]
classmethod validate_method(v)[source]
Return type:

str | Callable

Parameters:

v (str | Callable)

class eval_framework.suite.SuiteResult(**data)[source]

Bases: BaseModel

Parameters:
  • name (str)

  • task_results (dict[str, Self])

  • aggregates (dict[str, float | None])

aggregates: dict[str, float | None]
model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

name: str
task_results: dict[str, Self]
class eval_framework.suite.TaskSuite(**data)[source]

Bases: BaseModel

Parameters:
  • name (str | None)

  • tasks (Annotated[str | list[str | Self], BeforeValidator(func=~eval_framework.suite.parse_strings_to_task_or_suite, json_schema_input_type=PydanticUndefined)])

  • aggregates (list[SuiteAggregate])

  • temperature (float | None)

  • top_p (float | None)

  • top_k (int | None)

  • extra_llm_args (dict[str, Any])

  • num_samples (int | None)

  • num_fewshot (int | None)

  • max_tokens (int | None)

  • repeats (int | None)

  • batch_size (int | None)

  • task_subjects (list[str] | None)

  • hf_revision (str | None)

aggregates: list[SuiteAggregate]
batch_size: int | None
extra_llm_args: dict[str, Any]
get_hyperparam_overrides()[source]

Return hyperparam fields that were explicitly set in the suite definition.

Return type:

dict[str, Any]

hf_revision: str | None
property is_leaf: bool
classmethod load(path)[source]
Return type:

Self

Parameters:

path (Path | str)

classmethod load_from_py(path)[source]
Return type:

Self

Parameters:

path (Path | str)

classmethod load_from_yaml(path)[source]
Return type:

Self

Parameters:

path (Path)

max_tokens: int | None
model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

name: str | None
num_fewshot: int | None
num_samples: int | None
repeats: int | None
property task_name: str

The registered task name. Only valid for leaf tasks.

task_subjects: list[str] | None
tasks: Annotated[str | list[str | Self], BeforeValidator(func=parse_strings_to_task_or_suite, json_schema_input_type=PydanticUndefined)]
temperature: float | None
top_k: int | None
top_p: float | None
validate_suite()[source]
Return type:

Self

eval_framework.suite.compute_aggregates(aggregates, child_results)[source]

Compute suite-level stats from explicitly named (child, metric) sources.

For each SuiteAggregate, the value from each MetricSource is looked up by child name and exact metric key. Sources whose child is missing or whose metric is None or NaN are silently skipped. If no sources yield a valid value the aggregate is None.

Return type:

dict[str, float | None]

Parameters:
eval_framework.suite.parse_strings_to_task_or_suite(v)[source]

Expand bare strings in a list to leaf-suite dicts. Pydantic validates them into TaskSuite.

Return type:

str | list

Parameters:

v (str | list)

eval_framework.suite.resolve_to_evalconfig_kwargs(leaf, resolved_defaults, cli_kwargs)[source]

Build the kwargs dict expected by run_with_kwargs() for a single leaf task.

Merges CLI kwargs as the base, overlays resolved suite defaults, and routes temperature/top_p/extra_llm_args into the llm_args dict.

Return type:

dict

Parameters:
  • leaf (TaskSuite)

  • resolved_defaults (dict[str, Any])

  • cli_kwargs (dict[str, Any])

eval_framework.suite.run_suite(suite, cli_kwargs, parent_defaults=None, root_suite_name=None)[source]

Recursively run all tasks in a suite and compute aggregates bottom-up using post-order traversal.

For a leaf suite: runs the single task via _run_single_task and returns the aggregated results directly. For a composite suite: recurses into each child, collects results, then computes this suite’s aggregates.

Return type:

SuiteResult

Parameters:
  • suite (TaskSuite)

  • cli_kwargs (dict[str, Any])

  • parent_defaults (dict[str, Any] | None)

  • root_suite_name (str | None)

eval_framework.suite.save_suite_results(output_dir, results)[source]
Return type:

None

Parameters:
  • output_dir (Path)

  • results (dict[str, float | None])

Module contents