eval_framework.metrics.efficiency package

Submodules

eval_framework.metrics.efficiency.bytes_per_sequence_position module

class eval_framework.metrics.efficiency.bytes_per_sequence_position.BytesCompletion[source]

Bases: BaseMetric[Completion]

NAME: str = 'Bytes'
calculate(response)[source]
Return type:

list[MetricResult]

Parameters:

response (Completion)

class eval_framework.metrics.efficiency.bytes_per_sequence_position.BytesLoglikelihood[source]

Bases: BaseMetric[Loglikelihood]

NAME: str = 'Bytes'
calculate(response)[source]
Return type:

list[MetricResult]

Parameters:

response (Loglikelihood)

class eval_framework.metrics.efficiency.bytes_per_sequence_position.SequencePositionsCompletion[source]

Bases: BaseMetric[Completion]

NAME: str = 'SequencePositions'
calculate(response)[source]
Return type:

list[MetricResult]

Parameters:

response (Completion)

class eval_framework.metrics.efficiency.bytes_per_sequence_position.SequencePositionsLoglikelihood[source]

Bases: BaseMetric[Loglikelihood]

NAME: str = 'SequencePositions'
calculate(response)[source]
Return type:

list[MetricResult]

Parameters:

response (Loglikelihood)

eval_framework.metrics.efficiency.finish_reason module

class eval_framework.metrics.efficiency.finish_reason.FinishReason[source]

Bases: BaseMetric[Completion]

Why the backend ended the generation, as one 0/1 indicator per reason.

Averaged over a benchmark, each key becomes a rate: FinishReason/Repetition is the share of completions vLLM aborted through its repetition_detection sampling parameter (looping output), FinishReason/Length the share that hit max_tokens, and FinishReason/Stop the share that ended on their own (EOS or a stop sequence). Reasons outside these keys (e.g. tool_calls) count as 0 for all of them.

Reads the finish_reason backends already attach to the response. All values are None when the backend did not report one (e.g. HuggingFace) or when the sample errored, so such samples do not dilute the rates.

KEYS: list[str] | None = ['Stop', 'Length', 'Repetition']
NAME: str = 'FinishReason'
calculate(response)[source]
Return type:

list[MetricResult]

Parameters:

response (Completion)

eval_framework.metrics.efficiency.token_counters module

class eval_framework.metrics.efficiency.token_counters.TokenCounts[source]

Bases: BaseMetric[Completion]

Number of tokens the model generated for the completion, and how many of those were spent on reasoning (thinking).

Reads the token counts backends already attach to the response, so no extra tokenisation is performed. Each value independently falls back to None when the backend did not report that particular count or when the sample errored. Reasoning counts are only exposed by some backends (e.g. OpenAI reasoning models via usage.completion_tokens_details.reasoning_tokens, or vLLM started with –reasoning-parser); non-reasoning models and backends that do not surface a per-response count return None for that key.

KEYS: list[str] | None = ['Completion', 'Reasoning']
NAME: str = 'TokenCounts'
calculate(response)[source]
Return type:

list[MetricResult]

Parameters:

response (Completion)

Module contents