Decorators

Every function you pass to the plugin runs as a Temporal activity. A decorator marks the function for the plugin and sets how its activity runs: timeout, retries, and concurrency. evaluation.run() rejects any undecorated task, scorer, run evaluator, or metadata callback.

Use a decorator bare, or call it to override the defaults:

from datetime import timedelta

from mistralai.workflows.plugins.evaluations import evaluation
from mistralai.workflows.plugins.evaluations.types import Score, ScorerContext, TaskContext

@evaluation.task(execution_timeout=timedelta(minutes=30), max_concurrency=10)
async def slow_task(ctx: TaskContext) -> str: ...

@evaluation.scorer
async def accuracy(ctx: ScorerContext) -> Score: ...
@evaluation.task

@evaluation.task

Marks the function that produces an output for each record. It receives a TaskContext.

ParameterDefaultDescription
execution_timeout2 minutesMaximum time for one execution
retry_policy_max_attempts2Maximum attempts on failure
max_concurrency5Maximum task executions running at once

With multiple generations, max_concurrency is a total budget across records and generations: more generations per record means fewer records in flight, not more calls at once.

@evaluation.scorer

@evaluation.scorer

Marks a scorer passed to an Evaluator. It receives a ScorerContext and returns a Score or a number.

ParameterDefaultDescription
execution_timeout1 minuteMaximum time for one execution
retry_policy_max_attempts2Maximum attempts on failure
max_concurrency5Maximum scorers running at once for each record

When a run has several evaluators, the lowest max_concurrency among their scorers applies. A scorer that still fails after its retries records an error score for that record; the run continues.

@evaluation.run_scorer

@evaluation.run_scorer

Marks a scorer passed to a RunEvaluator. It runs once, after all records complete, and receives a RunEvaluatorContext.

ParameterDefaultDescription
execution_timeout1 minuteMaximum time for one execution
retry_policy_max_attempts2Maximum attempts on failure

When results are uploaded to Studio, reading the records back adds up to 10 minutes to this timeout. See Run evaluators.

Metadata callbacks

Metadata callbacks

Metadata callbacks store information derived from task outputs or scores on the records and the run in Studio.

@evaluation.record_metadata runs once per record, after its task and scorers. It receives a RecordMetadataContext, and the returned dict is stored on the record's metadata. Pass it to record_metadata:

from mistralai.workflows.plugins.evaluations.types import RecordMetadataContext

@evaluation.record_metadata
async def output_length(ctx: RecordMetadataContext) -> dict:
    return {"output_len": len(str(ctx.record.generations[0].output))}

result = await evaluation.run(..., record_metadata=output_length)

@evaluation.run_metadata runs once, after the whole run. It receives a RunMetadataContext, and the returned dict is added to the run's metadata. Pass it to metadata instead of a static dict:

from mistralai.workflows.plugins.evaluations.types import RunMetadataContext

@evaluation.run_metadata
async def record_count(ctx: RunMetadataContext) -> dict:
    return {"num_records": len(ctx.records)}

result = await evaluation.run(..., metadata=record_count)

Both callbacks accept execution_timeout (default 1 minute) and retry_policy_max_attempts (default 2). @evaluation.record_metadata also accepts max_concurrency (default 5).

@evaluation.mutator

@evaluation.mutator

Marks a custom mutator that replaces the optimizer's default reflective mutator. See Custom mutators.

ParameterDefaultDescription
execution_timeout2 minutesMaximum time for one mutation
retry_policy_max_attempts2Maximum attempts on failure