Workflow evaluation plugin API reference

Reference for evaluation.run(), its result, and the plugin's imports. The plugin reuses the SDK types: see the SDK API reference for Project, Evaluation, Evaluator, RunEvaluator, Goal, and System. For the other methods, see Decorators, Optimization, Rescoring, and Building blocks.

evaluation.run()

evaluation.run()

Runs an evaluation from a workflow. Call it from a @workflow.entrypoint.

from mistralai.workflows.plugins.evaluations import evaluation

result = await evaluation.run(
    dataset=...,
    task=...,
    evaluators=...,
    # optional parameters below
)
ParameterTypeDefaultDescription
datasetSequence[Mapping[str, Any]]requiredList of input records.
taskfunction, class, or strrequired@evaluation.task function, @workflow.define class, or workflow name. See Task modes.
evaluatorslist[Evaluator]requiredPer-record evaluators. Their scorers must be decorated with @evaluation.scorer. Pass [] for a task-only run.
run_evaluatorslist[RunEvaluator] | NoneNoneRun-level evaluators. Their scorers must be decorated with @evaluation.run_scorer.
projectProject | NoneNoneProject to save to (created if it doesn't exist).
evaluationEvaluation | NoneNoneEvaluation to save to (created if it doesn't exist).
namestr | NoneNoneRun name.
descriptionstr | NoneNoneRun description.
tagslist[str] | NoneNoneTags for filtering in Studio.
metadatadict | NoneNoneStatic run metadata, or an @evaluation.run_metadata callback.
record_metadatafunction | NoneNone@evaluation.record_metadata callback, run once per record.
systemSystem | NoneNoneSystem config passed to the task and scorers via context objects.
num_generationsint1Task executions per record, from 1 to 100.
localboolFalseIf True, skip every call to Studio.

Concurrency, timeouts, and retries are set on the decorators, not on run().

It returns an EvalResult:

FieldTypeDescription
run_idstrRun ID, or "local" in local mode.
run_urlstr | NoneLink to the run in Studio. None in local mode.
statisticsdict[str, EvaluatorStatistics]Per-evaluator aggregate statistics.
run_scoresdict[str, Score]Run-level evaluator scores, by evaluator name.
potentially_stale_run_evaluatorslist[str]Set by rescore() only. See Rescoring.

EvalResult is a Pydantic model: return result.model_dump() from your workflow to expose it as the workflow output.

Execution flow

Execution flow

  1. Setup: creates the project, evaluation, and run in Studio, and marks the run as running.
  2. Fan-out: starts one child workflow per record.
  3. Per record: uploads the input record, runs the task and its scorers for each generation, runs the record_metadata callback, and uploads the output record.
  4. Statistics: aggregates scores per evaluator.
  5. Run evaluators: runs the run-level evaluators and uploads their scores.
  6. Run metadata: runs the run_metadata callback and updates the run.
  7. Completion: marks the run as completed, or failed.

While the run is in progress, the workflow sends a heartbeat to Studio. If the worker dies, or the workflow is cancelled or terminated, Studio marks the run as cancelled instead of leaving it in progress.

A failed task or scorer is recorded as an error on its record; the other records continue.

Imports

Imports

Import types from mistralai.workflows.plugins.evaluations.types. This module only re-exports SDK types, so it's safe to import from workflow code running in the Temporal sandbox:

from mistralai.workflows.plugins.evaluations.types import (
    Evaluation,             # Evaluation identifier (name or slug)
    Evaluator,              # Per-record evaluator
    Goal,                   # Pass/fail gate: gte, lte, between
    Project,                # Project identifier (name or slug)
    RecordMetadataContext,  # Context for @evaluation.record_metadata
    RunEvaluator,           # Run-level evaluator
    RunEvaluatorContext,    # Context for @evaluation.run_scorer
    RunMetadataContext,     # Context for @evaluation.run_metadata
    Score,                  # Scorer return type
    ScorerContext,          # Context for @evaluation.scorer
    Statistic,              # Run-level statistic factories
    System,                 # System config (name + params)
    TaskContext,            # Context for @evaluation.task
    Tunable,                # Optimizable parameter slot
    TunableSystem,          # Search space for optimization
)

The evaluation namespace and the optimization types come from the package root:

from mistralai.workflows.plugins.evaluations import (
    evaluation,         # run, rescore, optimize, building blocks, and decorators
    EvalResult,         # Result of run() and rescore()
    GEPA,               # Pareto-based reflective optimizer
    SimpleOptimizer,    # Greedy hill-climb optimizer
    OptimizeResult,     # Result of optimize()
    OptimizeCandidate,  # One candidate in the trajectory
    OptimizeVariant,    # Baseline, winner, or best attempt
    MutatorRequest,     # Input of a custom mutator
    MutatorProposal,    # Output of a custom mutator
)
evaluation.optimize()

evaluation.optimize()

Same parameters as the SDK's client.evaluation.optimize(), except the ones that don't apply in a workflow. See Optimize prompts and parameters for the shared parameters and algorithms, and Optimization for the workflow specifics.

ParameterTypeDefaultDescription
systemTunableSystemrequiredThe search space. Must contain at least one Tunable slot.
datasetlist[dict[str, Any]]requiredInput records.
taskfunction, class, or strrequiredSame as in run().
evaluatorslist[Evaluator]requiredPer-record evaluators. They define the objective and the gates.
algoSimpleOptimizer | GEPArequiredThe search strategy.
run_evaluatorslist[RunEvaluator] | NoneNoneRecorded on each candidate's run. They don't drive selection.
projectProject | NoneNoneProject to save to.
evaluationEvaluation | NoneNoneEvaluation to save to.
steerstr | NoneNonePlain-language goal for the mutator, up to 4,000 characters.
namestr | NoneNoneOptimization name. Generated from steer when omitted.
descriptionstr | NoneNoneOptimization description. Generated from steer when omitted.
tagslist[str] | NoneNoneTags for filtering in Studio.
metadatadict[str, Any] | NoneNoneMetadata attached to each candidate's run.
localboolFalseNot supported: raises ValueError.