Building blocks

When your workflow already manages its own fan-out, for example with custom child workflow orchestration, call the steps of evaluation.run() yourself instead. You keep full control over execution and still get the results in Studio.

The building blocks map to what evaluation.run() does internally:

setup()  →  upload_inputs()  →  [your fan-out]  →  score()  →  upload()

evaluation.run() uploads each record from its own child workflow. With the building blocks, you own the fan-out, so you choose how many records each upload() call sends.

Example

Example

from mistralai.workflows.plugins.evaluations import evaluation
from mistralai.workflows.plugins.evaluations.types import (
    Evaluation, Evaluator, Project, Score, ScorerContext, System,
)

@evaluation.scorer
async def conciseness(ctx: ScorerContext) -> Score:
    ratio = len(str(ctx.output)) / max(len(ctx.input_record["text"]), 1)
    return Score(value=round(1.0 - ratio, 2))

evaluators = [Evaluator(name="conciseness", scorer=conciseness)]

# 1. Create the run in Studio
setup = await evaluation.setup(
    evaluator_names=["conciseness"],
    project=Project(name="Summarization"),
    evaluation=Evaluation(name="Summary quality"),
    system=System(name="small", params={"model": "mistral-small-latest"}),
)

# 2. Upload the input records
input_ids = await evaluation.upload_inputs(run_id=setup["run_id"], dataset=dataset)

# 3. Produce the outputs with your own fan-out
outputs = await my_custom_fan_out(dataset)

# 4. Score the outputs
scores = await evaluation.score(dataset=dataset, outputs=outputs, evaluators=evaluators)

# 5. Upload the outputs and scores
await evaluation.upload(
    run_id=setup["run_id"],
    evaluator_name_to_id=setup["evaluator_name_to_id"],
    input_record_ids=input_ids,
    outputs=outputs,
    scores=scores,
)
evaluation.setup()

evaluation.setup()

Creates the run in Studio. Returns a dict with run_id, run_url, evaluator_name_to_id, and run_evaluator_name_to_id.

ParameterTypeDefaultDescription
evaluator_nameslist[str]requiredNames of the evaluators
run_evaluator_nameslist[str] | NoneNoneNames of the run evaluators
projectProject | NoneNoneProject to create or attach to
evaluationEvaluation | NoneNoneEvaluation to create or attach to
namestr | NoneNoneRun name
descriptionstr | NoneNoneRun description
tagslist[str] | NoneNoneTags for filtering in Studio
metadatadict[str, Any] | NoneNoneMetadata attached to the run
systemSystem | NoneNoneSystem config recorded on the run
num_generationsint1Generations per record, up to 100. Each record you upload must carry this many generations, or Studio treats it as incomplete.
evaluation.upload_inputs()

evaluation.upload_inputs()

Uploads the input records. Returns the list of input record IDs.

ParameterTypeDescription
run_idstrRun ID from setup()
datasetSequence[Mapping[str, Any]]The input records
evaluation.score()

evaluation.score()

Scores the outputs, one per record. Returns a list of score dicts, one per record. Every scorer must be decorated with @evaluation.scorer.

ParameterTypeDescription
datasetSequence[Mapping[str, Any]]The input records
outputslist[Any]The task outputs, one per record
evaluatorslist[Evaluator]The evaluators to run
evaluation.upload()

evaluation.upload()

Uploads the outputs and scores to Studio.

ParameterTypeDescription
run_idstrRun ID from setup()
evaluator_name_to_iddict[str, str]Mapping returned by setup()
input_record_idslist[str]IDs returned by upload_inputs()
outputslist[Any] | NoneTask outputs, one per record. Use for single-generation runs.
scoreslist[dict] | NoneScore dicts from score(). Use with outputs.
generationslist[list[dict]] | NoneFor multi-generation runs: for each record, a list of generations, each a dict with output, scores (evaluator name to score dict), and an optional error.