Rescore persisted runs in a workflow

evaluation.rescore() scores the outputs already persisted in Studio without re-running the task, with your scorers running as activities. Produce an expensive task execution once, then iterate on your scorers as many times as you need.

The semantics are the same as the SDK's client.evaluation.rescore(). See Rescore persisted runs for the full model. The persisted run in Studio is the source of truth, not the workflow history.

Task-only runs

Task-only runs

To persist task outputs without scoring them, pass an empty evaluator list. evaluators stays required, so the empty list is an explicit opt-in:

result = await evaluation.run(
    project=Project(name="Support agent"),
    evaluation=Evaluation(name="Agent rollouts"),
    dataset=dataset,
    task=run_agent,
    evaluators=[],
)

run_evaluators still work with no evaluators: a run evaluator can aggregate outputs, errors, latency, or cost.

Rescore a run

Rescore a run

result = await evaluation.rescore(
    run_id=run_id,
    evaluators=[Evaluator(name="accuracy", scorer=accuracy_v2)],
    run_evaluators=[RunEvaluator(name="pass_rate", scorer=pass_rate)],
)

rescore():

  1. Loads the persisted run and its records: inputs and task outputs.
  2. Runs the scorers against the persisted outputs, as activities.
  3. Persists the new scores, then recomputes statistics and goals.
  4. Runs the run evaluators over the updated records.

Scorers and run evaluators must be decorated with @evaluation.scorer and @evaluation.run_scorer, as in run(). Pass at least one evaluator or run evaluator.

Add and replace

Add and replace

rescore() only changes the evaluators you pass. An evaluator with a new name is added; an evaluator with an existing name is replaced, definition and scores. Score uploads are idempotent, so activity retries never duplicate scores.

If generation scores change while a persisted run evaluator isn't recomputed in the same call, its name is logged and returned in result.potentially_stale_run_evaluators. Pass those run evaluators to run_evaluators to recompute them.

API reference

API reference

ParameterTypeDefaultDescription
run_idstrrequiredID of a persisted run
evaluatorslist[Evaluator] | NoneNoneEvaluators to add or replace
run_evaluatorslist[RunEvaluator] | NoneNoneRun evaluators to add, replace, or recompute
max_concurrencyint10Maximum generations scored at once

Returns an EvalResult, with potentially_stale_run_evaluators set when relevant.