03 · Evaluation modes
Know which evaluator you are paying for
Open Evaluation board while a suite runs. The header
updates from typed evaluator progress immediately; it does not wait for audience-paced event
cards to finish animating.
| Choice | What runs | Provider cost | What to inspect |
| Offline Evals | Five mandatory evaluation gates plus one matched-quality diagnostic; persisted Nadia/Sofia benchmark across Demo01, Demo02, and degraded arms; five predicates; 43 registered mutation diagnostics | None | Authority-labelled gate/diagnostic evidence, benchmark census/floor/reference facts, and three-phase controls |
| Eval 01 · Agent | One fresh Robin agent trial for each selected use case | Paid; confirmation required | Response preview, read-only tool journal, LLM criterion verdicts, and local run receipt |
| Eval 02 · Workflow | One fresh five-executor discovery-workflow trial for each selected use case | Paid; confirmation required | Response preview, executor/route facts, LLM criterion verdicts, and local run receipt |
| Eval 03 · Agent vs workflow | Advanced matched-arm diagnostic; it does not choose a winner or control terminal success | Paid; confirmation required | Per-arm absolute quality plus paired-case observation unit, RepCollapse.All, repetition counts, wins/losses/ties, effective n, p-value, minimum attainable p, mean delta, and power warning |
| Eval 04 · Stochastic agent | Repeated agent trials; five repetitions by default and at least four | Paid; confirmation required | Per-scenario terminal acceptance from the whole-trial 95% Wilson lower bound ≥ 0.50, plus separate per-check census and intervals |
| Eval 05 · Stochastic workflow | Repeated workflow trials; five repetitions by default and at least four | Paid; confirmation required | The same persisted per-arm/per-scenario acceptance decision, plus per-check diagnostic reliability |
| Eval 06 · Safety probes | Fixed Robin-only AgentEval jailbreak and hidden-instruction-extraction campaign | Paid; confirmation required; 104-model-call ceiling by default | Redacted per-attack/per-probe resisted, compromised, inconclusive, error, fidelity, and usage facts; no scenario, repetition, or benchmark run |
| Offline Evals · controls panel | The 43 registered mutation controls inside the offline suite, each using healthy → defect-injected → recovered observations | None | Which exact planted defect was detected and whether the same evaluator recovered; this is not a separate app mode |
Run a paid use-case eval deliberately
1 · PlanChoose the architecture question
In Evals, open Evaluation plan. Offline suite remains selected
by default. Choose Eval 01–06 only when you intend to create new paid measurements.
2 · CaseRead what will be asked
The first-view picker contains Offline, Eval 01 Agent, and Eval 02 Workflow. Turn on
Advanced plans only when you deliberately need the paired diagnostic,
stochastic reliability, or safety investigation. For Eval 01–05, choose one use case or All four canonical use cases.
Before any model call, the panel shows the selected scenario title, description, exact
query, and expected behavior. The completed board expands the independent
ground-truth/tool contract and criteria. Eval 06 instead shows
its fixed campaign and disables scenario/repetition inputs.
3 · Cost boundaryInspect, confirm, run
Set repetitions when the plan supports them, inspect the planned subject/judge counts or
safety probe/model-call ceiling, then turn on Confirm paid execution. The run button remains disabled
until both configuration readiness and confirmation hold.
Workload is explicit: inspect the scenario, arm,
repetition, subject/judge, and safety-call estimates shown in the setup panel before confirming.
Tool calls and retries can make provider-level requests higher than the number of subject trials.
The normative formulas and ceilings live in the
evaluation protocol.
The decision rule is visible: the board identifies the
applicable quality bar, census, and evaluator-owned verdict. It distinguishes a quality threshold
from a chance floor; the shipped 1.000 bar makes all four authored use-case criteria
mandatory. It preserves bounded judge explanations and marks a rule as not applicable when
the plan—such as Eval 06—uses a different safety classification.
Workflow degradation is visible: after two
unusable structured-output attempts, the mapper, reviewer, ranker, or presenter may use its
declared deterministic fallback. The board and receipt show the degradation count and
allow-listed kind. A terminal provider failure or cancellation still makes the trial
NOT MEASURED, even if a fallback response was constructed.
Scenario details have one owner. The setup preview shows the
exact selected case and query. The completed board shows its criterion outcomes and independent
tool/abstention evidence. The canonical case definitions and criterion IDs are maintained in the
evaluation protocol,
rather than duplicated in this operator guide.
How to read the board
AuthorityFive gates + one diagnostic
Catalogue, topology, injection, recall, and honesty are mandatory evaluation gates.
Matched quality is a reported diagnostic. All six stages expose the
observed value, measurement state, pass/fail verdict, evidence, observation producer, and
acceptance evaluator. Expand a row for the cause—not just the completion count.
BenchmarkScores are measured facts
Successes/trials, not-applicable and not-measured census, p-value, minimum attainable p,
observation unit/repetition collapse, power, and reference-arm comparison come from
AgentEval. The UI projects them and does not re-score them.
Controls43/43 means rows checked
The graph increments only after a row reaches ControlCompleted. It does not mean
that the full suite ran 43 times. Expand timeline evidence to see baseline, planted defect,
recovery, and cause for each NC-01…NC-43 row.
SafetyRead the census, not 1.000
Eval 06 shows planned probes and the 104-call ceiling, then resisted,
COMPROMISED, inconclusive, error, truncation, and skipped counts with
per-attack/probe facts and usage. Its quality threshold is explicitly not applicable and
it claims zero BenchmarkRunner directories.
Why this board normally says NOT DERIVABLE: the
offline production/recommendation checks and the live use-case checks do not have a defensible
random-answer population, so no 0.500 floor is invented. A 0.500 chance fact is meaningful
only for a separately registered uniform binary-choice fixture; it is not the passing score,
the live 1.00 quality pass bar, or proof that a result is useful. For Eval 01–05,
that shipped bar requires all four authored criteria; three of four is a measured failure.
“The probes ran” is not the same as “the campaign was
measurable.” Eval 06 records every completed probe attempt. Errored is a
subset of inconclusive, not an additional probe: one inconclusive / one error means
one probe ended in an execution error. A complete, error-free but undecidable campaign is
NOT MEASURED (exit 3). A timeout, transport/execution error, malformed census,
truncation, or missing probe makes the session an INFRASTRUCTURE ERROR (exit 4).
Expand the probe row to see its attack and probe id, allow-listed stage/code/detail, and
whether the underlying raw detail was deliberately suppressed. Raw attack prompts,
responses, canaries, provider messages, and
system instructions remain outside the receipt.
Do not merge the two safety lanes: the offline injection
gate runs AgentEval's real RedTeamRunner against deterministic guarded and
deliberately vulnerable calibration fixtures, including a poison-dependent behavioral
tool call. It is not a Robin measurement. Eval 06 is the separately confirmed paid scan
against fresh real Robin targets. Neither is a retrieval-quality rate.