← VITRINE documentation

Sanitized live-evaluation receipt

liveEval02Workflow

Session 20260910T170542623Z-9b556309c0f24741903290ea7b54b4cf · source schema 1.4 · source SHA-256 f4c494141f0e30eb69d83f969658f4fad6735318158f8e3e5d8dc4a66e0bd209

Terminal statuspassed
Exit code0
Started (UTC)2026-09-10T17:05:42.6358103+00:00
Completed (UTC)2026-09-10T17:07:31.8010688+00:00
Planned subject executions1
Planned judge evaluations1
Reported subject tokens14127
Subject estimated cost$0.174255
Reported judge tokens1998
Judge estimated cost$0.022035
Publication boundary. This receipt contains allow-listed aggregate evidence only. Exact queries, evaluator prompts, expected answers, ground truth, model responses, tool arguments, judge explanations, provider errors, endpoints, credentials, canaries, absolute paths and run directories are deliberately absent. Counts describe this session, not a universal model claim.

Configuration

Definition vitrine-live-use-cases 1.1.0+cases.sofia-capability-gap.rubric.e8da57b913f4df9e.threshold.3FF0000000000000.judge.configured-model.tokens.1200 · judge model configured-judge-model · rubric SHA-256 e8da57b913f4df9efd9a91e6279eb51bfce6f06e6ac61863017276c9f4ca35e0

Scenarios

Replenishment cadence and capability gap

Sofia repeatedly buys beans and cartridges, owns a canister, and has no grinder on file. The run should separate replenishment from discovery and identify the missing capability without recommending another owned durable.

sofia-capability-gap · persona USR-SK-03 · criteria finds-capability-gap, separates-lanes, avoids-owned-durable, grounds-advice

Per-check diagnostic census

These pooled rows explain each arm. When present, the separate per-scenario table owns terminal policy for stochastic plans.

ArmCheckCensusSuccessesEstimate95% Wilson interval
discovery-workflow-livevitrine.live.use-case-quality
Canonical use-case criteria
M 1 · N/A 0 · NM 01/11.000[0.207, 1.000]
discovery-workflow-livevitrine.live.response-observed
A non-empty customer response was instrumented
M 1 · N/A 0 · NM 01/11.000[0.207, 1.000]
discovery-workflow-livevitrine.live.agent-tool-journal
Agent tool journal is complete and read-only
M 0 · N/A 1 · NM 00/0not measured[not measured, not measured]
discovery-workflow-livevitrine.live.workflow-trace
Workflow executors and routes completed without failure
M 1 · N/A 0 · NM 01/11.000[0.207, 1.000]

Criterion census

ScenarioArmCriterionObservedMeasuredMetUnmetN/ANot measured
sofia-capability-gapdiscovery-workflow-livefinds-capability-gap111000
sofia-capability-gapdiscovery-workflow-liveseparates-lanes111000
sofia-capability-gapdiscovery-workflow-liveavoids-owned-durable111000
sofia-capability-gapdiscovery-workflow-livegrounds-advice111000

Trial usage and verdicts

sofia-capability-gap · discovery-workflow-live · repetition 1

Status completed · measurement measured · passed True

Subject usage: measured; calls 3; tokens 14127; estimated cost $0.174255
Judge usage: measured; calls not measured; tokens 1998; estimated cost $0.022035

Provider-stage reliability: failed attempts 0 · recovered failed attempts 0 · terminal stages 0

JSON companion: vitrine-live-eval02-sofia-2026-09-10.json. Generated deterministically from the private receipt by eng/export-live-evidence.ps1.