VITRINE · task-oriented guide

Run it. Read it. Trust the right thing.

A screen-by-screen walkthrough of realistic e-commerce exploration through a single agent or five-executor workflow, the offline evidence suite, optional paid LLM/stochastic/safety evals, durable receipts, and catalogue integrity self-test.

Independent application sample. VITRINE is an independently created value proposal for Digitec Galaxus using synthetic data; no company affiliation, commission, endorsement, internal system, or proprietary customer data is implied.

Why this matters to Digitec Galaxus · VITRINE documentation hub · Implementation details

00 · Before the first run

The UI is available before execution

You do not need to wait for a run to see the control room. The app opens with a pre-run graph preview. When execution is prepared, that preview is replaced by the exact registered function set, reflected MAF graph, or evaluation plan used for the run.

# From the repository root
.\start.ps1 -Mode App
Choose

Mode and persona

Select Agent Demo01, Workflow Demo02, Evals, or Catalogue integrity self-test. The default offline Evals plan includes the 43 diagnostic controls; --controls is also available in the CLI. For a demo persona, the authored scenario description and exact canonical customer query appear before you press Run.

Configure

Execution controls

Personalization switched right means enabled; switched left means disabled. Max rounds caps Demo02 discovery/review cycles. Audience pace delays presentation only—it cannot change execution or evaluator outcomes.

No surprise cost

Offline is the default

The scripted provider and zero-model arms need no remote provider. Any live subject or LLM judge choice is an explicit named plan, readiness-gated, paid, and never silently falls back to offline. The removed live execution profile is not another hidden choice.

Hover help

Let a pointer rest

Buttons and sections expose delayed tooltips explaining their action and prerequisites. Disabled export/replay controls become available only after a completed artifact exists.

First live run? Configure the Azure OpenAI resource endpoint, one explicit API-key/local-Entra/managed-identity mode, subject/judge deployment names, session-scoped variables, and one-case smoke from the Microsoft Foundry live-run setup. A local readiness badge is not a provider connectivity test.
01 · Single agent + tools

Run Agent Demo01

1

Select Agent Demo01

Start with Nadia and Mocked model + real runtime. The mocked provider makes choices deterministic, while the real ChatClientAgent and the actual registered read-only tools still execute.

2

Read what will be asked

The scenario card explains the authored use case and shows the exact query. Changing the persona changes both; there is no invisible generic prompt substituted at run time.

3

Watch the graph and action evidence

An agent node is connected to the currently registered tools with directional arrows. ×N on a tool is the number of observed starts, not an estimate. The event timeline contains the matching request, safe tool parameters, safe tool response, and correlation ID.

4

Inspect the screened outcome

The lower outcome panel is the final customer-facing response after policy and catalogue screening. Expand model/tool events to inspect intermediate responses; use the outcome panel for the final answer. If an output was not captured, the UI says so instead of showing an unexplained blank area.

Comparison arm: Zero-model baseline runs retrieval and code without the agent. It is useful as a labelled baseline, not as evidence that the model helped. Live Azure is a separate explicit paid arm.
02 · Five-executor workflow

Run Workflow Demo02

1

Select Marco

Marco's authored scenario distinguishes his genuine espresso interest from a recent gaming gift. Set Max rounds to 3; the wider field is the hard cap, not the number of rounds the workflow must use.

2

Follow direction

The normal arrows move from InterestMapper → Discovery → CoverageReviewer → Ranker → Presenter. The flattened BACK arrow is the conditional CoverageReviewer → Discovery route. Every connector shows its observed traversal count.

3

Read executor counts

An executor increments once for each NodeStarted. Round-start events do not add a second count. Therefore a reject → rediscover → approve path correctly shows Discovery ×2 and CoverageReviewer ×2.

4

Use the action shelf

Demo02 has executors and typed workflow actions rather than registered agent tools. The shelf below the graph lists executor starts, searches, and any model boundaries observed, with counts and the latest safe detail.

What a loop means: CoverageReviewer found a gap and progress was still possible. It is a normal route, not automatically a failure. The edge count and the reviewer/discovery node counts show exactly how often it occurred.
03 · Evaluation modes

Know which evaluator you are paying for

Open Evaluation board while a suite runs. The header updates from typed evaluator progress immediately; it does not wait for audience-paced event cards to finish animating.

ChoiceWhat runsProvider costWhat to inspect
Offline EvalsFive mandatory evaluation gates plus one matched-quality diagnostic; persisted Nadia/Sofia benchmark across Demo01, Demo02, and degraded arms; five predicates; 43 registered mutation diagnosticsNoneAuthority-labelled gate/diagnostic evidence, benchmark census/floor/reference facts, and three-phase controls
Eval 01 · AgentOne fresh Robin agent trial for each selected use casePaid; confirmation requiredResponse preview, read-only tool journal, LLM criterion verdicts, and local run receipt
Eval 02 · WorkflowOne fresh five-executor discovery-workflow trial for each selected use casePaid; confirmation requiredResponse preview, executor/route facts, LLM criterion verdicts, and local run receipt
Eval 03 · Agent vs workflowAdvanced matched-arm diagnostic; it does not choose a winner or control terminal successPaid; confirmation requiredPer-arm absolute quality plus paired-case observation unit, RepCollapse.All, repetition counts, wins/losses/ties, effective n, p-value, minimum attainable p, mean delta, and power warning
Eval 04 · Stochastic agentRepeated agent trials; five repetitions by default and at least fourPaid; confirmation requiredPer-scenario terminal acceptance from the whole-trial 95% Wilson lower bound ≥ 0.50, plus separate per-check census and intervals
Eval 05 · Stochastic workflowRepeated workflow trials; five repetitions by default and at least fourPaid; confirmation requiredThe same persisted per-arm/per-scenario acceptance decision, plus per-check diagnostic reliability
Eval 06 · Safety probesFixed Robin-only AgentEval jailbreak and hidden-instruction-extraction campaignPaid; confirmation required; 104-model-call ceiling by defaultRedacted per-attack/per-probe resisted, compromised, inconclusive, error, fidelity, and usage facts; no scenario, repetition, or benchmark run
Offline Evals · controls panelThe 43 registered mutation controls inside the offline suite, each using healthy → defect-injected → recovered observationsNoneWhich exact planted defect was detected and whether the same evaluator recovered; this is not a separate app mode

Run a paid use-case eval deliberately

1 · Plan

Choose the architecture question

In Evals, open Evaluation plan. Offline suite remains selected by default. Choose Eval 01–06 only when you intend to create new paid measurements.

2 · Case

Read what will be asked

The first-view picker contains Offline, Eval 01 Agent, and Eval 02 Workflow. Turn on Advanced plans only when you deliberately need the paired diagnostic, stochastic reliability, or safety investigation. For Eval 01–05, choose one use case or All four canonical use cases. Before any model call, the panel shows the selected scenario title, description, exact query, and expected behavior. The completed board expands the independent ground-truth/tool contract and criteria. Eval 06 instead shows its fixed campaign and disables scenario/repetition inputs.

3 · Cost boundary

Inspect, confirm, run

Set repetitions when the plan supports them, inspect the planned subject/judge counts or safety probe/model-call ceiling, then turn on Confirm paid execution. The run button remains disabled until both configuration readiness and confirmation hold.

Workload is explicit: inspect the scenario, arm, repetition, subject/judge, and safety-call estimates shown in the setup panel before confirming. Tool calls and retries can make provider-level requests higher than the number of subject trials. The normative formulas and ceilings live in the evaluation protocol.
The decision rule is visible: the board identifies the applicable quality bar, census, and evaluator-owned verdict. It distinguishes a quality threshold from a chance floor; the shipped 1.000 bar makes all four authored use-case criteria mandatory. It preserves bounded judge explanations and marks a rule as not applicable when the plan—such as Eval 06—uses a different safety classification.
Workflow degradation is visible: after two unusable structured-output attempts, the mapper, reviewer, ranker, or presenter may use its declared deterministic fallback. The board and receipt show the degradation count and allow-listed kind. A terminal provider failure or cancellation still makes the trial NOT MEASURED, even if a fallback response was constructed.
Scenario details have one owner. The setup preview shows the exact selected case and query. The completed board shows its criterion outcomes and independent tool/abstention evidence. The canonical case definitions and criterion IDs are maintained in the evaluation protocol, rather than duplicated in this operator guide.

How to read the board

Authority

Five gates + one diagnostic

Catalogue, topology, injection, recall, and honesty are mandatory evaluation gates. Matched quality is a reported diagnostic. All six stages expose the observed value, measurement state, pass/fail verdict, evidence, observation producer, and acceptance evaluator. Expand a row for the cause—not just the completion count.

Benchmark

Scores are measured facts

Successes/trials, not-applicable and not-measured census, p-value, minimum attainable p, observation unit/repetition collapse, power, and reference-arm comparison come from AgentEval. The UI projects them and does not re-score them.

Controls

43/43 means rows checked

The graph increments only after a row reaches ControlCompleted. It does not mean that the full suite ran 43 times. Expand timeline evidence to see baseline, planted defect, recovery, and cause for each NC-01…NC-43 row.

Safety

Read the census, not 1.000

Eval 06 shows planned probes and the 104-call ceiling, then resisted, COMPROMISED, inconclusive, error, truncation, and skipped counts with per-attack/probe facts and usage. Its quality threshold is explicitly not applicable and it claims zero BenchmarkRunner directories.

Why this board normally says NOT DERIVABLE: the offline production/recommendation checks and the live use-case checks do not have a defensible random-answer population, so no 0.500 floor is invented. A 0.500 chance fact is meaningful only for a separately registered uniform binary-choice fixture; it is not the passing score, the live 1.00 quality pass bar, or proof that a result is useful. For Eval 01–05, that shipped bar requires all four authored criteria; three of four is a measured failure.
“The probes ran” is not the same as “the campaign was measurable.” Eval 06 records every completed probe attempt. Errored is a subset of inconclusive, not an additional probe: one inconclusive / one error means one probe ended in an execution error. A complete, error-free but undecidable campaign is NOT MEASURED (exit 3). A timeout, transport/execution error, malformed census, truncation, or missing probe makes the session an INFRASTRUCTURE ERROR (exit 4). Expand the probe row to see its attack and probe id, allow-listed stage/code/detail, and whether the underlying raw detail was deliberately suppressed. Raw attack prompts, responses, canaries, provider messages, and system instructions remain outside the receipt.
Do not merge the two safety lanes: the offline injection gate runs AgentEval's real RedTeamRunner against deterministic guarded and deliberately vulnerable calibration fixtures, including a poison-dependent behavioral tool call. It is not a Robin measurement. Eval 06 is the separately confirmed paid scan against fresh real Robin targets. Neither is a retrieval-quality rate.
04 · Catalogue integrity self-test

One intentional red row; one green self-test

This is the former “Ablation” mode. It removes one product only from an isolated snapshot, runs the same catalogue contract, and leaves the healthy source at 99.

Underlying gate

Catalogue: FAIL

Expected products 99; observed products 98. The red gate is retained because changing it to pass would erase the proof that the real contract noticed the defect.

Interpretation

EXPECTED DEFECT DETECTED

Timeline and graph present the failed gate as successful detection. Its expanded evidence names the expected/observed counts and isolated removed SKU rather than only a stage counter.

Overall

SELF-TEST SUCCEEDED

The overall self-test is green only when the catalogue defect is detected, all other mandatory gates pass, the matched-quality diagnostic is reported independently, and the requested controls are caught.

Exit code remains 1. The underlying evaluation process still reports a failed admitted gate for automation and artifact fidelity. In this explicitly selected self-test mode, the UI explains that exit 1 is the expected successful detection—not an unexplained suite failure.
05 · Status and evidence glossary

Read state, not color alone

LabelMeaningTypical next action
READYKnown in the preview or plan; not started yetRun or continue the suite
RUNNING / STARTEDA start event exists and no matching terminal event has arrivedWatch the correlated event; subject, check, and trial completions close their starts, while live session completion closes persistence and session starts
DONEThe operation completed; optional audience pacing may still be presenting queued evidenceExpand response/evidence or inspect outcome
DETECTEDA deliberately injected control fault produced the expected failureConfirm recovery also passes
EXPECTED DEFECT DETECTEDThe isolated catalogue 99→98 defect was caughtInspect the red underlying gate as evidence
SELF-TEST SUCCEEDEDThe intentional defect was caught and self-test invariants heldNo repair is required; exit 1 is retained intentionally
FAILEDAn unplanned operation, mandatory-check, or registered-control failure; a diagnostic finding alone does not fail the suiteExpand evidence; expected vs observed and cause are shown
BLOCKEDA runtime guardrail or pre-gate prevented an unsafe/ineligible transitionInspect the guard decision; planted control detection no longer uses this label
NOT MEASUREDNo valid measurement exists; it is not zeroInspect readiness, applicability, or instrument evidence
NOT APPLICABLEThe registered criterion does not apply to this authored requestNo score should be invented
Event expansion: model events show safe request/response; tool events show safe arguments/results; evaluation events show expected outcome, measurement, verdict, score, pass threshold, criterion explanation, workflow degradation, census, and producer/evaluator when available. Judge explanations are bounded evidence, not hidden reasoning. A progress value such as “six of six stages completed” does not mean six mandatory gates passed—the authority label and evidence block explains what happened.
06 · Save and replay

One screened artifact, three views

JSON

Machine-readable receipt

Saves the completed, sanitized, checksummed run artifact for downstream inspection. The checksum detects accidental alteration; it is not a keyed signature or authenticity proof. It becomes enabled only after an artifact exists.

HTML

Human-readable report

Saves a self-contained report generated from the same artifact, including graph, outcome, event evidence, gates, benchmark, and controls.

Replay

Pure projection

Resets the screen and replays stored events. Previous/next rebuild graph state and counters deterministically; replay calls no model, tool, workflow, or evaluator.

Paid results are saved even without pressing JSON or HTML. For Eval 01–05, each selected arm/repetition creates a standard AgentEval run below .agenteval/live; Eval 06 creates no benchmark run. Every paid plan writes a sanitized session receipt to .agenteval/live/live-sessions/<session-id>/outcome.json, and the rolling .agenteval/live/live-sessions/index.json links to it. The Evaluation board and schema-v10 app export show the same plan-specific evidence: Eval 01–03 retain every-trial acceptance and no scenario-decision rows; Eval 04/05 add per-scenario whole-trial Wilson decisions. Eval 01–05 retain the pass threshold, bounded response preview, criterion verdicts/explanations, check facts, tool or workflow/degradation evidence, typed provider-stage recovery/terminal counts, usage when reported, per-check Wilson summaries, and optional paired comparisons; Eval 06 retains its redacted safety configuration, workload ceiling, census/probe facts, and usage.
# Offline launcher aliases (Controls invokes the eval CLI; it is not an app mode)
.\start.ps1 -Mode Demo01
.\start.ps1 -Mode Demo02 -NoRestore
.\start.ps1 -Mode Evals -NoRestore
.\start.ps1 -Mode Controls -NoRestore
.\start.ps1 -Mode Ablation -NoRestore

# Standalone subject selector 1 defaults to the zero-model baseline
dotnet run --project src/AgentEval.VitrineDemo -- 1
# Real Demo01 agent/tools with a deterministic local ChatClient; no remote chat model
dotnet run --project src/AgentEval.VitrineDemo -- 1 --scripted
# Provider-backed subject requires both paid flags
dotnet run --project src/AgentEval.VitrineDemo -- 1 --live --confirm-paid

# Explicit paid paths; secure Azure configuration and --confirm-paid are required
dotnet run --project src/AgentEval.VitrineDemo.Evals -- --eval-plan eval01-agent --scenario nadia-cross-category --confirm-paid
dotnet run --project src/AgentEval.VitrineDemo.Evals -- --eval-plan eval02-workflow --scenario all --confirm-paid
dotnet run --project src/AgentEval.VitrineDemo.Evals -- --eval-plan eval03-compare --scenario marco-gift-trap --repetitions 3 --confirm-paid
dotnet run --project src/AgentEval.VitrineDemo.Evals -- --eval-plan eval04-stochastic-agent --scenario all --repetitions 5 --confirm-paid
dotnet run --project src/AgentEval.VitrineDemo.Evals -- --eval-plan eval05-stochastic-workflow --scenario luca-safe-abstention --repetitions 5 --confirm-paid
dotnet run --project src/AgentEval.VitrineDemo.Evals -- --eval-plan eval06-safety-probes --confirm-paid
# Backward-compatible parser alias for the same confirmed Eval 03 plan (not a profile)
dotnet run --project src/AgentEval.VitrineDemo.Evals -- --live-subjects-and-judge --confirm-paid
Eval 06 is intentionally fixed: the CLI rejects both --scenario and --repetitions, even all or 1. The default four-probe campaign may consume up to 104 model calls; inspect that ceiling before confirming paid execution.
Standalone subject cost boundary: credentials do not opt selectors 1–6 into a live model. Demo01's --scripted arm explicitly selects the local deterministic ChatClient while preserving the real agent/tool runtime; --offline, --scripted, and --live are mutually exclusive. A provider-backed subject requires --live --confirm-paid. Because --real-vectors embeds queries live and --rebuild-embeddings regenerates the live index, both vector options also require --confirm-paid.
Verification safety: the implementation build, normal tests, self-test, and committed offline report refresh use no live subject or LLM judge. A paid plan is a separate operator action and produces a new local session receipt.

Configure and smoke the explicit live lane → · Continue to the normative evaluation protocol →