Build executable alternatives
A tool-using recommendation agent and a five-stage workflow operate over the same synthetic commerce domain, so control structures can be compared rather than debated in the abstract.
What I can help Digitec Galaxus achieve: move an AI-assisted shopping experience from an impressive demonstration to a capability whose relevance, safety, variability, cost, and failure modes can be inspected before and after every change.
A product-discovery model can sound convincing while recommending an owned product, following a gift-derived false preference, inventing commercial facts, or behaving differently on the next execution. My contribution is to make those risks explicit, testable, observable, and reviewable by product, engineering, and safety stakeholders.
A tool-using recommendation agent and a five-stage workflow operate over the same synthetic commerce domain, so control structures can be compared rather than debated in the abstract.
Ground truth, criteria, tool expectations, missing-measurement rules, repetitions, and adversarial probes live outside the system being evaluated. The subject never grades itself.
Each execution lane retains the evidence appropriate to it: app runs emit typed events and support checksummed export and replay; the offline benchmark and Eval01–Eval05 create standard AgentEval run directories; every paid plan writes a sanitized local session receipt.
VITRINE does not claim to know Digitec Galaxus production internals. It demonstrates an evaluation method that can be adapted to governed catalogue, customer, search, safety, and business data.
| Business question | VITRINE evidence | Decision enabled |
|---|---|---|
| Are recommendations grounded and useful? | Scenario-specific ground truth, LLM criteria, catalogue checks, and deterministic tool expectations. | Catch relevance, ownership, and fabricated-fact regressions before release. |
| Should this use case be an agent or a workflow? | When explicitly invoked, Eval03 runs matched cases through both architectures and retains native paired comparison facts. | Choose from observed quality, control, reliability, and cost trade-offs—not fashion. |
| Will the behavior remain stable? | When explicitly invoked, Eval04 and Eval05 repeat provider-backed executions. Every selected arm/scenario persists a whole-trial 95% Wilson interval and must clear a 0.50 lower-bound floor; per-check intervals remain diagnostic evidence. | Expose variability and decide whether the evidence is mature enough for a rollout. |
| Can the system resist misuse? | The offline suite calibrates injection detection on guarded and deliberately vulnerable fixtures. When explicitly invoked, Eval06 targets fresh Robin instances with AgentEval jailbreak and hidden-instruction-extraction attacks. | Prioritize containment work and rerun the same bounded safety contract after changes. |
| What did a run cost? | Paid work is explicitly confirmed, planned calls are shown before execution, and provider usage is retained when reported. | Compare quality per unit of spend and distinguish measured zero from missing telemetry. |
| Can we defend the decision later? | Sanitized receipts, run IDs, persisted benchmark output, checksummed artifacts, HTML/JSON export, and pure replay. The checksum detects accidental changes; it is not a signature or proof against malicious replacement. | Reproduce evidence for engineering review, incident learning, and model-change governance. |
The four cases are intentionally different. They prevent one generic “helpful recommendation” rubric from hiding distinct product and trust failures.
Connect hiking, pre-sunrise lighting, off-grid power, and photography evidence while avoiding products she already owns.
Recognize a missing coffee grinder while separating consumable replenishment from new durable-product discovery.
Suppress gifted gaming purchases as a preference signal and personalize only from Marco’s own espresso evidence.
Treat one low-information cable purchase as insufficient evidence, ask useful questions, and avoid speculative recommendations.
The goal is not to call an LLM for every assertion. Deterministic contracts protect invariants quickly; paid evaluation is reserved for semantic quality, stochastic reliability, architectural comparison, and real adversarial behavior.
Run the agent and workflow, inspect tool calls, graph routes, fallbacks, screening, and final customer response.
Apply catalogue, workflow, memory, honesty, injection, and recommendation predicates without provider spend.
Use scenario-routed LLM judging only after readiness and explicit cost confirmation.
Repeat runs for Wilson intervals; compare architectures; probe jailbreak and instruction leakage.
Keep bounded receipts and native run facts, fix a failure, then rerun the same contract.
A production programme would be collaborative: product and category experts define useful behavior, safety and privacy owners set boundaries, and engineering makes the evidence repeatable.
Choose a narrow product-discovery journey; map failure modes and decision owners; define a governed scenario registry; instrument tool, retrieval, screening, latency, cost, and fallback facts; establish deterministic baselines.
Connect read-only staging data; have domain experts review rubrics and failure examples; run blinded agent/workflow comparisons; calibrate judge agreement; add offline release gates and explicit missing-measurement policy.
Run scheduled model-backed canaries and bounded RedTeam probes; analyze reliability and spend; define model or prompt change gates; connect technical quality to privacy-approved business measures such as successful discovery, add-to-cart, support contacts, returns, and customer feedback.
The current receipts are engineering evidence from a synthetic sample. They demonstrate that the contracts execute and can detect planted defects; they are not Digitec Galaxus business metrics.
At the 2026-09-10 verification snapshot: Release build with zero warnings/errors; 450/450 credential-free tests; all five mandatory evaluation gates passed and the matched-quality diagnostic was reported; 43/43 registered mutations detected; catalogue defect detected and recovery verified.
The checked-in offline report records cases, arms, checks, census, paired comparisons, gates, controls, and the exact canonical run identity.
The live-plan orchestration and all selectors are exercised with injected fakes in ordinary tests. Intentional provider runs require readiness and one-shot paid confirmation, then write separate live receipts.