Independent job-application sample for Digitec Galaxus · AgentEval in practice

From AI product discovery to measurable value.

What I can help Digitec Galaxus achieve: move an AI-assisted shopping experience from an impressive demonstration to a capability whose relevance, safety, variability, cost, and failure modes can be inspected before and after every change.

Independent portfolio project. VITRINE was created independently as a job-application value proposal for Digitec Galaxus. It is not affiliated with, commissioned by, or endorsed by Digitec Galaxus. Its catalogue, customer histories, and personas are synthetic. Its measurements are observations of this sample, not company systems, proprietary data, or production results.
01 · The proposal

Make quality a system, not a launch-day opinion

A product-discovery model can sound convincing while recommending an owned product, following a gift-derived false preference, inventing commercial facts, or behaving differently on the next execution. My contribution is to make those risks explicit, testable, observable, and reviewable by product, engineering, and safety stakeholders.

Product engineering

Build executable alternatives

A tool-using recommendation agent and a five-stage workflow operate over the same synthetic commerce domain, so control structures can be compared rather than debated in the abstract.

Evaluation engineering

Define independent evidence

Ground truth, criteria, tool expectations, missing-measurement rules, repetitions, and adversarial probes live outside the system being evaluated. The subject never grades itself.

Decision support

Leave a durable answer

Each execution lane retains the evidence appropriate to it: app runs emit typed events and support checksummed export and replay; the offline benchmark and Eval01–Eval05 create standard AgentEval run directories; every paid plan writes a sanitized local session receipt.

The AgentEval connection. VITRINE is also an applied demonstration of my work on AgentEval: admitted atomic checks, benchmark arms and cases, honest missingness, native comparisons, Wilson reliability, RedTeam probes, filesystem evidence, and fail-closed execution are used here against a realistic e-commerce subject—not presented only as framework APIs.
The adjacent knowledge-systems connection. agent-memory-dotnet is my separate Neo4j-backed, graph-native memory provider for Microsoft Agent Framework, with GraphRAG and MCP integration. It demonstrates related knowledge-graph and semantic-memory platform work; it is not a hidden dependency of VITRINE and VITRINE does not claim to implement Digitec Galaxus's knowledge graph, semantic layer, or Agentic BI systems.
99authored catalogue records
14authored personas
2subject architectures
5 + 1mandatory gates + matched-quality diagnostic
43registered control mutations
6explicit paid Eval01–Eval06 plans
02 · Value for an e-commerce company

Turn recurring AI questions into inspectable decisions

VITRINE does not claim to know Digitec Galaxus production internals. It demonstrates an evaluation method that can be adapted to governed catalogue, customer, search, safety, and business data.

Business questionVITRINE evidenceDecision enabled
Are recommendations grounded and useful?Scenario-specific ground truth, LLM criteria, catalogue checks, and deterministic tool expectations.Catch relevance, ownership, and fabricated-fact regressions before release.
Should this use case be an agent or a workflow?When explicitly invoked, Eval03 runs matched cases through both architectures and retains native paired comparison facts.Choose from observed quality, control, reliability, and cost trade-offs—not fashion.
Will the behavior remain stable?When explicitly invoked, Eval04 and Eval05 repeat provider-backed executions. Every selected arm/scenario persists a whole-trial 95% Wilson interval and must clear a 0.50 lower-bound floor; per-check intervals remain diagnostic evidence.Expose variability and decide whether the evidence is mature enough for a rollout.
Can the system resist misuse?The offline suite calibrates injection detection on guarded and deliberately vulnerable fixtures. When explicitly invoked, Eval06 targets fresh Robin instances with AgentEval jailbreak and hidden-instruction-extraction attacks.Prioritize containment work and rerun the same bounded safety contract after changes.
What did a run cost?Paid work is explicitly confirmed, planned calls are shown before execution, and provider usage is retained when reported.Compare quality per unit of spend and distinguish measured zero from missing telemetry.
Can we defend the decision later?Sanitized receipts, run IDs, persisted benchmark output, checksummed artifacts, HTML/JSON export, and pure replay. The checksum detects accidental changes; it is not a signature or proof against malicious replacement.Reproduce evidence for engineering review, incident learning, and model-change governance.
03 · Representative customer risks

Evaluate behavior through concrete shopping situations

The four cases are intentionally different. They prevent one generic “helpful recommendation” rubric from hiding distinct product and trust failures.

Nadia

Cross-category intent

Connect hiking, pre-sunrise lighting, off-grid power, and photography evidence while avoiding products she already owns.

Sofia

Capability gap

Recognize a missing coffee grinder while separating consumable replenishment from new durable-product discovery.

Marco

Gift-history trap

Suppress gifted gaming purchases as a preference signal and personalize only from Marco’s own espresso evidence.

Luca

Safe abstention

Treat one low-information cable purchase as insufficient evidence, ask useful questions, and avoid speculative recommendations.

Why authored cases matter: these are evaluation hypotheses, not claims about real Digitec Galaxus customers. In a production pilot they would be reviewed, replaced, or expanded using privacy-governed evidence and domain-expert input.
04 · Improvement loop

Use the cheapest trustworthy evidence first

The goal is not to call an LLM for every assertion. Deterministic contracts protect invariants quickly; paid evaluation is reserved for semantic quality, stochastic reliability, architectural comparison, and real adversarial behavior.

1 · Explore

Observe the subject

Run the agent and workflow, inspect tool calls, graph routes, fallbacks, screening, and final customer response.

2 · Guard

Offline CI

Apply catalogue, workflow, memory, honesty, injection, and recommendation predicates without provider spend.

3 · Measure

Paid use-case evals

Use scenario-routed LLM judging only after readiness and explicit cost confirmation.

4 · Stress

Reliability and safety

Repeat runs for Wilson intervals; compare architectures; probe jailbreak and instruction leakage.

5 · Learn

Persist and compare

Keep bounded receipts and native run facts, fix a failure, then rerun the same contract.

What AgentEval contributes

  • First-class checks with explicit applicability and chance-floor policy.
  • Cases, arms, repetitions, census, reference comparisons, and observation units.
  • Wilson reliability and honest underpowered/undecidable states.
  • Real RedTeam orchestration and filesystem-backed evidence.

What VITRINE demonstrates

  • Mapping commerce risks to independent ground truth and evaluators.
  • Instrumenting a real tool-using agent and conditional-loop workflow.
  • Making costs, fallbacks, missing data, and failure causes visible.
  • Turning framework primitives into an operator-friendly quality workflow.
  • Feeding application-discovered needs—such as safe, structured red-team failure diagnostics without raw evidence—back into my AgentEval engineering work.
05 · A credible adoption path

Start bounded, then connect evaluation to real outcomes

A production programme would be collaborative: product and category experts define useful behavior, safety and privacy owners set boundaries, and engineering makes the evidence repeatable.

30

Days 1–30 · Frame and instrument

Choose a narrow product-discovery journey; map failure modes and decision owners; define a governed scenario registry; instrument tool, retrieval, screening, latency, cost, and fallback facts; establish deterministic baselines.

60

Days 31–60 · Calibrate and compare

Connect read-only staging data; have domain experts review rubrics and failure examples; run blinded agent/workflow comparisons; calibrate judge agreement; add offline release gates and explicit missing-measurement policy.

90

Days 61–90 · Pilot and govern

Run scheduled model-backed canaries and bounded RedTeam probes; analyze reliability and spend; define model or prompt change gates; connect technical quality to privacy-approved business measures such as successful discovery, add-to-cart, support contacts, returns, and customer feedback.

Deliverable: not merely a dashboard—a repeatable decision process with owned scenarios, versioned criteria, bounded execution, retained evidence, failure triage, and a clear path from a red result to remediation.
06 · Evidence available today

Inspect the implementation, not just the proposal

The current receipts are engineering evidence from a synthetic sample. They demonstrate that the contracts execute and can detect planted defects; they are not Digitec Galaxus business metrics.

Verified offline

Release confidence

At the 2026-09-10 verification snapshot: Release build with zero warnings/errors; 450/450 credential-free tests; all five mandatory evaluation gates passed and the matched-quality diagnostic was reported; 43/43 registered mutations detected; catalogue defect detected and recovery verified.

Durable

Run evidence

The checked-in offline report records cases, arms, checks, census, paired comparisons, gates, controls, and the exact canonical run identity.

Implemented, explicitly paid

Live measurement

The live-plan orchestration and all selectors are exercised with injected fakes in ordinary tests. Intentional provider runs require readiness and one-shot paid confirmation, then write separate live receipts.

07 · Honest boundary

What this proposal does not claim

Not claimed

Production business impact

  • The 99 products and 14 personas are synthetic, not production scale.
  • No conversion, revenue, margin, return-rate, or customer-satisfaction uplift is shown.
  • No Digitec Galaxus customer data, internal service, or proprietary prompt is used.
  • No paid live-plan result is presented here as portfolio evidence of product quality; each operator-run measurement keeps its own local receipt.
  • No current paid comparison establishes that the agent or workflow is superior.
Not certified

Statistical and security certainty

  • Configured small stochastic samples illustrate uncertainty; they do not establish production reliability.
  • The default four-probe Eval06 campaign is a bounded evaluation, not penetration testing or certification.
  • The 43 controls prove registered mutation reachability, not a natural defect rate.
  • Historical sample measurements are not Digitec Galaxus operational metrics.
The claim I am making: I can design and implement the engineering loop that makes an AI commerce capability easier to evaluate, compare, debug, govern, and improve—and I can communicate its evidence and limitations clearly.