Claim boundary
This is not a Digitec Galaxus measurement. It is a self-attested aggregate receipt from a historical paid-model observation of an earlier synthetic, same-author e-commerce sample. The public artifact validates its own schema, arithmetic, and checksum, but does not retain raw cases, provider receipts, or enough provenance to independently establish the observation's origin. The catalogue, personas, questions, subjects, and expected answers were synthetic. It establishes neither production quality nor company endorsement.
Measurement identity
| Measurement ID | vitrine-synthetic-live-2026-09-04-eval02b-02c |
|---|---|
| Scope retained | Aggregate stated-need and next-purchase observations only |
| Excluded | Raw prompts, responses, provider configuration, credentials, endpoints, and internal notes |
Stated-need result
Twelve authored cases (SN-01 through SN-12) were measured. The live synthetic subject scored 0.889; a zero-model tag-join oracle scored 1.000 at matched k. Mean matched k was 1.92, seven cases matched exact k, and 23 recommendation slots were observed.
The honest claim is deliberately narrow: the model-backed approach did not beat the simple oracle on this small authored set.
Next-purchase result
Thirteen pairs produced 3 wins, 1 loss, and 9 ties: only 4 informative pairs. The exact two-sided sign-test value was 0.625. At four informative pairs the minimum attainable two-sided p-value was 0.125, above α = 0.05. The experiment therefore cannot support a next-purchase superiority claim; more informative pairs are required.
Machine-checked copy
The aggregate values are duplicated in src/AgentEval.VitrineDemo.Evals/Data/HonestyEvidence.v1.json. The offline honesty gate pins that file by SHA-256, recomputes the exact-test quantities with AgentEval, and fails closed if the artifact or its claim wording is changed.