Matches the derived interest "multi-day trips, starts before sunrise, carried" because 'multi' meets 'multi' in the product's title, category or attributes. Selected by the baseline arm, with no model call.
Nadia Brunner USR-NB-01 · CH · de · customer since 2019-03
I'm planning a few multi-day trips this year — hut to hut, carrying everything, usually out before sunrise. What should I be looking at? I'd rather hear what you think fits me than browse a category.
Bought before — 5 order line(s), and how the code read each one
You might also consider · 1
Matches the derived interest "multi-day trips, starts before sunrise, carried" because 'multi' meets 'multi' in the product's title, category or attributes. Selected by the baseline arm, with no model call.
This turn, counted
| Proposed → shown | 2 → 1 |
| Removed by a guardrail | 1 |
| Moved to the second tray | 1 |
| Purchases ruled out as gifts | 0 |
| Prices re-read from the catalogue | 1 of 1 |
| Tool calls spent | n/a — this arm makes no refusable tool calls |
What this page does not measure
Everything above is one turn. The claims that need more than one turn — does the loop cover more of a customer's interests than the single agent, does a hostile product review change what gets recommended, is the answer stable when the same turn is repeated — cannot be read off this page in either direction. The credential-free --all command runs deterministic offline gates, a benchmark scoped to exactly 2 synthetic cases/personas × 3 deterministic arms × 2 repetitions, and registered mutation controls. That benchmark's admitted checks record chance floor NotDerivable because no defensible random-answer chance floor is derivable; it does not run the paid multi-scenario or repeated-model plans:
dotnet run --project src/AgentEval.VitrineDemo.Evals -- --all
For those claims, select the corresponding live eval plan and pass --confirm-paid. Live plans declare their actual scenario and repetition scope and do not manufacture a per-arm chance floor where none is derivable.