Playground
Things easier to understand by moving than by reading. Nothing here calls the API or needs a key, so it is all safe to poke at before you have set anything up.
The Northwind storefront
A working shop. Browse the gear, file a complaint about it, and watch your own words get classified live.
Scenario
Northwind brand
The mark, lockup, and palette for the company the labs are built around. Includes the 18px test that killed three other concepts.
Scenario
The escalation queue
Where requires_human actually goes. Seven fictional escalations, their reasons, and the states a reviewer moves them through. Read-only without a token.
Lab 8
Priya's operations dashboard
The KPIs a support director reports upward, across a staged rollout. Simulated history, clearly badged, with the real unit economics alongside.
Scenario
The inbound queue
Twenty real tickets, before and after triage. Try to spot the safety report before you flip the toggle.
Scenario
Batch planner
The Batches API is half price and cost 23% more on this workload. Move the prefix size and find the crossover for yours.
Lab 9
The trust boundary
Toggle the escaping off and watch a customer message write its way out of the data block. Then meet the attack that escaping does nothing about.
Lab 8
Model matrix
The same twelve cases across three tiers. The accuracy column is the one that misleads you; the calibration gap is the one that decides anything.
Lab 7
Cost explorer
Move the volume, flip caching off, watch the budget bar go red. Measured token counts, real pricing.
Lab 5
Agentic loop stepper
Three turns, four tool calls. Watch context accumulate and see why logging the last turn under-reports cost by 3x.
Lab 3
Spot the cache bug
Four prompt variants. One caches. All four return 200 with a correct answer, which is the whole problem.
Lab 5