Skip to main content

Playground

Things easier to understand by moving than by reading. Nothing here calls the API or needs a key, so it is all safe to poke at before you have set anything up.

The Northwind storefront

A working shop. Browse the gear, file a complaint about it, and watch your own words get classified live.

Scenario

Northwind brand

The mark, lockup, and palette for the company the labs are built around. Includes the 18px test that killed three other concepts.

Scenario

The escalation queue

Where requires_human actually goes. Seven fictional escalations, their reasons, and the states a reviewer moves them through. Read-only without a token.

Lab 8

Priya's operations dashboard

The KPIs a support director reports upward, across a staged rollout. Simulated history, clearly badged, with the real unit economics alongside.

Scenario

The inbound queue

Twenty real tickets, before and after triage. Try to spot the safety report before you flip the toggle.

Scenario

Batch planner

The Batches API is half price and cost 23% more on this workload. Move the prefix size and find the crossover for yours.

Lab 9

The trust boundary

Toggle the escaping off and watch a customer message write its way out of the data block. Then meet the attack that escaping does nothing about.

Lab 8

Model matrix

The same twelve cases across three tiers. The accuracy column is the one that misleads you; the calibration gap is the one that decides anything.

Lab 7

Cost explorer

Move the volume, flip caching off, watch the budget bar go red. Measured token counts, real pricing.

Lab 5

Agentic loop stepper

Three turns, four tool calls. Watch context accumulate and see why logging the last turn under-reports cost by 3x.

Lab 3

Spot the cache bug

Four prompt variants. One caches. All four return 200 with a correct answer, which is the whole problem.

Lab 5