Boss: operate under budget
Three requests come in, each with its own quality, cost, and speed budget. Can you set the levers so every one passes, without wasting money on the biggest model?
The idea inside
Operating a model is picking levers per request: model tier, extended thinking, and speculative decoding each move quality, cost, and latency a different way. Clear all three meters.
After this lesson
You can pick the operating levers for a request, model tier, extended thinking, and speculative decoding, to hit a quality, cost, and latency budget, instead of defaulting to the biggest model.
Where it leads
You can operate a model to fit any budget. But it's still frozen and remembers nothing between calls, so how does anything stateful get built around it?
Inside this lesson
That's the real lesson stage, paused. Claim your pass to operate it.
See how AI actually works, end to end.
This lesson is one stop on the full arc. Unlock all of it, and keep it for life.
What you get
- The 34-lesson main path, a finishable route from a word to agents
- Goal tracks for using AI at work and building AI features
- Boss labs that make you apply a whole act, not just recognize it
- Spaced recall that brings each idea back before you forget
- Course memory: every term defined, with links to where it first appears
- A shareable capability card when you finish the main path
- Lifetime access on every device, every future lesson included
Not videos to watch. You predict, operate the machine, then prove it. That is why it stays.
99 interactive lessons and challenges. No videos, no code.
Free launch pass: lifetime access, no card needed
New here? The first lessons are free to try. Start with lesson 0.1
What this lesson shows
Operating a model is picking levers per request: model tier, extended thinking, and speculative decoding each move quality, cost, and latency a different way. Clear all three meters.
The question it opens with
Three requests come in, each with its own quality, cost, and speed budget. Can you set the levers so every one passes, without wasting money on the biggest model?
The walkthrough, in the lesson's own words
- You didn't reach for the biggest model. You matched each request to its budget.
- Right call. No lever setting was safe enough, so a human takes this one. Send the next request.
- All three meters cleared. This config fits the budget. Send the next request.
- Set the levers so quality, cost, and latency all clear this request's budget.
- Toy meters, invented numbers that keep the trade-offs honest in shape, not scale.
- This request fits a budget: some lever setting clears all three meters. Escalation is for when no setting is safe enough. Keep working the levers.
- That's the strongest config on the bench and quality still misses the bar. When no lever clears it, the answer isn't a lever.
- Quality is the blocker: patient safety puts the bar at 99. Push the levers as high as they go and watch the ceiling.
- Red meter = over budget. Adjust the levers until all three read green.
- Same model, four calls. Cheap bulk went to the small tier, the hard case earned the frontier model plus thinking, realtime leaned on speculative decoding for speed, and medical triage had no safe lever at all, so you escalated it to a human. Operating a model is fitting levers to a budget, and knowing when not to automate.
Key takeaway
You ran one model three ways, matching each request's quality, cost, and latency budget, instead of reaching for the biggest model every time.
What you can do after this lesson
You can pick the operating levers for a request, model tier, extended thinking, and speculative decoding, to hit a quality, cost, and latency budget, instead of defaulting to the biggest model.
Where it leads: You can operate a model to fit any budget. But it's still frozen and remembers nothing between calls, so how does anything stateful get built around it?
This is the written summary. The lesson itself is interactive: you predict, drag and operate the mechanism above, and the reveal answers you.