Skip to content
all lessons
Operating & Scaling4.10Locked

The thinking dial: when more hurts

If thinking helps, just crank it to max, right?

The idea inside

More thinking has diminishing then negative returns, so match the budget to the task.

After this lesson

You can reason about thinking budgets: when more reasoning helps, and when it wastes time and money.

Where it leads

Generating tokens is the bottleneck, can we make it faster?

Inside this lesson

That's the real lesson stage, paused. Claim your pass to operate it.

See how AI actually works, end to end.

This lesson is one stop on the full arc. Unlock all of it, and keep it for life.

What you get

  • The 34-lesson main path, a finishable route from a word to agents
  • Goal tracks for using AI at work and building AI features
  • Boss labs that make you apply a whole act, not just recognize it
  • Spaced recall that brings each idea back before you forget
  • Course memory: every term defined, with links to where it first appears
  • A shareable capability card when you finish the main path
  • Lifetime access on every device, every future lesson included

Not videos to watch. You predict, operate the machine, then prove it. That is why it stays.

99 interactive lessons and challenges. No videos, no code.

Free launch pass: lifetime access, no card needed

New here? The first lessons are free to try. Start with lesson 0.1

What this lesson shows

More thinking has diminishing then negative returns, so match the budget to the task.

The question it opens with

If thinking helps, just crank it to max, right?

The walkthrough, in the lesson's own words

  • Crank the thinking dial on an easy task. Watch cost and latency climb, does quality?
  • Thinking longer made the model smarter last lesson. So just max the dial, always?
  • Match the budget to the task: a lot for hard problems, a little for medium ones, almost none for easy ones.
  • Invented curves. Overthinking an easy task is a real effect, but usually a modest dip, not this cliff.
  • How many reasoning tokens the model is allowed before it answers.
  • On an easy task quality starts high and stays there, but cost and latency climb the whole way. Drag past halfway and watch quality dip. Then try the other tasks.
  • This task needed a little thinking and it has already flattened: past its sweet spot, extra budget buys almost nothing.
  • A medium task climbs fast: a small budget buys most of the quality. Watch for where the curve flattens.
  • Too much budget on a trivial task: the model can second-guess itself into an edge-case answer, while charging you more and making you wait.
  • A genuinely hard problem needs room to reason. Here the extra budget bought real quality, the cost and latency paid off.
  • On a hard task quality starts low. Keep dragging right and it climbs, this is where the budget earns its cost.
  • More thinking bought more accuracy last lesson, so just max it out, always?
  • More thinking is a dial, not a switch: it helps a hard problem, but on an easy one the model can talk itself out of the right answer. Every task has its own sweet spot, the point where its curve flattens, and it sits somewhere different each time.
  • Last lesson: spending more compute at answer-time, letting the model think longer, bought more accuracy. The obvious move is to turn that dial all the way up, all the time. It's not free, but smarter is worth it, right?
  • An easy task is already at full quality with zero thinking, it needs no working-out, so more reasoning just burns
  • You switch a reasoning model to its highest thinking setting for everything, and a simple yes/no question now comes back slower and sometimes wrong. Why?
  • Thinking is a dial, not a switch: past the sweet spot the model can talk itself out of an already-correct answer while you pay more cost and latency. Match the budget to the task, a lot for hard problems, almost none for easy ones.
  • Cost and latency rise the whole way. Does quality?

Key takeaway

You can set the thinking budget deliberately, more for hard problems, little for easy ones.

What you can do after this lesson

You can reason about thinking budgets: when more reasoning helps, and when it wastes time and money.

Check yourself: Is more 'thinking' always better?
  • No, it adds cost/latency and can make it overthink simple tasks(correct)
  • Yes, always crank it to max
  • Only ever for easy tasks
  • It never changes the answer

More thinking costs time and money and can make the model overthink easy tasks. Match the budget to how hard the question actually is.

Where it leads: Generating tokens is the bottleneck, can we make it faster?

This is the written summary. The lesson itself is interactive: you predict, drag and operate the mechanism above, and the reveal answers you.