Skip to main content
Antal István

Thoughts6 min read

How LLM reasoning works

You're billed for every thinking token, including the ones the vendor hides from you, and more thinking doesn't always mean a better answer.

An iceberg drawn in flat shapes. The small peak above the waterline carries a single short line, the answer, while the far larger mass below is packed with rows of blacked-out text: the working you never see. A yellow price tag hangs from the hidden part, its price drawn as a red bar.

Some problems need pen and paper. You write down the steps because there's too much to hold in your head at once. Reasoning models do something similar, using tokens to write out their working before answering.

On a paid API, you pay for that working too, even if all you see is a summary. A short answer can still have a lot of working behind it.

What reasoning actually is

Without a separate thinking step, the model starts writing the answer straight away. It predicts each token from your prompt and what it has written so far. It can still solve problems this way, but it doesn't write out any working first.

With a thinking step, it writes out its working first. It can recall more relevant facts, break the problem into smaller parts, try different approaches and check the results before answering. That gives it more to work with when it writes the answer, though a wrong step can also lead it astray.

It still predicts one token at a time, just as it does when writing the answer. Each new token becomes part of the input for the next one. Writing more tokens lets it spend more computing time on the problem and keep track of its results. Reasoning models are trained to do this well, without needing "think step by step" in your prompt.

Watching a model think

On a local reasoning model, we can read the working. Here's ornith-1.5-9b-mlx, a 9B model running in LM Studio on my machine:

Prompt

2 + 2 * 3 = ?

Only tell me the result.

Thinking

The user is asking a simple math problem: 2 + 2 * 3 = ?

Using order of operations (PEMDAS/BODMAS), multiplication comes before addition.

So 2 + 2 * 3 = 2 + 6 = 8

Answer

8

The trace reads like someone muttering under their breath. It recalls the rule it needs: multiplication before addition, giving 8 rather than the 12 you'd get by reading left to right.

Look at what "only tell me the result" did. It kept the answer to a single character without stopping the thinking. The working is much longer, and on a paid reasoning API you pay for that as well as the final answer.

Why the big vendors hide it

Anthropic and OpenAI use thinking tokens too, but show you a summary at best. They keep the full trace on their servers or send it back encrypted, so you can't read it.

I think the main reason is distillation: training one model on another model's answers. The full working gives it more to learn from. When OpenAI hid o1's chain of thought, it listed competitive advantage among its reasons.

Anthropic has reported distillation attempts by DeepSeek, Moonshot and MiniMax. One method asked Claude to write out the steps behind a finished answer, creating large amounts of training data. Anthropic now looks for these requests, and its newest models can refuse to include their internal reasoning in an answer.

A long scroll covered in lines of writing, locked inside a glass case with a yellow padlock. Beside the case sits a small card with just two lines: the summary you get instead.

Anthropic uses a separate model to summarise the thinking, and on its newest models even that summary is off by default. The full trace comes back encrypted in a signature field. You pass it back on the next request so the model can use its earlier working.

OpenAI doesn't show raw reasoning tokens either. You have to ask for a summary, and newer models may require you to verify your organisation first. For later requests, it keeps the reasoning on its servers or returns encrypted_content for you to send back.

Every thought is billed

Thinking tokens count as output, usually the priciest tokens on your bill. You pay whether you can read them or not: OpenAI bills reasoning you never see, and Anthropic bills the full trace, not the shorter summary it shows you.

A small speech bubble holding one short line, the answer, beside a long receipt listing line after line of thinking and ending in a red total.

So the tokens you're billed for won't match what you see on screen. Both APIs report how many tokens went on thinking in usage.output_tokens_details: reasoning_tokens at OpenAI and thinking_tokens at Anthropic.

The effort setting nudges how much work the model does. Higher effort usually means more reasoning tokens and a longer wait. At lower effort, some models can skip thinking on easy requests. Anthropic describes it as guidance rather than a fixed token limit, and so do OpenAI and Google.

In a conversation, the cost adds up. Newer Claude models keep the thinking from earlier replies and bill it again as input on later requests, at a lower rate if it's cached. The next post looks at the cost of each effort level and how caching changes the bill.

What hidden traces revealed

If a vendor reports 800 reasoning tokens and bills you for 800, you still have to trust that count. In August 2026, researchers found a way to read those hidden traces.

They sent an encrypted trace from Claude Opus 4.8 to Haiku 4.5 through Anthropic's API. The API decrypted it so Haiku could read it, and the researchers got Haiku to copy it into its answer. That example took two API calls. Similar attacks worked against OpenAI and Google, though they were harder and needed more attempts.

A torch shines onto a blue panel, and inside the beam, rows of hidden writing become visible.

Across 120 Codeforces programming problems, the recovered text often had about as many tokens as the API reported. But the researchers assumed the API counts were right and used them to judge how much of the trace they'd recovered. For OpenAI, they picked the versions with the closest token counts.

So this doesn't prove the bills are right. But finding all that working behind the answers is at least a hint that, for now, the labs are likely giving us honest token counts.

The traces also showed what summaries leave out. In one case, Claude gave the answer first and then checked it, but the summary made it look as though Claude had worked it out from scratch.

The team also found API keys and passwords in encrypted traces from agent logs people had shared publicly. The vendors blocked these attacks after the researchers reported them, but I'd still treat those blobs as private data.

More thinking isn't always better

You know the sitcom bit: a character has one simple message to send. They rehearse every way it could be read, spiral through every possible reply, and end up sending something worse than their first draft, which was fine all along.

Three paths leave the same dot: a straight green arrow hits the centre of the bullseye, while two red lines loop and knot their way across, one ending on the target's edge and the other falling short of it.

Models can make the same mistake. Researchers designed tasks where more thinking led to worse answers, with models getting distracted by details that didn't matter or spotting misleading patterns.

Even when it gets the answer right, a model can spend more time thinking than the task needs. One early study, Do NOT Think That Much for 2+3=?, found o1-style models spending lots of tokens on simple problems for little benefit.

Choose the lowest effort that works

Reading the full trace lets you spot repeated checks and dead ends. But getting the right answer early doesn't mean the checks that followed were a waste. Even the full trace may not tell you exactly how the model reached its answer.

With a closed model, you get a token count and perhaps a summary. Summaries focus on the main ideas and can leave out repeated steps. That makes it harder to tell whether those 800 tokens went on useful checks or on second-guessing a correct answer.

For either kind of model, try less effort and see whether the answers are still good enough. Run your usual tasks at low, medium and high effort, repeating each a few times. Compare the answers, the cost and how long they took, then keep the lowest level that works well. It's what the vendors recommend, and you can do it even when you can't read the working.

A dial with three marks of increasing size, like low, medium and high effort, turned all the way down to the smallest.

Related reading