Careful or Adventurous
Should a model play it safe or take a risk on its next word? The temperature dial, explained.
A quick note on how we teach here. You won't find equations or numbers in this series, and that is on purpose. The goal is understanding, not arithmetic, so every idea is explained in plain words and simple pictures. When you want the math later, the rest of Internals Decoded is ready for you.
Welcome back. Last time we saw how a model can turn words into places in a vast meaning space, so that words with similar meanings sit close together. Now we have one more idea to explore before all the pieces of the groundwork are in place. It is about a choice the model makes every time it writes the next word, and how we can nudge that choice.
The heart of it is this. A model can be careful and pick the most expected next piece, or it can be a little adventurous and sometimes pick a less expected one. That single decision shapes whether its writing feels steady and plain or surprising and creative.
The favorite dish and the surprise
Imagine you are at a restaurant you know well. You look at the menu. Some days you order your reliable favorite, the dish you have had a dozen times and always enjoy. Other days you point at something you have never tried, just to see what it is like. Neither choice is wrong. It depends on whether you want the comfort of the familiar or the thrill of something new.
A language model faces the same kind of moment every time it picks the next piece of a sentence. It always has a list of possible next pieces, each with a sense of how likely it is. The most likely piece is like your favorite dish. The less likely ones are the unfamiliar menu items. The model can always reach for the top choice, or it can sometimes reach a little further down the list.
How a model picks the next piece
You might remember from earlier in the series that a model does not think in whole words. It works in tiny pieces, little fragments of words that we called tokens. When the model is writing, it looks at the tokens it has already produced and asks itself: given all of that, which token should come next? It does not know the answer for sure. Instead, it assigns a likelihood to every possible next token. Some tokens feel very natural and get a high likelihood. Others feel possible but less obvious, so they get a lower likelihood.
If the model always picks the token with the highest likelihood, its writing will be safe. It will follow the most worn path every time. The sentences will make sense. They will be steady and predictable. But they might also feel a little flat or repetitive, like a friend who always tells the same story the same way.
If the model sometimes picks a token that is not the very top one, its writing becomes more varied. It might choose a word that is a little unexpected but still fits. That can lead to a delightful turn of phrase, a fresh way of saying something. It can also lead to a sentence that feels a bit odd or wanders off track. The model is taking a risk.
A dial for safe or surprising
Now picture a dial you can turn. Turn it all the way to one side, and the model becomes extremely careful. It almost always picks the most likely next token. The writing is solid, but it may use the same patterns again and again. Turn the dial the other way, and the model gets adventurous. It is willing to reach far down the list of possible tokens, even to ones that are only a tiny bit likely. The writing can become wildly creative, sometimes poetic, sometimes nonsensical.
The wonderful thing is that the model itself does not change. Its knowledge stays exactly the same. The only thing that changes is how willing it is to step off the beaten path. Same model, same understanding of language, same memory of everything it learned. Only its appetite for risk has been adjusted.
You can think of it like a kitchen with a single cook. The cook knows all the recipes. On careful days, the cook sticks to the classics, exactly as written. On adventurous days, the cook adds a pinch of this, a dash of that, maybe tries a combination nobody has asked for. The skill is the same. The results just feel different.
What people call this dial
When people who work with language models talk about this idea, they call the dial temperature. You first met that word way back when we talked about likelihood, and now it has a clear picture to live in. A low temperature means the model is careful. It hugs the most likely tokens tightly. A high temperature means the model is more adventurous. It spreads its attention across many possible tokens, even the long shots.
So when you hear someone say they turned the temperature down, they mean they asked the model to play it safe. When they turned it up, they asked the model to surprise them. That is all there is to it.
Where this is heading
The temperature dial is the last new idea in this series. You now have a home in your head for every building block the main AI series leans on. What comes next is a single final piece that ties all of these ideas together and sends you forward, proudly, into the world of large language models. After that, you will be ready to begin the main series, starting with a piece called "What an LLM Actually Is". Everything you have learned here will be right there with you, like a well packed bag for a journey you are finally ready to take.