When you ask a chatbot a question, the answer seems to arrive as a thought-out reply. Underneath, it is built by repeating one small step: look at all the text so far, guess what comes next, add it, repeat. This page shows that loop with a tiny model you can drive yourself, explains what "temperature" does, and why a chatbot can sound completely sure of something that isn't true.

One small step, repeated

The chatbots you've heard of (ChatGPT, Claude, Gemini and friends) are built on a large language model, or LLM. "Model" here just means a program whose behaviour was learned from examples instead of written out as rules. "Language model" means the thing it learned is: given some text, what is likely to come next?

That's all it does. To write a whole reply, the chatbot runs a loop:

The generation loop

1Read all the text so far
2Score every possible next piece
3Pick one piece
4Stick it on the end, go to 1

The loop stops when the model picks a special "end of reply" piece, or hits a length limit. This is why answers often appear word by word on screen: you are watching the loop run.

Two words in that loop need explaining: what is a "piece", and how does it "pick"?

Tokens: the pieces a model reads and writes

A model doesn't work with letters, and not quite with words either. Before anything else, text is chopped into tokens: common chunks of characters. A frequent word like the is usually one token (often including the space in front of it). A rarer word gets split into several pieces, like un + believ + ably. Each token in the model's vocabulary has a number, its token ID, and those numbers are what the model actually sees.

The program that does the chopping is called a tokenizer. Real ones learn their vocabulary automatically from huge amounts of text, keeping the chunks that appear most often, and end up with somewhere between tens of thousands and a couple of hundred thousand tokens. The one below is a toy: a hand-written list of a few hundred pieces, so the exact splits won't match any real chatbot. But it behaves in the same spirit. Type something and watch it get cut up.

Toy tokenizer · a simplified illustration

0characters
0words
0tokens
0chars per token

Each coloured box is one token; · marks a space that belongs to the token. A red dashed box is a chunk our toy didn't know, so it fell back to small letter groups. Real tokenizers handle capitals, emoji and every language, and usually treat Hello and hello as different tokens. This one ignores case to stay small.

A few things you can see, and that are true of real tokenizers too:

Tokens matter in practice too: chatbot and API limits ("this model can read up to N tokens") and API prices are counted in tokens, not words.

Guessing the next token

Now step 2 of the loop. Given the tokens so far, the model doesn't produce one answer. It produces a probability for every token in its vocabulary: a number between 0 and 1 saying how likely that token is to come next. All of them add up to 1 (100%). After "The capital of France is", a well-trained model puts a very high probability on Paris and tiny probabilities on everything else.

Where do the probabilities come from? A real LLM is a huge mathematical function, a neural network with billions of adjustable numbers, tuned by showing it enormous amounts of text and nudging it, again and again, to give a higher probability to the token that actually came next. (Chatbots then get extra training on example conversations and human feedback so they behave like a helpful assistant, but the next-token machinery stays the same.)

We can build a much simpler version that works on the same idea: read a small text, and for every word, count which words came right after it. That's called a bigram model ("bigram" = a pair of neighbouring words). Our model learns from just these twelve sentences (the full stop counts as a word, so it can learn where sentences end):



After the word sat it has only ever seen on, so it's 100% sure. After the it has seen many words, so the probability is spread out: each candidate's share is how often it followed divided by how often the appeared. Press Next word to run the loop one step at a time.

Toy next-word predictor · bigram model

1.0
Show what the model "knows" (all its counts)

Notice how little this model has to go on: it only ever looks at the one word before. That's why it can wander into loops like "the cat chased the dog chased the ball". A real LLM looks at the whole conversation, often many thousands of tokens, and has learned far subtler patterns than "which word follows which". The loop is the same; the guessing is enormously better.

Picking a token: sampling and temperature

Once the model has its probabilities, something has to choose. The simplest rule, called greedy decoding, is to always take the top one. That sounds sensible, but it makes the text repetitive and the same every time. So most chatbots sample instead: they pick at random, but weighted by the probabilities. A token with 60% gets picked about 60% of the time. In the predictor above, the striped bar under the probabilities is exactly that: a random number between 0 and 1 is "thrown" at the bar, and whichever word's stretch it lands in wins.

The temperature setting reshapes the probabilities before the throw:

Drag the temperature slider in the predictor and watch the bars after the squash or stretch. (For the curious: each probability is raised to the power 1/temperature and then they're rescaled to add up to 100% again. That's the same as the usual definition, which divides the model's raw scores by the temperature.) Many chatbots let developers set it; the apps you chat with choose a value for you, and often combine it with other tricks, such as ignoring the long tail of very unlikely tokens.

Because of sampling, asking the same question twice can give different answers. That's not the chatbot "changing its mind". It's a different roll of the dice at some step, and every later step builds on whatever got picked.

Why it can sound confident and still be wrong

Here's the important part. Look at what the model is trained to do: produce text that is likely, given the text before it. Not text that is true. Most of the time those overlap, because the text it learned from is mostly sensible. But they are not the same thing, and our tiny model shows the gap clearly. Generate some sentences:

Write 5 sentences · same bigram model

temperature

copied appears word for word in the training text. new is a fresh mix that the training text never says. false is a new mix that says something we know is untrue. The model can't tell these apart: all it knows are the counts.

Sooner or later you'll get something like the moon is made of milk . or the sun orbits the earth . Every single step was a perfectly reasonable guess: "made of" really is followed by "milk" in the training text. The sentence is fluent, grammatical, and wrong. Nothing in the loop ever checked it against the world.

Real LLMs are vastly better at this, but the same basic gap is there, and it has a name: hallucination, when a model states something false or made up as if it were fact, like a quote nobody said or a book that doesn't exist. A few reasons it happens:

Chatbot makers fight this in several ways: training models to say "I'm not sure", letting them search the web or read documents you provide (and quote them), and testing them hard on facts. These help a lot, but none makes a chatbot a trustworthy source by itself. The practical rule: use chatbots for drafts, explanations and ideas, and check anything that matters, like numbers, quotes, names, laws, medical facts and code, against a real source or by running it.

Is it "just autocomplete"? Partly. The loop really is next-token prediction, but to predict well across billions of examples, a big model has to pick up a great deal about grammar, facts, and how reasoning steps usually follow each other. That's why it can be genuinely useful and confidently wrong. Both come from the same machinery.

Check yourself

A chatbot has written "The capital of Italy is". What does the model produce next?

Each step gives a probability for every possible next token (here " Rome" would get a very high one). Then one token is picked, added, and the loop runs again. There's no fact database being looked up.

You set temperature to 0 and ask the same question three times. What happens?

At 0 it always picks the top token, so the output is (in practice, very nearly) repeatable. But "most likely" is not the same as "true": a greedy answer can be confidently wrong too.

Why might "Antidisestablishmentarianism" count as several tokens while "the" is one?

The vocabulary keeps chunks that appear often in the training text. "the" is everywhere; a rare long word has to be spelled out with several common pieces. The "4 characters" figure is only an average for English.

A chatbot gives you a detailed, confident answer with a reference to a research paper. What's the sensible next step?

Confidence is a writing style the model learned, not a measure of truth, and made-up references are a classic hallucination. Asking it again is just another round of the same prediction. Check the real source.

The short version

Next time you watch a reply appear word by word, you'll know what you're seeing: the same small guess, made over and over, very, very well.