Back in Lesson 2 we admitted something: our boards were five slots wide and three deep so you could actually see them, but a real one carries billions of coloured patches, across far more boards than three — not a different idea, just a bigger one. Now that we know weights and tokens, we can put real numbers on that.
Every real AI is built from exactly the parts we already have: a model (the board and its pegs), weights (the painted floor), layers (boards stacked on boards), tokens for an alphabet. Nothing changes at scale except the amount of it.
That's worth holding onto, because the numbers ahead are large enough to sound impressive on their own. They're not a different kind of magic — they're the same model, the same weights, the same training, just multiplied enormously.
Each coloured patch on the floor — each weight — counts as one parameter. Add up every patch across every stacked board, and you get the number people quote when they talk about a model's size: three billion, thirty billion, and up.
Remember the arithmetic from Lesson 2: our own board, a simple five-by-five grid, already added up to about a hundred parameters — and that was one small board, before stacking a second one on top. Real ones aren't built differently. They're just built at a size where "a hundred" becomes "a hundred billion."
Three broad tiers are worth having a feel for:
| Hardware tier | Roughly | What it looks like |
|---|---|---|
| A laptop, even a phone | ~3 billion | A small, genuinely useful assistant, running entirely on your own device. Open, published examples at this size come from families like Llama, Qwen, and DeepSeek — you can download one today. |
| A strong home PC or Mac | ~30 billion | A noticeably more capable model, still private and offline, but now needing a serious graphics card or a Mac built for it. Larger versions of the same open families live here too. |
| Data-centre scale | ~500 billion – several trillion | Where the largest models live — general-purpose services like ChatGPT and Claude, and the largest open-weight releases. Closed models' exact sizes are still never published. Open-weight releases do publish theirs, though, and the largest published figures are now well into the trillions of parameters — a number that keeps climbing. |
Three billion to thirty billion is a real, meaningful jump — about ten times the floor to paint. Thirty billion to a few trillion is a far bigger jump again. Same model. Vastly more of it.
Two very different big numbers both get measured in tokens — worth telling apart.
What fits in view at once is the size of the funnel — how many tokens of conversation the model can hold and reread in a single pass. Ordinary chats use a few hundred to a few thousand. But the funnel itself, in the newest systems, can now be enormous: some can hold several hundred thousand tokens at once, and a few stretch past a million. That's the capacity, not what a normal conversation actually uses.
What it read during training is a categorically bigger number. That's the mountain of text — books, articles, code, conversations — the model was shown once, patiently, while its floor was being painted. That figure is usually measured in the trillions of tokens: a thousand times bigger again than the largest funnel. It's also finished and frozen, exactly like we said in Lesson 2 — read once, long before the model ever reached you.
It's tempting to treat a model with trillions of parameters as something categorically other than the small one on your desk. It isn't.
It's more patches, more layers, more tokens of training — the exact same idea, carried out at a size no drawing could ever hold. That's also the answer to "why is the free model on my phone worse than the one in the cloud": a smaller floor, painted with fewer patches, over less reading. Not a lesser kind of intelligence — a smaller model.
You've had all three words for a while without anyone naming the thing they spell.
Model — the board itself, from Lesson 1. Language — the tokens it reads and writes, from Lessons 3 and 4. Large — this lesson: the billions of parameters, the trillions of tokens of training. Put the three together and you get the term you've almost certainly heard: a Large Language Model, or LLM. That's all it is — the exact model you've been building this whole time, just with its full name attached.
Billions of parameters, trillions of tokens of training — and it is still, underneath, a model making a prediction. Predictions can be wrong.
The next lesson goes back to that pile of runners-up from Lesson 3 — mouse tallest, but fish and bird right behind it — and shows exactly where a confident mistake comes from.
Everything from the lessons before this one, plus what this lesson added.
| The board | The word for it |
|---|---|
| The board and its pegs | The structure — the model |
| A single peg | A neuron |
| A single step from one peg to the next | A connection |
| How red or blue a patch of floor is | A weight — a parameter |
| The whole painted floor | The trained model |
| Boards stacked, one on the next | Layers — "deep" |
| Smearing the floor in, coin by coin | Training |
| Dropping one coin through a finished board | Using it — inference |
| Which slot you pour into | Your input — your prompt |
| The bins, relabelled with tokens | The vocabulary — every token it could pick |
| The tallest pile among the bins | The prediction — the likeliest next token |
| Feeding the growing line back in and running again | The loop — how one guess becomes a sentence |
| One bin's label, precisely | A token — not always a whole word |
| How many bins there are, in total | Vocabulary size — tens of thousands |
| The size of the mountain of text it was shown | Training data — measured in tokens |
| A bin's own address, plotted in space | An embedding — meaning turned into a list of numbers |
| How close two addresses sit | Similarity — how related two things are |
| The measurement's real name, out in the wild | Cosine similarity — same idea as similarity, different name |
| About meaning, not exact spelling | Semantic |
| A neighbourhood of addresses that share a topic | A domain — a cluster of related meaning |
| Scanning every address for whichever one sits closest to a new point | Nearest neighbor — the closest real match to a computed point |
| A small, separate machine that only plots words as points — it never writes a sentence | An embedding model — not the same as a language model |
| + New in this lesson | |
| The total number of coloured patches, across every board | Parameters — the number people quote |
| How many tokens the funnel can hold in one pour | Input tokens — the context window |
| How many tokens come back out | Output tokens — what it writes back |
| Everything you've met so far, built at real-world size | A Large Language Model — an LLM |
26 rows now. It keeps growing as the course goes on.