Session 1 · Part A · intuition, math-light
How language models actually work
I'll spend this session on one trick: predicting the next token, over a library of text bigger than any lifetime of reading. By the end you can explain the trick to your own lab, and you know where it breaks.
Section 01
A language model is a very well-read autocomplete
Everything you'll see this fall, from code that writes itself to summaries of a thousand papers, rests on one core mechanism. The model looks at the text so far and predicts what comes next. It reads and writes one chunk at a time, over and over.
That sounds too small to matter. The surprise of the last few years is what the trick looks like when you do it extremely well, over more text than any person could read in a hundred lifetimes. My goal is that you leave able to trace any answer a model gives you back to this one mechanism.
How would you finish this sentence? Call it out, then let the model take its turn; we'll see what four short paragraphs of reading gets it.
Try it. The model below read four short paragraphs of biology text and counted which letters follow which. Give it the start of a sentence, then let it take over.
The continuation will appear here.
It has no idea what PCR is. It still writes biology-shaped words, because four short paragraphs of counting were enough to teach it which letters tend to follow which. Hold onto the gap between what it knows and what it produces; hallucination is where that gap bites.
Section 02
Text becomes numbers
Before a model can predict anything, it has to read. Models read in chunks called tokens. A token is a common cluster of characters: everyday words get their own token, rarer words get split into pieces. By the time text reaches the model, it has become a list of numbers, one ID per token, and that list is the model's entire view of everything you typed.
Tokenization explains a family of quirks that otherwise look like stupidity. Ask a model how many letters are in deoxyribonucleic and it may miscount, because it never saw letters at all; it saw opaque ID numbers. Exact strings from the lab bench hit the same wall: the model sees token IDs, not characters, and one wrong letter changes what a reagent or a variant means. Watch what a production tokenizer does to vocabulary your lab uses daily.
Where do the boundaries land inside deoxyribonucleic? How many pieces? Guess before you tokenize it.
Type anything, then tokenize it. Watch where the boundaries land, especially on long bio words.
The tokens will appear here.
Every chip is one token, with its ID number underneath. The model's whole world is sequences of these IDs. Our toy model reads one letter at a time so you can watch it work; production models read tokens. It's the same trick with a different chunk size.
Section 03
A probability list, then a dice roll
At every step, the model does one thing: it assigns a probability to every possible next token. Prediction means picking from that list, and picking, called sampling, means rolling a weighted die. That die is why the same prompt can give you a different answer tonight than it gave this morning.
Temperature is the knob that reshapes the list before the roll. Low values make the model cautious and repetitive. High values make it adventurous and strange. When you use a chat app, the app hides this knob from you. When you call a model through its API, the way software talks to a model directly instead of through a chat window, you set it yourself, and it belongs in your methods section.
Roll it twice at the same temperature. Hands up if you expect the same pick both times.
Roll it yourself. The bars show the model's top candidates for the next character after "The", reshaped by the temperature you set. The highlighted one is what the die picked this time.
The probability list and the pick will appear here.
Roll it several times without touching the slider. The bars stay put and the pick moves. That, in miniature, is what "the model is probabilistic" means for the reproducibility paragraph of your next paper.
Section 04
Where the numbers come from
The model's probabilities come from training. During training, the model makes a guess at every position in the text, compares its guess with what the text actually says next, and nudges itself a little toward the truth. Billions of nudges later, the patterns of the library live in the model's parameters, the stored numbers training adjusts.
This is why training data matters so much: the model can only mimic the distribution of the text it was shown. Our toy corpus is four short paragraphs of biology. A frontier model's library is closer to the whole readable internet. The mechanism is identical, and so is the catch: whatever patterns sit in the library, flawed and biased and outdated alike, end up in the model. Using the trained model afterwards is called inference.
Call it before the drag: at fifteen percent of the library, will it write soup, English-shaped soup, or biology?
Drag the slider to decide how much text the model has read. Then look at two things: how surprised it is by text it has never seen, and what it writes when prompted.
The falling bar is training. The changing text is the payoff: soup, then English-shaped soup, then biology-shaped sentences with the facts slightly scrambled. Scale that same loop up by a factor of a billion, and you get the frontier models.
Section 05
It looks like magic because the library is bigger than a lifetime of reading
A language model mimics the training distribution of text, and that distribution is bigger than anything any of us could read in a lifetime. So when you ask about a niche proteomics method at midnight and a fluent answer comes back, the question feels out of the ordinary to you. For the model, the question is ordinary: millions of pages in the training text say similar things.
That's why the right mental picture is interpolation, not retrieval. The model itself never looks anything up. Every answer lands somewhere inside the cloud of training text, and the cloud is so vast that almost everything a scientist can think to ask already lives inside it. The exceptions live at the edge of that cloud, and the last section walks through them. (When a chat app searches the web mid-answer, that's a separate tool stapled onto the same predictor, and Part B covers where that line sits.)
Give me the most niche question in your field, the one you're sure nobody has written down. That's the one I'll ask.
Ask something that feels out of the ordinary and watch where it lands. Each dot is a piece of text: the small dense cluster is what one careful researcher reads in a career, the vast cloud is what the model read during training.
Questions that feel novel to us still land inside the cloud; that's what in-distribution means. The cloud has an edge, though: training stopped on a date, and that edge is where the exceptions begin.
Section 06
Confidently wrong is the mechanism working normally
A hallucination is the model doing exactly what it was trained to do, on a question where the fluent continuation happens to be false. There's no lookup table to consult and no alarm that rings when a prediction strays past what the model actually knows. There's only the next most plausible token.
This is what makes the failure mode dangerous in research: fabricated references arrive in perfect formatting, with plausible author lists, real journals, and tidy page numbers. Two of the five references below are fabricated; the other three come straight from the bootcamp's reading list.
A reviewer sent you this paragraph. Call each reference real or fabricated on its card below, then check your answers.
Hepatic JNK signaling has become a central node in metabolic disease research. Loss of JNK1/2 in the liver reshapes the transcriptome through the PPARα–FGF21 axis , and chronic JNK blockade has been reported to protect against diet-induced steatosis through FGF21-independent routes . In inflammatory skin disease, single-cell profiling of vitiligo showed an expansion of effector CD8⁺ T cells, with CCR5 positioning regulatory T cells to restrain disease . More recently, a single-cell atlas of photosensitive skin described MMP9⁺ myeloid circuits organized by keratinocyte–fibroblast crosstalk , consistent with the spatial mapping of a keratinocyte→fibroblast→myeloid axis driving photosensitivity .
-
Vernia S., et al. (2014). "The PPARα–FGF21 hormone axis contributes to metabolic regulation by the hepatic JNK signaling pathway." Cell Metabolism 20(3):512–525.
-
Marquez-Herrera R., et al. (2019). "FGF21-independent hepatoprotective effects of chronic JNK inhibition in diet-induced steatosis." Cell Metabolism 29(4):912–925.
-
Gellatly K.J., et al. (2021). Vitiligo single-cell study: effector T cell expansion and CCR5-dependent positioning of regulatory T cells. Science Translational Medicine.
-
Okonkwo A. & Reyes T. (2022). "Single-cell atlas of photosensitive skin reveals MMP9⁺ myeloid circuits." Nature Immunology 23(11):1745–1758.
-
Wang Y., et al. (2026). Spatial mapping of a keratinocyte→fibroblast→myeloid circuit driving photosensitivity. Nature Immunology.
From inside the model, every citation is equally fluent. Fluency is its native output, so fluency can never be your evidence. Checking each reference against the primary source is the habit that catches fakes; Part B turns that habit into concrete rules.
Section 07
What changes when everything gets bigger
Same recipe, more of everything: more parameters to hold patterns, more text to learn them from, more compute to push the training loop. For years the gains looked politely incremental: better grammar, fewer nonsense words. Then, somewhere past a few billion parameters, abilities showed up that nobody had explicitly programmed: following instructions in plain language, picking up a task from a couple of examples in the prompt, working through a problem step by step: what researchers call emergence. Researchers still argue how sharp those jumps really are; that they arrived with scale rather than with code is the part everyone agrees on.
Drag across ten orders of magnitude and watch what the same mechanism picks up along the way. The counts are public through GPT-3; newer labs keep their numbers quiet, so treat the right edge as "more, undisclosed".
-
1940s–now
Common letter patterns
It counts which chunks follow which. The toy model on this page lives here, at a few thousand counted parameters.
-
2013
Word meanings as geometry
Words become long lists of numbers, and similar words get similar lists, so word math starts to work: king − man + woman lands near queen.
-
2018
Fluent paragraphs
GPT-1, 117 million parameters. The transformer recipe starts writing decent prose on almost any prompt.
-
2019
Whole essays, wobbly facts
GPT-2, 1.5 billion. It writes multi-paragraph text that reads clean, though the facts arrive loose.
-
2020
Learns from examples in the prompt
GPT-3, 175 billion. Show two examples, ask for a third. Nobody programmed this; researchers still debate the shape of the jump, but it arrived with scale.
-
2022
Follows plain-language instructions
Chat-era models are tuned to do what you ask instead of completing text. The apps in your browser tabs live here.
-
frontier
Reasons step by step, drives tools
Sizes stay undisclosed, but multi-step reasoning and tool use show up at this end of the slider.
One thing scale didn't buy: the dice and the library's edge survive every order of magnitude, so fluent and true stay different words at any slider position. Emergence at scale is the bridge to foundation models, and the reason your field builds one shared model instead of training a fresh network for every question.
Section 08
Train once, use everywhere: foundation models
You now own every piece of the vocabulary: tokens, next-token prediction, training, sampling, distribution, scale. A foundation model is what you get when you push that recipe far enough on broad data that one model adapts to many jobs. Pretrain once on the giant library, then adapt: ask in plain language, show a few examples, or fine-tune on your own data.
The same idea left text behind. Swap what counts as a token and the same sequence-modeling machinery learns biology: amino acids, genes, regulatory DNA. Open the cards that matter to your science.
Hands up: proteins, genes, regulatory DNA, or single cells? Your votes pick which cards I open.
Open a card to see what goes in and what comes out; the block in the middle is the same in every card, which is the whole point.
Text LLMsGPT, Claude, Llama
Sections 1 through 7 are this card. Everything else here is the same recipe pointed at a different kind of sequence.
AlphaFold 2protein structure, 2021
The model treats a protein's sequence as a language and reads the structure out of the patterns across related sequences.
Geneformernetwork biology, 2023
Tokens are genes; predicting missing genes becomes predicting which genes belong together in a working cell.
scGPTsingle-cell analysis, 2023
One pretrained model handles many single-cell chores. The Session 5 and 6 analyses run on models like this one.
Enformergene regulation, 2021
It reads a stretch of regulatory DNA, the genome's control panel, and predicts how the nearby genes respond in each cell type.
The session Manuel Garber leads takes these tools into spatial biology. When you get there, it assumes the vocabulary you've built today, and you already have it.
Section 09
Know where the trick breaks, then use it anyway
Every failure mode you'll hear about this fall traces back to something you just worked through. Once you own the mechanism, failures become predictable, and each one comes with a habit that catches it.
| What you'll see | Where it comes from | The habit that catches it |
|---|---|---|
| The model makes something up | Prediction without lookup | Verify claims against primary sources |
| Same prompt, different answer | Sampling, the dice roll | Log model, settings, and prompts; expect wobble |
| Confident about a world that changed last month | The library has an edge date, and nothing you type adds to it | Check current sources for anything time-sensitive |
| Repeats the field's old mistakes | Mimics the training distribution | Read outputs with the literature's biases in mind |
| Echoes whatever sensitive text you pasted | Its entire view of you is the text you hand it; nothing in the mechanism keeps a secret | Keep patient and unpublished data out of consumer tools |
Before you go, I'll show you four failure reports, and you pick the idea that explains each one.
1. You ask the same question twice and get two different protocols.
2. Your paragraph cites a paper that turns out not to exist.
3. You ask about a preprint posted yesterday and get a summary that's fluently wrong.
4. Your analysis suggestion uses a statistical method the field moved away from decades ago.
The frame I find most useful: treat the model as a collaborator. Collaborators do brilliant work, make honest mistakes, and expect their work to be checked. The mistake you catch is the one you looked at.
A language model predicts the next token over a library bigger than anyone could read in a lifetime, so fluent answers come from interpolating within that library's patterns. They land outside the truth when the most plausible continuation happens to be false, because nothing in the mechanism checks truth.
Part B, led by Alper, turns the habits in that table into concrete rules: protecting patient data, verification discipline, disclosure, reproducibility, and the cases where you leave the model out of it entirely. He has war stories, and they're worth staying for.
Appendix
Take it with you
Every session is recorded, so nothing here depends on having been in the room. Here is the vocabulary in one place, the questions we expect in the room, and where to go deeper if you want more.
The vocabulary
- Token
- The chunk of text a model reads: a word, a piece of a word, or a punctuation mark, mapped to an ID number.
- Next-token prediction
- The whole trick: look at the text so far, assign a probability to every possible next token.
- Probability distribution
- The full list of probabilities the model assigns, one for each possible next token. The model's answer to every question is one of these lists.
- Sampling
- Rolling the weighted die: picking one token from the distribution. This is why answers vary.
- Temperature
- The knob that reshapes the distribution before the roll. Low is cautious, high is adventurous. Set explicitly in API use, hidden in consumer apps.
- Parameters
- The stored numbers inside the model; training adjusts them and every answer runs through them, so scale just means more of them.
- Training
- The expensive loop: guess the next token everywhere in the library, compare with reality, nudge the parameters.
- Inference
- Using the trained model to answer. Cheap compared with training; this is what happens when you chat.
- Training distribution
- The shape of everything the model read. The model mimics it, so it can only be as current and fair as its library.
- Interpolation
- Landing inside the cloud of patterns the model has seen. The right default picture for how it answers; it retrieves nothing.
- Hallucination
- The same mechanism, fluently wrong. A plausible continuation that happens to be false.
- Fine-tuning
- Extra training on your own data to specialize a pretrained model.
- Foundation model
- One model, trained on broad data at scale, adapted to many jobs: text, protein structure, gene expression, regulation.
Questions we expect in the room
Is it just autocomplete?
The mechanism is next-token prediction; the interest comes from doing it superbly over a library the size of the readable internet. Be careful with "just" in both directions: it deflates the hype and undersells the risk.
Does the model learn from our conversation?
The parameters freeze after training. Within one chat, the model rereads the whole conversation each turn, which feels like learning but is really a longer prefix. Whether the vendor trains on what you type is a separate data-policy question; the short rule is to assume your prompts are retained unless you're on a plan or deployment that says otherwise, and Part B walks through how to check for the tools your lab actually uses.
Which model should my lab use?
There's no single right answer; the honest one depends on your task and how sensitive your data is, which is why the tools segment and Part B get concrete. One principle that holds no matter which model you use: its outputs need the same two things from you, verification and disclosure.
What about top-p, seeds, the other knobs?
Top-p is a sibling of temperature: another way to reshape the distribution before the roll. A seed is different: it pins the dice, so the same prompt and settings give you the same draw, which is what you want when a result has to reproduce. Like temperature, all of these surface in the API and stay hidden in consumer apps.
Will it take my job?
The line this bootcamp draws: routine work, the kind where you can check and redo the output, gets faster; the judgment that defines research stays yours. Framing the question, calling the outlier, answering the reviewer: those remain yours, and AI cannot bear responsibility for them.
Further reading
- Welch Labs on YouTube, the deep-dive visual series on how LLMs work.
- 3Blue1Brown on YouTube, neural networks and transformers as animations.
- Language models for biological research: a primer, Nature Methods, 2024.
- The bootcamp repo, all sessions and materials.
- This site's source, one container, runs anywhere.
Run it yourself: docker compose up, then open port 8000. The model, corpus, styles, and scripts all live in the image.
Before you go
Thank you for reading
If this page taught you the mechanism in a way that stuck, that's the smallest version of the thing I care about most: learning how AI can actually help you learn. In February 2027 I'm running a one-week retreat on exactly that, called Learn Anything. You bring something you want to learn, and a small group of curious people works alongside AI as the accelerator.
The timing is the point. What AI can help you learn is on a curve that is accelerating away from us today, and getting on that curve is a skill worth practicing in person, together. If that sounds like your kind of week, I'd love to have you: learn-anything.nonlinearlabs.ai.