next‑token AI for Bioinformatics Bootcamp
Session 1 · Part A

Session 1 · Part A · intuition, math-light

How language models actually work

I'll spend this session on one trick: predicting the next token, over a library of text bigger than any lifetime of reading. By the end you can explain the trick to your own lab, and you know where it breaks.

The toy model behind every demo you'll see is a real language model, just a very small one. It read four short paragraphs of biology text and counted which letters follow which. Nothing here is faked, which is exactly the point.

Section 01

A language model is a very well-read autocomplete

Everything you'll see this fall, from code that writes itself to summaries of a thousand papers, rests on one core mechanism. The model looks at the text so far and predicts what comes next. It reads and writes one chunk at a time, over and over.

That sounds too small to matter. The surprise of the last few years is what the trick looks like when you do it extremely well, over more text than any person could read in a hundred lifetimes. My goal is that you leave able to trace any answer a model gives you back to this one mechanism.

How would you finish this sentence? Call it out, then let the model take its turn; we'll see what four short paragraphs of reading gets it.

Try it. The model below read four short paragraphs of biology text and counted which letters follow which. Give it the start of a sentence, then let it take over.

The continuation will appear here.

It has no idea what PCR is. It still writes biology-shaped words, because four short paragraphs of counting were enough to teach it which letters tend to follow which. Hold onto the gap between what it knows and what it produces; hallucination is where that gap bites.

Section 02

Text becomes numbers

Before a model can predict anything, it has to read. Models read in chunks called tokens. A token is a common cluster of characters: everyday words get their own token, rarer words get split into pieces. By the time text reaches the model, it has become a list of numbers, one ID per token, and that list is the model's entire view of everything you typed.

Tokenization explains a family of quirks that otherwise look like stupidity. Ask a model how many letters are in deoxyribonucleic and it may miscount, because it never saw letters at all; it saw opaque ID numbers. Exact strings from the lab bench hit the same wall: the model sees token IDs, not characters, and one wrong letter changes what a reagent or a variant means. Watch what a production tokenizer does to vocabulary your lab uses daily.

Where do the boundaries land inside deoxyribonucleic? How many pieces? Guess before you tokenize it.

Type anything, then tokenize it. Watch where the boundaries land, especially on long bio words.

The tokens will appear here.

Every chip is one token, with its ID number underneath. The model's whole world is sequences of these IDs. Our toy model reads one letter at a time so you can watch it work; production models read tokens. It's the same trick with a different chunk size.

Section 03

A probability list, then a dice roll

At every step, the model does one thing: it assigns a probability to every possible next token. Prediction means picking from that list, and picking, called sampling, means rolling a weighted die. That die is why the same prompt can give you a different answer tonight than it gave this morning.

Temperature is the knob that reshapes the list before the roll. Low values make the model cautious and repetitive. High values make it adventurous and strange. When you use a chat app, the app hides this knob from you. When you call a model through its API, the way software talks to a model directly instead of through a chat window, you set it yourself, and it belongs in your methods section.

Roll it twice at the same temperature. Hands up if you expect the same pick both times.

Roll it yourself. The bars show the model's top candidates for the next character after "The", reshaped by the temperature you set. The highlighted one is what the die picked this time.

0.8

The probability list and the pick will appear here.

Roll it several times without touching the slider. The bars stay put and the pick moves. That, in miniature, is what "the model is probabilistic" means for the reproducibility paragraph of your next paper.

Section 04

Where the numbers come from

The model's probabilities come from training. During training, the model makes a guess at every position in the text, compares its guess with what the text actually says next, and nudges itself a little toward the truth. Billions of nudges later, the patterns of the library live in the model's parameters, the stored numbers training adjusts.

This is why training data matters so much: the model can only mimic the distribution of the text it was shown. Our toy corpus is four short paragraphs of biology. A frontier model's library is closer to the whole readable internet. The mechanism is identical, and so is the catch: whatever patterns sit in the library, flawed and biased and outdated alike, end up in the model. Using the trained model afterwards is called inference.

Call it before the drag: at fifteen percent of the library, will it write soup, English-shaped soup, or biology?

Drag the slider to decide how much text the model has read. Then look at two things: how surprised it is by text it has never seen, and what it writes when prompted.

1 of 6

The falling bar is training. The changing text is the payoff: soup, then English-shaped soup, then biology-shaped sentences with the facts slightly scrambled. Scale that same loop up by a factor of a billion, and you get the frontier models.

Section 05

It looks like magic because the library is bigger than a lifetime of reading

A language model mimics the training distribution of text, and that distribution is bigger than anything any of us could read in a lifetime. So when you ask about a niche proteomics method at midnight and a fluent answer comes back, the question feels out of the ordinary to you. For the model, the question is ordinary: millions of pages in the training text say similar things.

That's why the right mental picture is interpolation, not retrieval. The model itself never looks anything up. Every answer lands somewhere inside the cloud of training text, and the cloud is so vast that almost everything a scientist can think to ask already lives inside it. The exceptions live at the edge of that cloud, and the last section walks through them. (When a chat app searches the web mid-answer, that's a separate tool stapled onto the same predictor, and Part B covers where that line sits.)

Give me the most niche question in your field, the one you're sure nobody has written down. That's the one I'll ask.

Ask something that feels out of the ordinary and watch where it lands. Each dot is a piece of text: the small dense cluster is what one careful researcher reads in a career, the vast cloud is what the model read during training.

your career of reading the model's training distribution

Questions that feel novel to us still land inside the cloud; that's what in-distribution means. The cloud has an edge, though: training stopped on a date, and that edge is where the exceptions begin.

Section 06

Confidently wrong is the mechanism working normally

A hallucination is the model doing exactly what it was trained to do, on a question where the fluent continuation happens to be false. There's no lookup table to consult and no alarm that rings when a prediction strays past what the model actually knows. There's only the next most plausible token.

This is what makes the failure mode dangerous in research: fabricated references arrive in perfect formatting, with plausible author lists, real journals, and tidy page numbers. Two of the five references below are fabricated; the other three come straight from the bootcamp's reading list.

A reviewer sent you this paragraph. Call each reference real or fabricated on its card below, then check your answers.

Hepatic JNK signaling has become a central node in metabolic disease research. Loss of JNK1/2 in the liver reshapes the transcriptome through the PPARα–FGF21 axis , and chronic JNK blockade has been reported to protect against diet-induced steatosis through FGF21-independent routes . In inflammatory skin disease, single-cell profiling of vitiligo showed an expansion of effector CD8⁺ T cells, with CCR5 positioning regulatory T cells to restrain disease . More recently, a single-cell atlas of photosensitive skin described MMP9⁺ myeloid circuits organized by keratinocyte–fibroblast crosstalk , consistent with the spatial mapping of a keratinocyte→fibroblast→myeloid axis driving photosensitivity .

  1. Vernia S., et al. (2014). "The PPARα–FGF21 hormone axis contributes to metabolic regulation by the hepatic JNK signaling pathway." Cell Metabolism 20(3):512–525.

  2. Marquez-Herrera R., et al. (2019). "FGF21-independent hepatoprotective effects of chronic JNK inhibition in diet-induced steatosis." Cell Metabolism 29(4):912–925.

  3. Gellatly K.J., et al. (2021). Vitiligo single-cell study: effector T cell expansion and CCR5-dependent positioning of regulatory T cells. Science Translational Medicine.

  4. Okonkwo A. & Reyes T. (2022). "Single-cell atlas of photosensitive skin reveals MMP9⁺ myeloid circuits." Nature Immunology 23(11):1745–1758.

  5. Wang Y., et al. (2026). Spatial mapping of a keratinocyte→fibroblast→myeloid circuit driving photosensitivity. Nature Immunology.

From inside the model, every citation is equally fluent. Fluency is its native output, so fluency can never be your evidence. Checking each reference against the primary source is the habit that catches fakes; Part B turns that habit into concrete rules.

Section 07

What changes when everything gets bigger

Same recipe, more of everything: more parameters to hold patterns, more text to learn them from, more compute to push the training loop. For years the gains looked politely incremental: better grammar, fewer nonsense words. Then, somewhere past a few billion parameters, abilities showed up that nobody had explicitly programmed: following instructions in plain language, picking up a task from a couple of examples in the prompt, working through a problem step by step: what researchers call emergence. Researchers still argue how sharp those jumps really are; that they arrived with scale rather than with code is the part everyone agrees on.

Drag across ten orders of magnitude and watch what the same mechanism picks up along the way. The counts are public through GPT-3; newer labs keep their numbers quiet, so treat the right edge as "more, undisclosed".

1 thousand
  • 1940s–now Common letter patterns

    It counts which chunks follow which. The toy model on this page lives here, at a few thousand counted parameters.

  • 2013 Word meanings as geometry

    Words become long lists of numbers, and similar words get similar lists, so word math starts to work: king − man + woman lands near queen.

  • 2018 Fluent paragraphs

    GPT-1, 117 million parameters. The transformer recipe starts writing decent prose on almost any prompt.

  • 2019 Whole essays, wobbly facts

    GPT-2, 1.5 billion. It writes multi-paragraph text that reads clean, though the facts arrive loose.

  • 2020 Learns from examples in the prompt

    GPT-3, 175 billion. Show two examples, ask for a third. Nobody programmed this; researchers still debate the shape of the jump, but it arrived with scale.

  • 2022 Follows plain-language instructions

    Chat-era models are tuned to do what you ask instead of completing text. The apps in your browser tabs live here.

  • frontier Reasons step by step, drives tools

    Sizes stay undisclosed, but multi-step reasoning and tool use show up at this end of the slider.

One thing scale didn't buy: the dice and the library's edge survive every order of magnitude, so fluent and true stay different words at any slider position. Emergence at scale is the bridge to foundation models, and the reason your field builds one shared model instead of training a fresh network for every question.

Section 08

Train once, use everywhere: foundation models

You now own every piece of the vocabulary: tokens, next-token prediction, training, sampling, distribution, scale. A foundation model is what you get when you push that recipe far enough on broad data that one model adapts to many jobs. Pretrain once on the giant library, then adapt: ask in plain language, show a few examples, or fine-tune on your own data.

The same idea left text behind. Swap what counts as a token and the same sequence-modeling machinery learns biology: amino acids, genes, regulatory DNA. Open the cards that matter to your science.

Hands up: proteins, genes, regulatory DNA, or single cells? Your votes pick which cards I open.

Open a card to see what goes in and what comes out; the block in the middle is the same in every card, which is the whole point.

Text LLMsGPT, Claude, Llama IN OUT your prompt as ··· sequence model same machinery next token chopped into tokens sampled, over and over

Sections 1 through 7 are this card. Everything else here is the same recipe pointed at a different kind of sequence.

AlphaFold 2protein structure, 2021 IN OUT M V L S G ··· sequence model same machinery amino-acid sequence, with related sequences the folded 3-D structure

The model treats a protein's sequence as a language and reads the structure out of the patterns across related sequences.

Geneformernetwork biology, 2023 IN OUT 1 · GENE-A 2 · GENE-C 3 · GENE-B sequence model same machinery genes, ranked by expression which genes belong together

Tokens are genes; predicting missing genes becomes predicting which genes belong together in a working cell.

scGPTsingle-cell analysis, 2023 IN OUT sequence model same machinery A B A C A B C A expression profiles, one per cell annotated, integrated cells

One pretrained model handles many single-cell chores. The Session 5 and 6 analyses run on models like this one.

Enformergene regulation, 2021 IN OUT A T G C A T G G C T A sequence model same machinery one long stretch of regulatory DNA expression of nearby genes, per cell type

It reads a stretch of regulatory DNA, the genome's control panel, and predicts how the nearby genes respond in each cell type.

The session Manuel Garber leads takes these tools into spatial biology. When you get there, it assumes the vocabulary you've built today, and you already have it.

Section 09

Know where the trick breaks, then use it anyway

Every failure mode you'll hear about this fall traces back to something you just worked through. Once you own the mechanism, failures become predictable, and each one comes with a habit that catches it.

Failure modes and the habits that catch them
What you'll seeWhere it comes fromThe habit that catches it
The model makes something up Prediction without lookup Verify claims against primary sources
Same prompt, different answer Sampling, the dice roll Log model, settings, and prompts; expect wobble
Confident about a world that changed last month The library has an edge date, and nothing you type adds to it Check current sources for anything time-sensitive
Repeats the field's old mistakes Mimics the training distribution Read outputs with the literature's biases in mind
Echoes whatever sensitive text you pasted Its entire view of you is the text you hand it; nothing in the mechanism keeps a secret Keep patient and unpublished data out of consumer tools

Before you go, I'll show you four failure reports, and you pick the idea that explains each one.

1. You ask the same question twice and get two different protocols.

2. Your paragraph cites a paper that turns out not to exist.

3. You ask about a preprint posted yesterday and get a summary that's fluently wrong.

4. Your analysis suggestion uses a statistical method the field moved away from decades ago.

The frame I find most useful: treat the model as a collaborator. Collaborators do brilliant work, make honest mistakes, and expect their work to be checked. The mistake you catch is the one you looked at.

The short version

A language model predicts the next token over a library bigger than anyone could read in a lifetime, so fluent answers come from interpolating within that library's patterns. They land outside the truth when the most plausible continuation happens to be false, because nothing in the mechanism checks truth.

Part B, led by Alper, turns the habits in that table into concrete rules: protecting patient data, verification discipline, disclosure, reproducibility, and the cases where you leave the model out of it entirely. He has war stories, and they're worth staying for.

Appendix

Take it with you

Every session is recorded, so nothing here depends on having been in the room. Here is the vocabulary in one place, the questions we expect in the room, and where to go deeper if you want more.

The vocabulary

Token
The chunk of text a model reads: a word, a piece of a word, or a punctuation mark, mapped to an ID number.
Next-token prediction
The whole trick: look at the text so far, assign a probability to every possible next token.
Probability distribution
The full list of probabilities the model assigns, one for each possible next token. The model's answer to every question is one of these lists.
Sampling
Rolling the weighted die: picking one token from the distribution. This is why answers vary.
Temperature
The knob that reshapes the distribution before the roll. Low is cautious, high is adventurous. Set explicitly in API use, hidden in consumer apps.
Parameters
The stored numbers inside the model; training adjusts them and every answer runs through them, so scale just means more of them.
Training
The expensive loop: guess the next token everywhere in the library, compare with reality, nudge the parameters.
Inference
Using the trained model to answer. Cheap compared with training; this is what happens when you chat.
Training distribution
The shape of everything the model read. The model mimics it, so it can only be as current and fair as its library.
Interpolation
Landing inside the cloud of patterns the model has seen. The right default picture for how it answers; it retrieves nothing.
Hallucination
The same mechanism, fluently wrong. A plausible continuation that happens to be false.
Fine-tuning
Extra training on your own data to specialize a pretrained model.
Foundation model
One model, trained on broad data at scale, adapted to many jobs: text, protein structure, gene expression, regulation.

Questions we expect in the room

Is it just autocomplete?

The mechanism is next-token prediction; the interest comes from doing it superbly over a library the size of the readable internet. Be careful with "just" in both directions: it deflates the hype and undersells the risk.

Does the model learn from our conversation?

The parameters freeze after training. Within one chat, the model rereads the whole conversation each turn, which feels like learning but is really a longer prefix. Whether the vendor trains on what you type is a separate data-policy question; the short rule is to assume your prompts are retained unless you're on a plan or deployment that says otherwise, and Part B walks through how to check for the tools your lab actually uses.

Which model should my lab use?

There's no single right answer; the honest one depends on your task and how sensitive your data is, which is why the tools segment and Part B get concrete. One principle that holds no matter which model you use: its outputs need the same two things from you, verification and disclosure.

What about top-p, seeds, the other knobs?

Top-p is a sibling of temperature: another way to reshape the distribution before the roll. A seed is different: it pins the dice, so the same prompt and settings give you the same draw, which is what you want when a result has to reproduce. Like temperature, all of these surface in the API and stay hidden in consumer apps.

Will it take my job?

The line this bootcamp draws: routine work, the kind where you can check and redo the output, gets faster; the judgment that defines research stays yours. Framing the question, calling the outlier, answering the reviewer: those remain yours, and AI cannot bear responsibility for them.

Further reading

Run it yourself: docker compose up, then open port 8000. The model, corpus, styles, and scripts all live in the image.

Before you go

Thank you for reading

If this page taught you the mechanism in a way that stuck, that's the smallest version of the thing I care about most: learning how AI can actually help you learn. In February 2027 I'm running a one-week retreat on exactly that, called Learn Anything. You bring something you want to learn, and a small group of curious people works alongside AI as the accelerator.

The timing is the point. What AI can help you learn is on a curve that is accelerating away from us today, and getting on that curve is a skill worth practicing in person, together. If that sounds like your kind of week, I'd love to have you: learn-anything.nonlinearlabs.ai.