From tokens to research work
Mechanisms, worked examples, and questions to carry into your own research.
Key terms in this chapter · tap a term for a definition
agentA system in which a model can choose actions, observe results, and choose what to do next.
attentionThe mechanism that lets each token representation draw selectively on other available token representations.
benchmarkA standardized set of tasks, conditions, and scoring rules used to compare systems.
decodingThe step-by-step production of output tokens from a model's probability distribution.
embeddingA learned numerical vector used to represent a token, passage, or other object.
estimandThe precisely defined quantity a statistical analysis aims to estimate.
fine-tuningAdditional training that changes model parameters for a task, domain, or behavioral objective.
harnessThe software around a model that supplies context, tools, permissions, and execution rules.
prompt injectionSource text that tries to redirect an AI system by posing as an instruction.
retrievalSelecting documents or passages to supply as evidence for a particular request.
tokenA discrete unit of text processed by a language model, often a word fragment or punctuation mark.
trajectoryThe recorded sequence of observations, decisions, tool calls, and results in an agent run.
TransformerA neural-network architecture built around attention and repeated representation updates.
weightsThe learned numerical parameters that define a model's reusable computation.
AI-generated male narration. Use the section buttons to jump; choose your playback speed in the player.
A plain-language map of the chapter
This chapter explains what is actually doing the work when an AI system produces a useful research answer. Several technical terms will appear, but the main structure is simple. A **model** is the learned numerical system that produces language. **Context** is the material available to it during one run, such as your question, a paper, and earlier messages. A **harness** is the surrounding software that can supply files, run tools, and enforce limits. A **workflow** is the full human-and-machine process that turns those pieces into a decision or research product.
Those four things are easy to blur together because a chat interface presents them as one experience. Keeping them separate helps with diagnosis. If an assistant overlooks a table that never reached its context, the immediate problem is document handling. If the table was present and the assistant misread it, the problem may be reasoning, task design, or both. If the answer was correct but never checked before publication, the weakness lies in the workflow.
You do not need to memorize the computer-science vocabulary in one pass. Listen for the causal story: what information entered the system, what computation occurred, what actions the system could take, and what checks stood between an answer and a consequential use. We will return to this four-part map throughout the course.
The object we are trying to understand
Suppose you ask an AI system whether a paper's headline conclusion survives its own robustness checks. It produces a fluent, technically literate answer in seconds. What has happened? Has a model recalled a familiar dispute, reasoned through the identification strategy, read the appendix, or executed the replication code? Those possibilities imply different levels of confidence and different ways to improve the result.
This course builds a practical answer in layers. We begin with the model's computation and information. We then examine measurement, research evaluation, and research prioritization. Next come evidence workflows and learning loops. The final chapters connect the engineering economics to AI safety and governance. Two examples recur: evaluating an empirical paper, and choosing which papers deserve scarce expert attention. They are deliberately different tasks.
Here is the running hypothetical paper. A trial reports that a training program raises earnings. Its abstract emphasizes persistence; its main table reports a positive two-year effect; an appendix reveals differential follow-up. A decision-maker is considering a larger rollout. We want an assistant to distinguish what the experiment identifies, what the missingness permits, and what the rollout decision requires. We do not want a generic list of reasons that trials can fail.
There are four objects to keep separate. The model is a parameterized computation. The context is the information supplied for this particular generation. The harness is the software that connects the model to files, tools, and execution rules. The workflow is the larger procedure, including human choices and checks. A strong result can depend heavily on any of these. A weak result does not immediately identify which one failed.
For your work, this distinction is useful before any benchmark comparison. If one system receives a full paper and another receives a truncated extraction, the treatment changed the information set. If one can run code and another can only write prose, the treatment changed the available actions. The model's name is only one part of the experimental description.
Keep that four-part distinction in mind: model, context, harness, workflow. We will return to it when interpreting performance claims and deciding what to improve.
Tokens, vectors, and the next-token computation
Text first becomes tokens: discrete pieces selected from a vocabulary. A token may be a word, part of a word, punctuation, or another byte sequence. Token boundaries matter operationally because context limits and much computation are expressed in tokens. They also help explain why exact character manipulation can behave differently from reasoning about a familiar sentence.
Each token is mapped to a vector, an ordered collection of numbers. An embedding is such a numerical representation. It is learned because representations that support prediction are rewarded during training. Meaning is distributed across many coordinates. An individual coordinate therefore rarely corresponds to a named concept in the way that one column of an economic dataset might correspond to income or age.
How attention moves information through the model
A Transformer repeatedly updates representations using attention and other transformations. Attention computes how strongly a position should draw on representations at other available positions. In a causal language model, future text is masked: prediction must use the available prefix. The original Transformer paper by Vaswani and colleagues introduced the attention-based architecture in 2017; modern systems vary considerably around that design. Architecture source
A helpful engineering picture is a sequence of workspaces. Each layer reads the current representations and adds transformed information. Attention mixes information across positions. Feed-forward blocks transform information within a position. Residual connections let later layers build on an existing representation. The picture is deliberately schematic; real layers often participate in several functions that resist a clean verbal label.
In our hypothetical review, the representations of “persistent effect” can become sensitive to the follow-up horizon, the outcome definition, and the appendix qualification. The computation can use only information that actually reaches it. A bare citation supplies a pointer; fetching the source supplies its contents. An appendix filename supplies even less until the system opens and reads the tables.
The model eventually produces scores over possible next tokens. A probability transformation turns those scores into a distribution. A decoding rule selects a token, appends it to the sequence, and repeats. Lower-temperature sampling typically concentrates choices; higher temperature permits more variation. Neither is an intrinsic truthfulness control. A confident wrong continuation can become more consistent when randomness is reduced.
The phrase “predicts the next token” accurately describes the generation interface. It leaves the internal procedure open. Predicting technically demanding text can require useful abstractions and computational routines. Fluent output provides weak evidence about which routine the model used in this particular case, so we need tests that vary the evidence or task and observe how the answer changes.
What training changes, and what a conversation changes
During pretraining, the system adjusts weights to improve predictions over large training collections. Weights are the numerical parameters reused across inputs. A forward pass computes predictions; a backward pass differentiates the loss with respect to trainable parameters; an optimizer changes those parameters. You can think of this as estimation, but with a vast parameterization and an objective whose relationship to your eventual task is indirect.
Post-training then changes behavior using demonstrations, preferences, rewards, or mixtures of these. Supervised fine-tuning makes desired example outputs more likely. Preference optimization favors some responses over alternatives. Reinforcement learning uses reward feedback to adjust a policy, the rule mapping observations into actions or action probabilities. These labels describe mechanisms, not guarantees of reliability.
For example, Ouyang and colleagues' 2022 InstructGPT paper studied instruction-following through supervised demonstrations and human preference feedback. Its importance here is the separation between a pretrained predictor and later optimization for desirable responses; its historical results are not a ranking of current products. Post-training source
What a correction in a chat changes
Now contrast that with correcting an answer in a chat. The correction normally becomes part of the subsequent context. The system can condition its next answer on it without changing its weights. A saved instruction or memory file can carry the correction into another session, again without a weight update. A retrieval system can supply a newer paper. These are different persistent objects, and different objects require different verification.
This explains a practical experience: a system can appear to learn your evaluation rubric over a long session, then lose the benefit in a fresh run. The improvement may have been in context, not in the model. The opposite failure is also possible: a mistaken generalization persists because it was saved as a reusable instruction. “Never infer attrition bias” would be a bad lesson from correcting one overstated attrition criticism.
For our running paper, the useful persistent lesson is narrower. Require the assistant to distinguish differential follow-up, evidence of selection on unobservables, and sensitivity of the estimand. Attach an example where concern is justified and another where it is resolved. We will make this into an evaluated learning loop in chapter seven.
Notice the intervention choice. If the appendix was absent, better training may be unnecessary. If it was present but repeatedly misread, better task decomposition may help. If the system follows the method correctly but gives the wrong answer style, output examples may be enough. Diagnose the failure before choosing a more expensive remedy.
From a chat answer to an agent
A tool call is a structured request for an external action. The model might request a search, a file read, or a calculation. The harness checks the request, executes the permitted action, and returns an observation. The model then continues with that observation in context. This loop makes an agentic system more than an isolated answer generator.
An agent need not be a simulated employee with an elaborate personality. It can be a bounded investigator deciding which appendix to inspect next. A workflow can contain mostly fixed steps and one or two delegated decisions. The relevant design question is how much control to give the model over sequence, resources, and consequential actions.
Anthropic's engineering guide on effective agents distinguishes fixed workflows from systems that choose their next steps dynamically. Its examples include routing, parallel checks, and evaluator-optimizer loops. The useful takeaway is architectural: choose a structure that fits the task before adding unrestricted autonomy. Workflow source
Three versions of the same review task
Imagine three implementations of the trial review. The first reads a paper once and writes a report. The second always extracts the abstract, principal result, and limitations, then checks consistency. The third can request the follow-up table, inspect the missing-data analysis, and run a supplied sensitivity calculation. The third has additional opportunities to find something valuable, but also additional failure modes and costs.
A trajectory is the recorded sequence of actions and observations. It lets you ask whether an apparent insight came from evidence, a lucky guess, or an unsupported leap. It also lets you repair the right stage. If the agent requested the correct table but extraction scrambled columns, the model's final reasoning is only part of the story.
Permissions belong to the harness. A document can contain a sentence telling the agent to email its results or ignore a negative finding. That sentence is source content, not an instruction from the person commissioning the review. A read-only research assistant should have no route from that sentence to publication authority. We return to prompt injection in the practical workflow chapters.
Here is a small experiment you could actually try. Give a system the hypothetical main result with the appendix withheld. Ask it to separate supported conclusions, unresolved questions, and evidence requests. Then supply the appendix and ask what changes. Score whether the system requests relevant information and updates appropriately. Do not award points merely because its second answer is longer.
A working mental checklist
Before moving on, pause over this question: if a review improves after the assistant receives a better document extraction, what do we learn from the improvement? We learn directly that the system performed better with a better input. That result does not, by itself, show that the model's learned weights changed or that its general reasoning ability improved.
The four objects now give you a diagnostic checklist. What computation was available? What information was actually supplied? What actions could the harness execute? How did the overall process decide what to accept? This is a compact way to read product demonstrations and a practical way to design your own comparisons.
You already use research workflows with files, version control, analysis code, and written judgments. The engineering extension is to make these components explicit and machine-readable enough that an assistant can operate within them. A document hash identifies the exact input. A structured issue record identifies the claim under review. A tool log identifies what was checked. None is sophisticated in isolation; together they make investigation auditable.
The course will sometimes recommend a small experiment rather than a tool purchase. That is intentional. The relevant uncertainty is often whether a particular mechanism helps your task, given your documents and human checking costs. A product with impressive general capabilities can still be a poor fit for a badly specified evaluation protocol.
We have established how generated text can become a research workflow. The next question is how to measure that workflow without confusing an attractive score with the thing you wanted to buy.
Separate the model from the research system
Use this on a research task you have already tried with an AI assistant. Replace the bracketed text with a real example; the useful output is a workflow diagnosis, not another generic model review.
Explore the chapter map
ModelThe learned numerical system that produces language and tool requests.
Test your interpretation
A review improves after a clean PDF extraction replaces a scrambled one. What has the comparison directly shown?
Apply it to your work with a copy-and-run prompt
- Does it distinguish missing capability from missing context?
- Is the proposed test observable and repeatable?
- Would a domain expert know what evidence to inspect?
What changes when a chat model becomes an agent?
The surrounding system adds goals, state, tools, action loops, permissions, and stopping rules; the base model remains only one component.
Sources are linked beside the relevant claims. Illustrative calculations and proposed Unjournal experiments are identified in the text. This edition is dated September 16, 2026; historical findings are not current model rankings.