Which World Model?

Representation, Reliability, and the Uses of Prediction

The previous essay, “Which AGI?”, distinguished a running model, a deployment built around it, and the lineage sustained by a laboratory and its users. Much of the disagreement about general intelligence turned out to concern which of these systems should possess the capacities being demanded. Continual learning made the distinction particularly visible: a laboratory can improve its models without the individual deployment learning from its own life.

There is a parallel disagreement about world modeling. To one observer, an AI’s ability to diagnose a software failure, follow a complicated argument, or work out what changes when a premise changes is compelling evidence that it models the systems it discusses. Critics point to failures on spatial or causal tasks as evidence that LLM-derived systems lack world models and that the research programme is heading in the wrong direction. Again, the subject matters: a text-only model, a multimodal predictor, and a robot controlled by one are different systems. So does the timescale: a shortfall in one generation says little about the direction in which successive generations are moving.

Our starting intuition is the first observer’s. Competence needs an explanation. A system that repeatedly succeeds at unfamiliar tasks, under variations that defeat memorized answers, must have acquired something about the relations that make those tasks work. Calling this “pattern matching” leaves the explanatory question untouched: which patterns, represented how, and combined by what operations? A learned pattern can be a causal dependency, a program invariant, or the structure of an argument. The word does not settle the sophistication of the mechanism.

Competence is evidence, but it does not uniquely identify a mechanism. Successful behavior can arise from a policy that does not simulate consequences, from a shortcut that works in the tested domain, or from an external tool. The efficacy of a whole deployment does not locate its intelligence inside one component.

Diverse, transferable competence is strong evidence for sophisticated representations. Targeted experiments can establish that particular representations model environmental structure and causally guide behavior. A coherent, continually maintained model that an agent can consult reliably for prediction and action is a further achievement. A failure to maintain and reliably use a world model can then be mistaken for an absence of world modeling altogether.

The negative claims need similar discipline. “This system falls short” reports a level of performance. “This approach is a dead end” makes a claim about a trajectory. To support the latter empirically, we need a relevant metric, its movement across successive systems or increasing research resources, and an expectation against which to judge that movement. A persistent gap below a demanding threshold can coexist with substantial progress toward it. Whether that progress is fast and affordable enough is a question about the trend.

In “Which AGI?” we placed a structured world model among the roles filled inside the network. Here we owe that entry a defense. We also need to say what filling the role leaves unresolved.

Richard Carrier gives the objection an unusually forceful form. In the passage from his continuously updated “AI Is Garbage and a Bubble”, he describes a progression through five achievements. A machine must model its surroundings and body well enough to act; invent imaginary environments in which to develop skills; model causal systems, including imagined ones; model its own reasoning and errors; and integrate these capacities with a model of its whole mind, a narrative history, and an enduring but revisable organization of desires. His conclusion is categorical:

That model-building and model-navigating pathway is the only way to real AI.

And of systems descended from language modeling:

And why LLM-based AI can’t and never will.

This is a useful challenge because its positive proposal has content. It asks for an agent that can locate itself in a situation, anticipate consequences, learn through imagined and actual action, and revise its conduct in light of what happens. We share much of that aspiration. The problem is the jump from demanding this package for “real AI” to denying that present systems think, learn, or understand anything—and to treating one proposed developmental route as an impossibility theorem about other routes.

Carrier’s list moves between several standards. Knowing the structure of an environment is one achievement. Using that knowledge to control a body is another. Retaining discoveries across a lifetime is a third. Having a narrative self whose projects organize that lifetime is a fourth. An absence at the fourth level cannot establish an absence at the first. Otherwise an animal’s uncertain status as an autobiographical self would cast doubt on whether it represents its surroundings at all.

Nor does the order in which capacities developed biologically establish the order in which they must be engineered. Even if Carrier’s proposed sequence were the correct account of our evolution, it would demonstrate a successful route. Establishing uniqueness would require an additional argument. A machine trained on the products of human culture begins with access to information that the first organisms possessing these capacities had to acquire by other means. That inheritance may permit a different developmental order.

A particularly important possibility is that an agent acquires extensive knowledge of causal systems before it acquires a robust model of its own embodiment. A language model can encounter descriptions of experiments, interventions, failed plans, mathematical structures, and other people’s experiences. Those descriptions are imperfect evidence, but they carry information about the world. We should ask what can be recovered from them, where that recovery fails, and what further interaction repairs the failure. Declaring the channel disqualified in advance answers none of those questions.

The empirical evidence Carrier invokes includes a narrower comparison across model sizes. In Learning the Wrong Lessons, Chantal Shaib and colleagues fine-tune OLMo-2 variants on a synthetic dataset. Spurious associations between syntax and subject domain persist in both base and instruction-tuned variants across the tested 1B–13B sizes. Yet in-domain paraphrase accuracy improves from 53 to 84 percent for the instruction-tuned models. The negative result concerns transfer across domains under that recipe. It is evidence against expecting those changes alone to remove this shortcut. The study excludes reasoning models and chain-of-thought-trained models; it does not measure the trajectory of successive frontier systems. A local scaling result can identify a bottleneck without establishing that an entire research programme is pointed away from the goal.

We therefore need a meaning of world model that is demanding enough to exclude arbitrary successful behavior, while modest enough to describe partial achievements.

Consider a representation of a situation: which objects are present, how they relate, what can happen next. It earns the name model when its internal organization supports inferences about the represented system. It earns explanatory significance when the agent actually uses those inferences. An intersection represented somewhere in a neural network is more interesting if changing that representation changes which turns the network takes to be possible. A representation of a program’s state is more interesting if it supports predictions about the effects of a change to the program.

In control and reinforcement learning, a world model often has a more specific job. Given an estimated state and a candidate action, it predicts a subsequent state or observation. Schematically,

\hat{s}_{t+1}=F(\hat{s}_t,a_t).

Here the state may be a learned vector, and the prediction may be a distribution over possibilities. In a partially observed environment, the relevant state may summarize uncertainty about what is happening. This is already more than possessing useful concepts: it supplies transitions that can be used to evaluate actions before taking them. Yet it still leaves open how the agent chooses which actions to evaluate, judges their desirability, or discovers that its model is wrong.

The distinctions can be made operational:

Capacity What would count as evidence?
Representing structure Internal variables encode relevant entities and relations, including on held-out cases.
Tracking the current situation Those variables update appropriately as observations and actions arrive.
Predicting consequences The system forecasts transitions under specified actions, including relevant alternatives.
Using a model Altering or disabling the representation changes decisions in the predicted way.
Maintaining a model Experience leads to durable corrections that improve later performance.

These capacities can vary independently. A system may know the map and lose its position. It may predict the result of an action but choose a bad action. It may correct itself within a conversation and forget the correction afterward. Its internal representations may be causally active without being available through ordinary verbal questioning. None of those distinctions is captured by asking whether it has a world model as though the answer were a single bit.

The familiar description “next-token predictor” does not resolve the question either. It identifies an objective and an output interface. It does not enumerate the computations learned to perform the objective.

Suppose a training corpus contains sequences generated by some structured process: moves in a game, operations on a machine, or journeys through a city. The next observation depends on a state that the sequence does not always name explicitly. A successful predictor can benefit from inferring that state and its transition rules. It can also exploit shortcuts. The objective makes both routes possible; the data, architecture, optimization, and available resources influence which route is learned.

For language, the generative process is more complicated because people write for many purposes. They describe reality, deceive, invent, speculate, quote others, and make mistakes. Predicting their text can reward representations of the world and of speakers’ beliefs about it. It can also reward reproducing a persuasive falsehood. Learning from language therefore provides neither a prohibition on world modeling nor a guarantee of truthfulness. The mixture of purposes in the data is part of the problem.

Token prediction also operates through continuous internal representations. The choice between a discrete output vocabulary and a continuous latent space is not a choice between a machine containing only words and one containing ideas. A transformer producing words already computes with vectors. The substantive question is what those vectors preserve and what the training objective encourages them to do.

This matters to the argument from efficacy. If successful performance survives changes to names, surface descriptions, irrelevant details, and combinations of familiar components, an explanation based on memorized wording becomes less plausible. If predictions about a hidden state transfer between tasks, the case strengthens further. If interventions on the putative representation produce the expected behavioral changes, we have begun to identify the mechanism. Evidence accumulates through these increasingly discriminating tests.

There remains a real causal limitation. Observations alone can leave different causal explanations indistinguishable. A predictor can fit the observed distribution while giving the wrong answer about an intervention. But the limitation concerns the information and assumptions available to the learner. It does not establish that a transformer cannot represent a causal explanation, learn from descriptions of interventions, or run an experiment through its tools. An action-conditioned predictor is not automatically a correct causal model either: the data may fail to distinguish the effect of an action from the circumstances in which it was chosen.

The proper question is whether the system has enough evidence to identify the relevant dependency, and whether its internal computation preserves and uses that dependency. This question applies to every proposed architecture.

An unusually revealing example comes from Pierre Beckmann, Matthieu Queloz, and André Freitas’s September 2026 preprint, World Modeling in Transformers. It revisits a transformer trained to predict turns on routes through Manhattan. Earlier behavioral tests had produced an apparently devastating result: despite good next-turn predictions, maps reconstructed from its generated routes contained incoherent connections. Forced detours also damaged performance. The natural interpretation was that the model had failed to learn a faithful map.

The new investigation separates the map from the mechanisms that use it. In the model trained on random walks, the researchers identify representations of intersections, legal moves, and connections between intersections. Their readout recovers the exact set of legal moves for 99.4 percent of intersection features at the reported layer and threshold. Connection recovery is less complete: in a steering test, the correct successor is the most strongly activated intersection for 76.6 percent of tested streets, and among the top five for 93.3 percent. This is evidence of substantial relational structure, with measurable imperfections.

Crucially, the study goes beyond decoding. An external probe can sometimes read information that the model itself does not use; finding a readable feature is therefore insufficient. Here, changing the encoded position changes the model’s choices appropriately. The authors also identify a separate directional representation, a goal compass. Steering it changes the direction of travel. Ablating it largely preserves legal moves while impairing progress toward distant goals. Knowing where movement is possible and knowing which direction advances the journey are supported by distinguishable mechanisms.

Why, then, the incoherent routes? The intersection features share representational space. Under the tested distribution shifts, the signal for the current location weakens and interference from other locations becomes more consequential. The model can apply information belonging to the wrong intersection. Researchers can substantially improve the relevant behavior by restoring the correct position signal or reducing interference.

The distinction is familiar outside AI. A person with a good street map can take an impossible route if they keep locating themselves on the wrong street. Observing the route would reveal a navigation failure; it would not tell us whether the map was wrong. Here the distinction has an internal causal account, rather than relying on an analogy with people.

There is a further complication. Intersections admitting the same moves tend to have similar representations. Confusing two such intersections can leave the next move legal, giving the system an opportunity to recover. What looks like a shortcut—grouping situations by immediate affordances—can coexist with finer structural knowledge and help protect behavior against representational interference. The choice between shortcuts and world models is not always exclusive.

These experiments use a finite world with deterministic transitions, and some interventions supply the experimenter’s knowledge of the correct state. Those interventions diagnose the failure; an autonomous agent would need its own way to recover that state information.

The study undermines a categorical inference from unreliable behavior to absent modeling. The model’s failure was real. So was its map. The important discovery was how the two coexisted.

That discovery gives training research a more precise target: making learned world structure easier to maintain and use.

Jayden Teoh and colleagues’ Next-Latent Prediction Transformers Learn Compact World Models addresses this question by adding supervision on internal states. Ordinary next-token training lets a transformer consult many earlier positions whenever it needs information. The current hidden vector need not contain everything required to continue the computation: attention can look backward again at the next step. This flexibility is useful, but it does not by itself impose a compact state with a consistent update rule.

NextLat jointly trains a small dynamics model to predict the transformer’s next hidden state from its current hidden state and the next token. The next-token loss remains. The additional objective encourages information needed later to be carried forward in a form that the dynamics model can update. The transformer backbone and its ordinary inference procedure can remain unchanged.

The paper proves that if the model achieves both the correct next-token distribution and the specified consistency of latent transitions, its hidden state is sufficient for predicting the future sequence: a belief state in the paper’s sense. The result establishes predictive sufficiency under exact conditions. Whether training attains those conditions is an empirical question, and a causal interpretation requires additional evidence.

On the Manhattan detour test, the metric is the fraction of valid traversals for unfamiliar origin–destination pairs when random legal turns replace the model’s preferred turn 75 percent of the time. Next-token training scores 85 percent and NextLat scores 95 percent. But multi-token prediction also reaches 95 percent: on this measure, predicting additional tokens supplies the same improvement as predicting latents.

The more distinctive gains appear on synthetic planning problems. On the Path-Star task, which requires choosing a route through a graph, NextLat achieves nearly 100 percent accuracy in all three tested configurations, while the compared methods fail on at least some configurations. This is evidence that the training recipe can overcome particular learning difficulties encountered by the token-prediction baselines.

The permutation-tracking experiment tests a different contribution. The ordinary and NextLat-trained transformers fail to generalize from 12-token training sequences to 36-token sequences, while NextLat’s auxiliary model exceeds 95 percent accuracy when run recurrently. Recurrence supplies a computation that grows with sequence length. The interesting training result is that this recurrent model learns its update rule from transformer states produced in parallel. Without a separately trained recurrent baseline, the experiment does not isolate an advantage of latent prediction over conventional recurrent learning.

The language-modeling results are much less striking. At 1.3 billion parameters trained on 100 billion tokens, ordinary next-token training averages 58.82 percent on the multiple-choice benchmarks. NextLat scores 58.79 with one-step supervision and 59.21 with two-step supervision, with gains and losses across individual tasks. Ordinary next-token training remains a strong baseline even against a recipe explicitly designed to improve the organization of its hidden states.

NextLat’s strongest evidence concerns particular synthetic tasks. Its language results give little reason to regard ordinary language-model training as a mistaken direction. The useful question is where the added constraint earns its cost and whether those gains extend to broader tasks.

Joint Embedding Predictive Architectures, or JEPAs, make a related design choice: predict representations of targets rather than reproduce every detail of the targets themselves. A useful forecast for manipulating a cup may need its position, orientation, and relation to the hand, while leaving countless details of the image unspecified. Learning an appropriate abstraction can reduce the burden of prediction. But learning what to discard is a substantive problem. A small detail irrelevant to one task may determine the outcome of another.

Lukas Kuhn and colleagues’ LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives makes that problem concrete. It learns from paired images and captions: an image representation predicts a text representation, and a text representation predicts an image representation. Separate prediction heads perform these mappings; the targets are held fixed for each prediction’s gradient calculation, and a regularizer pushes each modality’s embedding distribution toward a standard Gaussian.

The measured gains concentrate on features for individual image regions. With the same vision backbone and training data, LeVLJEPA improves semantic segmentation over the contrastive image–text methods CLIP and SigLIP: on ADE20K, its mean intersection-over-union score is 23.15%, against 20.90% and 19.24% respectively. Its visual features also improve question-answering accuracy when connected to either of two frozen language models. Classification features summarizing the whole image are roughly tied, while zero-shot ImageNet classification at the larger training scale is worse: 42.45%, against 47.32% and 50.78%. The comparison concerns ways of training visual representations; it does not test a replacement for autoregressive language pretraining.

The ablations reveal why predicting abstractions is easier to propose than to make work. Directly matching the image and text representations almost completely collapses the visual representation. Adding the distributional regularizer spreads the representations out but still yields poor features: in this reduced experiment, ImageNet accuracy with a linear classifier is only 2.56%. Combining the regularizer with prediction heads and stop-gradient targets raises it to 27.51%. Avoiding a constant representation is not enough; the distinctions that survive must be useful. The paper’s achievement lies in a recipe that makes those distinctions learnable.

This difficulty matters to the claim that the LLM-derived trajectory is misguided. Next-token prediction supplies an externally fixed target: the learner cannot make its task trivial by changing the tokens it is supposed to predict. Jointly learned latent targets introduce another freedom that must be controlled. Comparing functioning language models with a latent predictor assumed to retain exactly the right information gives the proposed alternative credit for a central problem it has to solve. LeVLJEPA offers evidence of a particular solution and its tradeoffs. Its clearest practical advantage is a better visual component for the language-model-based systems in which the authors test it.

Alexander Mattick’s conversation with Tim Scarfe on Machine Learning Street Talk puts this in a useful perspective. Around 34 minutes, Mattick reports that DINO embeddings have worked much better than JEPA-style embeddings in his own applications. He immediately complicates the comparison: the JEPA label is broad enough that DINO could itself be presented under it. Around 2:01:22, he returns to the substantive difficulty: an internal state can be consistent yet useless, and its usefulness can be hard to evaluate.

Mattick also endorses a concrete development in this family. Around 37:38–39:58, he agrees with LeCun that the contrastive instruction to pull similar examples together and push others apart insufficiently constrains a useful representation. He praises LeJEPA’s statistical tests for matching embeddings to a fixed reference distribution, which he sees as improving stability over a target determined by the current batch. This is the kind of distributional regularization used in LeVLJEPA.

LeVLJEPA uses a vision transformer and a GPT-2-style text encoder. It changes the learning objective within familiar computational machinery. In the downstream tests, the resulting visual encoder feeds an existing language model: research motivated by a criticism of current training becomes part of an LLM-derived system.

The scope of the proposed dead end matters here. Scaling a fixed next-token training recipe is a narrower programme than developing systems descended from language models. A successful change of objective might count against the sufficiency of the fixed recipe while counting as progress within the broader programme. The boundary should be stated before the result is known. Otherwise every improvement can be reassigned to a different path, making the original path’s alleged sterility impossible to test.

Emmanuel Dupoux, Yann LeCun, and Jitendra Malik address a more ambitious problem in Why AI Systems Don’t Learn and What to Do About It: Lessons on Autonomous Learning from Cognitive Science: how an agent could organize its own learning.

They distinguish three systems. System A learns from observation, acquiring representations and predictive structure. System B learns through action, pursuing goals and receiving feedback. System M controls their interaction: it selects inputs, routes information, changes operating modes, and regulates learning using signals such as uncertainty and error. An episodic memory provides material that can be revisited and replayed. The paper recognizes substantial existing achievements in the first two systems. The central problem is that human engineers still organize much of their cooperation.

A child can watch someone use a toy, try it, notice failure, watch again, experiment with a particular movement, and revisit what happened. In contemporary machine learning, analogous transitions are often divided among separately designed training stages. A team curates a dataset, trains a representation, supplies an action objective, evaluates the result, and changes the recipe. The proposal asks how an agent could perform the relevant selection and coordination within its own activity.

That is a close technical relative of Carrier’s aspiration. It connects observation to action, permits learning through communication and imagination, and gives internal monitoring a causal role. The authors present a tentative blueprint that can be explored in simple environments, without first building Carrier’s autobiographical self. Their proposed System M even has a core routing policy fixed over an individual’s lifetime, shaped on an outer evolutionary timescale, while the learning systems it regulates change within the lifetime.

This makes autonomy a matter of organization. A fixed rule can direct learning in response to changing conditions. Conversely, a collection of plastic components can remain dependent on an external operator for every significant adaptation. Whether the system learns autonomously turns on where control of learning resides, what it can change, and how effectively those changes improve later conduct.

The connection to “Which AGI?” is direct. At the lineage level, the laboratory already supplies data selection, evaluation, and changes to the training recipe. The proposal internalizes some of those functions within a deployment. A tool-using language agent that notices uncertainty, runs a test, updates a file, and consults that file later supplies parts of the same functional loop. Extending it to the A–B–M proposal would require control over learning in its perceptual and procedural components as well.

The earlier essay’s distinction between interpreted and compiled knowledge helps locate the remaining work. An interpreted record can preserve a new lesson while leaving the agent to work out its application each time. Retrieval, executable artifacts, and selective context can reduce that cost. The model already has implicit skills acquired during pretraining; integrating a new lesson into those skills could make it available without retrieving and interpreting the record. The question is how new experience becomes part of future performance, across which tasks and at what expense.

The two objections to current AI—insufficient world modeling and insufficient continual learning—meet here. A model that cannot durably revise itself may retain a sophisticated but increasingly unsuitable account of its environment. A learner with poor representations may fail to recognize what an experience should change. Better state representations and better control of learning can reinforce each other.

Mattick and Scarfe’s conversation also draws attention to the cost of using knowledge. Specifying an objective or possessing a representation leaves open how efficiently a system can obtain a useful answer.

In the discussion of energy-based models, this distinction concerns the work required to find or sample states favored by a learned energy. In the later discussion of control, it becomes the gap between predicting consequences and selecting actions. Mattick illustrates this with chess: give someone perfect knowledge of the game’s transition rules and they still have an enormous problem deciding how to play. A perfect one-step dynamics model leaves search, evaluation, and the allocation of computation to be solved. The relevant passages occur around 28–33 minutes and 2:04:38.

This is a limitation of the strongest world-model rhetoric too. A learned simulator can make experiments cheap, enable planning, and supply training data for a policy. It does not follow that control becomes easy, or that an agent knows which imagined experiments are worth conducting. A simulator can also be exploited by a planner that discovers actions appearing successful only because the simulator gets them wrong. Model quality has to be assessed over the situations that the resulting policy actually visits.

Mattick’s emphasis on prior structure adds another correction. Information can be expensive to acquire. A known physical limit or established constraint can spare a learner repeated exploration of possibilities already ruled out. Around 1:38:38, he argues for incorporating knowledge for which someone has already paid the price of discovery. Around 1:42–1:50, the discussion develops the advantages of stating constraints explicitly rather than hiding every requirement in a weighted reward.

This should make us take language pretraining more seriously as one possible source of inherited structure. A corpus contains the products of costly investigations, distilled into explanations, procedures, and warnings. A system that can use them has access to an enormous prior. The question is how much of that prior is accurate, how it is represented, and how new evidence can correct it. Our inference from Mattick’s argument is that discarding this inheritance merely because it arrived through language would itself be wasteful.

The robotics discussion brings these issues back to physical deployment. Beyond a successful demonstration, we need to know the distribution of failures, the quality of sensing, and how the controller behaves when its assumptions cease to hold. Mattick presses this point in the final part of the conversation. Explicit constraints can help, but a constraint enforced in an inaccurate model does not automatically become a guarantee about the physical world. Reliability requires the modeling assumptions to survive contact with the task.

With these distinctions in hand, we can assess the much weaker argument in Rekhi’s February 2026 Medium article, “Why Transformers Are Wrong for AGI and Why Scaling Them Higher Makes No Sense”. Its central errors are failures of inference. It repeatedly takes a limitation of a specified computation or training regime as a universal limitation of systems built using transformers.

The opening presents function-composition results as proof that transformers cannot reliably combine even two elementary relational facts. But the cited work by Binghui Peng, Srini Narayanan, and Christos Papadimitriou proves its central composition bound for a single attention layer, under an explicit relationship among domain size, number of heads, embedding dimension, and numerical precision. It distinguishes this theorem from suspicions about multilayer behavior and acknowledges that intermediate generated steps can mitigate the simple composition problem. The condition specifying the lower bound is part of the result. Removing it does not strengthen the theorem; it changes the claim. See On Limitations of the Transformer Architecture, especially Sections 1 and 3.

Another passage turns Sanford, Hsu, and Telgarsky’s triple-detection problem into three-digit modular addition. The actual task asks about matching triples drawn from a sequence. The paper establishes resource requirements for a single attention layer, gives positive results for other tasks and structured variants, and separates proved results from a conjecture about deeper networks. This is valuable evidence about which operations attention computes efficiently. It supplies no blanket impossibility of arithmetic or world modeling. See Representational Strengths and Limitations of Transformers, Section 1.1.

The broader complexity discussion similarly treats restrictions on a fixed computation as restrictions on everything an agent can do over time. Merrill and Sabharwal’s The Parallelism Tradeoff places the formal class of transformers they analyze in logspace-uniform TC⁰. That is an upper bound under specified assumptions, not an equivalence between deployed AI and unrestricted log-space computation. Their subsequent work on chain of thought explicitly studies how intermediate generation extends computational power, with results depending on the number of steps and normalization assumptions.

An argument about computational limits must therefore specify the computation: one pass, multiple generated steps, recurrent latent updates, or a controller invoking external tools. A fixed-size network asked to answer immediately and a deployment allowed to run an algorithm are different computational systems. “Which AGI?” has a mathematical application here.

The article also infers an absence of reasoning from the next-token objective. The inference fails because a training objective does not prescribe all intermediate computations. TaxiGPT provides a concrete counterexample to the corresponding claim about environmental representation. LeVLJEPA and NextLat show how changing objectives can further shape those computations within transformer-based systems. An alternative objective can be better without its predecessor having learned nothing of the relevant structure.

Frozen deployment weights are then treated as evidence of architectural inability to learn. But freezing is a choice about when updates occur. The serious challenges concern stable adaptation, retention, interference, and cost. Contextual adaptation and external memory supply other routes by which experience can affect behavior. The A–B–M proposal asks how to coordinate such processes with changes to the learned representations themselves.

Finally, the article moves from missing bodily experience and imperfect causal generalization to categorical absence of understanding. Those premises support narrower conclusions. Understanding what heat will do to a material does not require personally touching it, although manipulating it reliably may require additional sensorimotor knowledge. An account of coffee’s chemistry, preparation, and effects can be informative without reproducing its taste. The distinction between knowing a system and undergoing an experience mattered throughout our earlier essays; it should not disappear merely because the subject is artificial.

Shallow computations have limits, distribution shifts expose shortcuts, and long computations accumulate errors. Turning those findings into a rule that every success is imitation and every failure exposes the system’s essence prevents us from investigating either. It assigns the verdict before the experiment.

A useful evaluation starts with a novel environment whose states and interventions can be controlled. Test whether the system predicts consequences when irrelevant surface features change and when familiar elements appear in unfamiliar combinations. Compare success with plausible shortcuts, including access to memorized examples. Probe the candidate representations, then intervene on them to determine whether they guide behavior. Measure separately whether the system knows the structure, tracks its state, selects useful actions, and retains corrections after its temporary context is gone.

For a deployed agent, repeat the tests with its actual tools and memory, and account for their contributions. Include the cost of elicitation: examples, additional training, search, tool calls, and experimenter-supplied information. A capacity available only after extensive intervention is still a finding about the system, but it is a different practical resource from a capacity the agent can recruit on its own. Both negative and positive results become more informative when these conditions are stated.

To assess the research direction, turn these evaluations into comparisons across successive systems. For navigation, one could measure the fraction of held-out routes completed under specified disturbances, or the longest route completed at a fixed required success rate. For consequence prediction, one could measure error as the prediction horizon grows and conditions change. Specify the observations, available actions, task distribution, and inference budget. Track training resources as well. A plot against release date describes progress in the field; a plot against compute or data tests a more specific scaling claim.

Then state the expected progress. A proposal that greater scale should overcome a limitation predicts improvement on the relevant measure as resources increase; a proposed alternative may predict better performance at the same budget. A sustained plateau, deterioration under controlled conditions, or costs that grow prohibitively as errors shrink would weigh against the corresponding proposal. The expectation must be justified by the capability hypothesis, prior trend, or practical target. Demanding immediate perfection supplies none of those justifications. Conversely, a positive slope does not guarantee that a target will be reached within feasible resources. A claim of practical unviability needs that cost argument; a claim of architectural impossibility needs a stronger one still.

This is the missing comparison in an appeal to what current systems still cannot do. The relevant question is how the specified incapacity changes as the approach develops, and whether that change is consistent with reaching the stated goal. The studies considered here identify particular bottlenecks and particular improvements. They do not, taken together, establish a longitudinal failure of LLM-derived systems to progress on the capabilities Carrier requires. Investment in the AI ecosystem fuels these ideas too, even when scaling is the main attraction.

Transformers already learn structured representations of environments and use them to guide behavior. The studies reviewed here help explain how that achievement can coexist with failures of state tracking and control. They also show how changes in training and in access to learned representations can improve performance.

The efficacy intuition recognizes an achievement; the demand for better world models identifies ways to extend it. We already have systems that learn how parts of the world work. The task is to make that knowledge dependable in action and open to revision through experience.


Sources and scope. This essay follows “Which AGI?” while revisiting its attribution of a world model to the network. The supplied empirical papers are preprints; their results are reported within their tested settings. The supplied paper versions were Dupoux, LeCun, and Malik, arXiv:2603.15381v1; Kuhn and colleagues, arXiv:2607.00784v1; Teoh and colleagues, arXiv:2511.05963v4; and Beckmann, Queloz, and Freitas, arXiv:2609.21748v1. The Carrier passage was supplied as an extract from his article, originally dated October 29, 2025 and subsequently updated. The MLST discussion with Alexander Mattick and Tim Scarfe was consulted in the supplied caption transcript for the September 21, 2026 episode; timestamps refer to that transcript. The interpretive connections among these sources are our argument, rather than claims jointly made by their authors.