Where Does Seeing Become Thinking?
Representations, Operations, and the Perception–Cognition Boundary in AI
Consider a photograph of a crowded room. We ask an AI assistant whether a wheelchair could pass between the sofa and the table. The assistant identifies the furniture, describes the arrangement, and produces a fluent explanation of accessibility. Then it recommends a route through a gap that is too narrow.
Where did the failure occur? Perhaps the system never extracted the relevant dimensions. Perhaps it represented them but lost track of which distance belonged to which gap. Perhaps it knew the dimensions but failed to work out how rotating the wheelchair would change its clearance. Perhaps it reasoned correctly about a room assembled from its expectations rather than the one in the photograph. The image might also lack the information needed to decide: a convincing perspective view is not a measured floor plan.
These possibilities concern different achievements. They also cross a boundary we often assume we already understand: the boundary between seeing a situation and thinking about it. Once the sofa and table have been recognized, has perception finished its work? When the system estimates clearance, is it still seeing, or has it begun to reason? When it imagines moving the table, what keeps the imagined room distinct from the actual one?
In “Which World Model?”, we argued that failures of action do not by themselves establish an absence of structured environmental representation. A model can possess a map and lose its position. The distinction was supported by interventions on a transformer trained on Manhattan routes: researchers could distinguish learned street structure from mechanisms that maintained location and directed travel. But having separated representation from successful use, we owe a fuller account of what using a representation requires.
An earlier essay, “What Is a Concept?”, approached this question through representational format. A concept can be characterized by its role in categorization, inference, and action while we investigate what kind of structure fills that role. The resulting questions concern constituents, composition, similarity, and the operations a representation supports. They also concern whether perception and conceptual thought use different formats.
The two essays meet here. What an artificial system can think about depends partly on what it can represent. What it can do with those representations depends on their organization, the transformations it has learned, and how it keeps them connected to evidence. Investigating that connection may tell us whether frontier multimodal systems have a perception–cognition boundary, and what kind of boundary it is.
What Visual Reasoning Requires
Andrew Dai’s September 2026 conversation on The Information Bottleneck gives the question a practical setting. Around 27–34 minutes, he distinguishes object recognition and segmentation from counting, interpreting assembly diagrams, and working out spatial arrangements. Around 37–42 minutes, he discusses access to fine visual detail, more efficient video representations, and training examples in which people edit images to express intermediate visual thinking. The proposal connects a familiar complaint about AI vision to a particular need: representations on which useful operations can be performed.
Counting illustrates the difference. Recognizing an apple is insufficient for counting the apples in a bowl. A successful procedure must individuate candidates, resolve ambiguous boundaries, avoid counting one object twice, and maintain progress. Procedures may combine or separate these steps, but naming the objects correctly leaves the counting unfinished. An account of the failure that stops at “the model knows what apples look like” has not explained the relevant competence.
Dai sometimes makes stronger claims: that reasoning is fundamentally visual, that visual reasoning should form the foundation on which linguistic reasoning is built, and that current text-based reasoning leaves a central human capacity unimplemented. These claims need to be separated from his diagnosis of particular failures. The evolutionary priority of vision does not establish a necessary engineering order. Nor does a textual reasoning trace tell us that every computation producing it has a linguistic format. Those are precisely the distinctions that made the earlier world-model debate tractable.
Christopher Manning’s interview on the same podcast supplies a useful counterweight. Around 29–34 minutes, he argues that language helps organize thought within an individual and coordinate knowledge across a society. Around 40 minutes, he distinguishes what can be learned through language from how efficiently it can be learned: images and video can make spatial knowledge easier to acquire. On this view, language contributes cognitive tools as well as descriptions of perceptual discoveries. For our room example, a phrase such as “the narrowest point along the route” can organize a sequence of visual comparisons. The question becomes how language directs those operations and remains answerable to their results.
A reading conversation about “Which World Model?” arrived at a more discriminating phrase: modality-specific representational manipulation. It shifted attention from whether a system possesses spatial knowledge to how perceptual information survives conversion, maintenance, and transformation during reasoning. But modality itself may be too coarse a distinction. A photograph, an annotated photograph, a line drawing, and a coordinate table can concern the same scene while making different operations easy. Prose, algebra, and executable code likewise offer different computational resources despite all being expressible as text.
We need to distinguish the source of information, the format of its internal representation, and the operations available over that representation. None uniquely determines the others.
Which Perception–Cognition Boundary?
Ned Block’s The Border Between Seeing and Thinking provides a demanding account of format. His conception of iconic representation emphasizes analog tracking and mirroring: degrees of difference in what is represented correspond to degrees of difference in the representation, and relevant relations are preserved. Iconicity does not require a literal picture, mathematical continuity, or an absence of parts. Block argues that perception is constitutively iconic, while cognition can also employ iconic representations—for example, in reasoning about moving furniture. His disagreement with perceptual pluralists includes whether object-tracking evidence requires a discursive constituent identifying an object, and whether some experiments concern perception or conceptual working memory. This is a dispute about format and functional role, not a simple division between visual input and verbal output.
Jake Quilty-Dunn’s perceptual pluralism allows both iconic and discursive representations within perception. Its object representations can have a structure more like an individual bearing separately represented properties than an undifferentiated picture. Yet shared formats need not dissolve the distinction between perception and cognition. His dissertation defends a distinction based on encapsulation and stimulus control instead of format. A system may therefore use related representational structures in perception and thought while differing in what information those processes admit and what governs their operation.
This qualifies a suggestion in our concept survey. Finding compositional structure in perception would undermine a boundary defined exclusively by iconic versus discursive format. It would leave other boundaries available. The question for AI is whether several distinctions that are often bundled together actually coincide.
One is architectural. A system might contain an image encoder, a component that connects visual representations to a language model, and a backbone that produces answers. These divisions locate computations and constrain information flow. They do not establish that the first component only perceives or that the last only thinks. A model can infer object identity within its visual processing and resolve visual details during answer generation. Conversely, putting all processing into a common architecture does not establish that it has no functional divisions.
Another is representational. Does a computation preserve spatial relationships through corresponding relationships among internal states? Does it distinguish an object from the properties currently attributed to it? Can a relation be applied to new participants while preserving its role? Answering these questions requires examining the organization and use of representations. Their origin in an image or a sentence does not answer them.
A third distinction concerns what governs the representation. A representation of the room currently shown should remain sensitive to the photograph. A representation of how the room could look after moving the table should change according to the proposed intervention. Both may use spatial machinery. Their different responsibilities become visible when the agent must compare them or return from the hypothetical arrangement to the original one.
A fourth distinction concerns access and control. Some information may influence a response without being available for deliberate comparison, retention, or report. Other information can be recruited across many tasks. An agent that can answer a question about a visible object may still fail to retain that object’s identity while reasoning about a transformation. Conversely, an inability to describe a representation does not show that the representation is inactive. We need to investigate which processes can use it and under what conditions.
These distinctions give the boundary question empirical content. We can ask whether a change of format occurs at an architectural interface, whether the same format supports stimulus-directed and hypothetical processing, and whether information becomes more flexibly available as it passes through the system. The result might be a single boundary or several partially overlapping divisions.
When Geometry Does Computational Work
Interpretability research lets us investigate how a model represents information and uses it. Our question about preserving spatial relationships leads naturally to the geometry of neural representations: the distances, directions, and shapes formed by patterns of activation. That geometry could help explain how a model represents the world and reasons about it. But any activation pattern can be treated as a point in a mathematical space; placing it there does not explain what the model does with it. We need to establish which relationships in the world the geometry captures and how the model’s computations use them. A projection can reveal structure, conceal it, or make an investigator’s preferred variables unusually conspicuous.
Manning makes a related methodological point around 52:19: linear accounts of concepts may receive disproportionate attention because they are easy to study. He acknowledges the success of low-dimensional subspaces while emphasizing superposition and more complex organization. This sharpens our concept survey’s concern about the assumptions built into interpretability tools. A useful method of reading a representation need not exhaust its structure.
Kiho Park, Yo Joong Choe, and Victor Veitch’s work on the linear representation hypothesis addresses part of this problem. It formalizes relationships among conceptual directions, probing, and steering, and introduces an inner product intended to respect specified causal relationships among concepts. The methodological lesson is that the geometry used to analyze a representation needs justification. An angle between two activation directions is not automatically a cognitive relationship merely because it can be plotted.
More concrete examples show geometry participating in computation. Joshua Engels and colleagues find circular representations of weekdays and months. They intervene on these representations in Mistral 7B and Llama 3 8B while the models solve calendar arithmetic problems. Changing the circular state affects the answer, supplying evidence that the structure contributes to the computation. The intervention procedure also controls other activation components to isolate the relevant subspace.
The example is revealing because the input can be entirely linguistic while the internal organization reflects a cycle. Moving several days forward has a wraparound structure that ordinary distance along a line would fail to preserve. Here the geometry organizes an abstract calendar relation. The interventions connect that organization to the model’s answers, showing how geometric structure can contribute to computation even on a task that requires no visual input.
An especially relevant study is Wes Gurnee and colleagues’ When Models Manipulate Manifolds. They examine how Claude 3.5 Haiku predicts line breaks in fixed-width text. Character counts and line-width information occupy curved, low-dimensional structures within the larger activation space. Attention heads transform these structures to compare current position with the approaching boundary. The researchers support their account with interventions and ablations. They also show how the same representation can be described through discrete features or through a continuous manifold, with different explanatory advantages. Their account leaves aspects of the complete linebreaking computation unresolved.
This is a useful complication for concept theory. A decomposition into features and a geometric account of their collective behavior can describe the same mechanism. Discovering apparently discrete features therefore does not automatically establish a language of thought; discovering a smooth manifold does not automatically establish a wholly different kind of cognition. The contrast that matters may concern the operations supported by an organization rather than whether our preferred description uses points or curves.
Three levels of evidence should remain distinct. A researcher may recover a variable from internal states. Intervening on those states may establish that they affect the answer. Explaining the transformation performed by the model requires more: identifying how the relevant state changes, what remains invariant, and how the operation behaves on new combinations. Decoding orientation, causally altering an orientation-sensitive judgment, and establishing a general rotation procedure are progressively more demanding achievements.
A recent study by Junzhe Ren and Jiayuan Zhang offers a lead at the second of these levels. On a controlled rotated-versus-mirrored object task, they report an angle-related direction in LLaVA-1.5-7B and projection ablations that progressively flatten the relationship between angle and decision margin. Their abstract interprets this as evidence for spatial processing with a causal role, while withholding the stronger claim of human-like mental rotation. Establishing the latter would require distinguishing a transformation process from other computations sensitive to angular difference.
From Representations to Reliable Operations
The visual failures that motivated our inquiry should be treated with the same care. Sai Srinivas Kancheti and colleagues evaluate seventeen models across thirteen spatial benchmarks. They find substantial problems with textual chain-of-thought in their tested systems. A control replaces the image with an uninformative gray field and adds an option acknowledging insufficient visual evidence; several models still produce detailed visual claims. But the proprietary-model comparison is mixed: some improve slightly with chain-of-thought, while GPT-5 and GPT-5-nano perform slightly worse. The paper therefore supports concern about particular reasoning procedures and training regimes, not a universal claim that additional verbal reasoning damages vision.
Zhuoran Jin and colleagues’ Look Light, Think Heavy similarly finds that chain-of-thought can hurt visual grounding and object counting while helping mathematical, scientific, and multi-image reasoning. The task dependence matters. A procedure useful for drawing out an implication may be unhelpful when the required information must first be extracted accurately from a crowded image.
These findings do not yet locate a perception–cognition boundary. If a model confidently describes a nonexistent object, its output leaves open whether it failed to extract the evidence, lost it, outweighed it with expectations, or generated an answer through a pathway insufficiently constrained by it. The vividness of a fabricated visual description does not resolve those alternatives. A verbal trace is behavior to explain, including when the trace claims to report what the system saw.
Internal interventions can help locate the failure. Qiming Li and colleagues’ cross-modal causal tracing study examines visual object information across decoder layers, attention, feed-forward components, and visual and textual token positions. They report a role for middle-layer attention in aggregating information and develop an intervention that reinforces intermediate object representations. Evaluation includes LLaVA and Qwen-family models and InternVL2-8B, with improvements on hallucination-related measures. This supplies evidence that some failures can be addressed through the handling of visual information within the decoder. A longtime reader of our blog will notice a connection to our evolving account of phenomenal consciousness.
Manning’s discussion of representation finetuning, around 45–49 minutes, connects this to the distinction between stored knowledge and its use. In Zhengxuan Wu and colleagues’ ReFT, the base model remains frozen while additional parameters learn interventions on hidden representations. The resulting changes in task performance show how much behavior can depend on how internal states engage the model’s existing machinery. These interventions are themselves trained, so their success cannot simply be credited to a fully formed capacity waiting to be uncovered. The relevant mechanism includes both the learned computation and the way it is accessed and directed.
We can now sharpen the role of representational manipulation. Two systems may retain enough information to answer a question in principle while differing greatly in the computation needed to answer it. An adjacency list and a drawing can encode the same graph. A sequence of coordinates and an image can describe the same arrangement. Recoverability establishes an informational possibility; a learned operation makes that possibility available at a particular cost and reliability.
This also explains why adding a verbal description may fail to help. A description can discard precisely the relationships the next operation needs. It can preserve them but make their use expensive. It can be accurate locally while leaving the agent to reconstruct a global arrangement from many statements. The difficulty then belongs to the interaction between a format and a procedure. Merely extending the procedure supplies no guarantee that it will preserve the relevant information more faithfully.
Consider describing every piece in a mechanical assembly. Even a correct inventory leaves unanswered which surfaces meet, which openings align, and what motion is possible without collision. An exploded diagram, a contact graph, and a geometric model would support different parts of the inquiry. Choosing among them is itself a cognitive achievement. The agent must recognize what its present representation makes difficult and obtain another representation that preserves the relevant constraint.
Clearance illustrates what it means to characterize a concept by its role. Having this concept should let a system connect object dimensions, possible movement, and obstruction across different situations. Using it may draw on a geometric estimate, a remembered rule, or a measured diagram. No single format need be the concept’s sole vehicle. We can investigate how these resources cooperate, which transformations preserve the relation, and what happens when the system encounters an unfamiliar arrangement.
Three Ways to Test the Boundary
To test for a perception–cognition boundary, we need to compare how systems implement potentially shared capacities. A system operating entirely through cognition in Block’s sense could match the task performance of one with distinct perceptual processes, while differing in efficiency, interference patterns, representational decomposition, and causal dependencies. We therefore need to investigate those differences and ask whether they recur across tasks. Whether the resulting organization qualifies as perceptual depends on the account of perception being tested. Three families of experiments would be particularly informative.
The first concerns object identity. Construct sequences in which objects move, change visible properties, and pass behind occluders. Ask questions that distinguish tracking an individual from matching its current appearance. Changing the color of an object should not necessarily change which object it is. Exchanging the paths of two similar objects should matter when the question concerns their histories. Different manipulations put pressure on different candidate representations.
With internal access, we could test whether identity and properties can be altered separately, and whether the resulting combinations influence several tasks consistently. Merely decoding color and shape independently would be insufficient: separate probes can recover variables that the model’s own computation never treats as independently reusable constituents. A successful intervention would selectively change which property belongs to which tracked object, producing the predicted changes in tracking, comparison, and transfer to unfamiliar combinations while preserving unrelated object–property assignments and object continuity. Controls would distinguish this selective effect from general damage to the representation. This would establish an object–property association used by the system, without yet establishing perceptual binding in the narrower philosophical sense.
To investigate the perception–cognition boundary, we would then ask when this association is constructed, what maintains it, and which processes can modify it. Binding maintained through stimulus-driven updating may have different costs and dependencies from an equivalent object–property assignment reconstructed during deliberate reasoning, even when both support the same answers. Does the structure participate in maintaining a representation of the scene before an explicit question is asked? Does it arise only when an instruction demands deliberate tracking? Does it persist when the image is no longer available? These distinctions would let the artificial case engage the dispute about perception and working memory rather than merely borrowing its terminology.
The second family concerns the influence of expectations. Keep the visual stimulus fixed while changing a caption or prior instruction. Compare a neutral description with one that suggests an absent object or a misleading spatial arrangement. If the answer changes, trace where the difference enters. It may alter selection of image regions, the representation of the object, later interpretation, or the final decision. Those are different forms of influence with different implications for perceptual encapsulation.
The converse manipulation matters too. Keep the instruction fixed while changing a small but decisive visual detail. A useful system should sometimes resist linguistic expectations because the image contradicts them. Yet sensitivity to context is not automatically corruption: context can identify which of several visible objects matters or resolve an otherwise ambiguous instruction. We need conditions in which the evidence warrants different degrees of reliance on prior information. The question is how that reliance is controlled.
Learning visual representations through language supervision does not by itself establish that current linguistic expectations can alter perceptual processing. In the proposed experiment, a misleading caption might change the visual representation, or only how subsequent reasoning uses it. Locating that influence is essential to deciding whether the result bears on perceptual encapsulation.
The third family concerns the difference between observation and supposition. Show an arrangement, ask the model to reason about moving one object, then ask separately about the hypothetical and original arrangements. Vary whether the image remains accessible, whether the agent has an external scratchpad, and how many transformations intervene. Measure whether changes meant for the imagined scene contaminate answers about the observed one.
This is more demanding than producing a plausible transformed picture. The agent must maintain the identity of objects across two situations while attributing different properties to them. It must know which situation a question concerns and recover when the distinction is lost. A representation can be geometrically rich while lacking reliable ways to keep these alternatives apart. Conversely, an explicit symbolic record may preserve their distinction even when the model’s visual manipulation remains weak.
Human imagination complicates this comparison: it can recruit perceptual resources while being experienced as imagined rather than currently observed, with substantial individual variation in vividness and detail. The experiment therefore asks how a system distinguishes the source and status of representations that may share a format or processing resources. In AI, we can investigate that distinction functionally without presupposing a phenomenal marker. Imagery also cautions against assuming that the perception–cognition boundary and the distinction between phenomenal experience and cognitive access coincide. Block allows perceptual materials to enter cognition; imagery raises the question of what remains perceptual when those materials are generated and manipulated in the service of thought.
Together, these experiments could reveal a division between stimulus-dependent updating and controlled counterfactual use. They might instead reveal several procedures whose responsibilities vary with the task. Either outcome would improve on assigning “perception” to everything before a module boundary and “reasoning” to everything after it. The point is to discover an organization that predicts behavior under new interventions.
What Architectural Differences Change
The experiments above concern capacities that different architectures might implement in different ways. Architecture constrains which representations can influence which computations, when that influence arrives, and how much work is required to use it. These constraints also shape learning. Our earlier essay on dynamics distinguished the processes that change a model during training from those that operate on its representations during inference. Here we need that distinction to ask a narrower question: how can the demands of a task affect the processing of perceptual evidence?
Biological vision supplies concrete examples. In Li, Piëch, and Gilbert’s experiments, neurons in macaque primary visual cortex responded differently to an identical stimulus depending on which visual discrimination the animal performed. Task dependence reached the processing of visual information itself. It need not operate solely by selecting another object to look at or interpreting a completed percept. For our comparison, the relevant architectural question is how information about the task reaches the machinery processing the stimulus.
Action supplies another kind of anticipatory signal. An efference copy carries information about an outgoing motor command that can be used to predict its sensory consequences. Sommer and Wurtz identified a corollary-discharge pathway through which information about an impending eye movement reaches frontal eye fields. Interrupting the pathway impaired the shifts in visual receptive fields that normally precede the movement. This offers a mechanism relevant to maintaining visual stability as the eyes move. For this comparison, what matters is the routing and timing of information: the visual system can prepare for a change in input before that change arrives. Whether such sensorimotor coordination crosses a perception–cognition boundary is a further question.
What can a feedforward architecture support? Within a conventional fixed-depth pass, a later activation cannot return to change an earlier activation already computed. But an instruction available at the start can shape subsequent processing, and later layers can construct refined representations of the same scene. In an architecture with a separate, task-independent image encoder, the instruction cannot alter that encoder’s current representation unless there is an additional route for it to do so. It can still affect the handling of visual information downstream. The distinction between revising an encoding and refining a representation built from it matters, especially when the original encoding discarded information. Neither “feedforward” nor “multimodal” alone tells us how much task-directed perceptual processing is available.
A bounded recurrent computation can be unrolled into a feedforward computation, but that construction preserves the recurrent circuit’s parameter sharing. It does not establish that an independently trained stack of layers will acquire the same organization. Even under one global loss, separate layers encounter different inputs and gradients and can specialize in different stages. Recurrent circuitry must work across successive states using shared parameters. The Universal Transformer explicitly explores this recurrent inductive bias. For our purposes, the hypothesis is that these different learning constraints affect the acquisition of reusable refinement procedures. Reuse can encourage such procedures without guaranteeing convergence, correction, or fidelity to evidence.
The arrangement of the recurrence matters too. Geiping and colleagues demonstrate a language model that iterates a recurrent block, allowing additional computation in latent space before producing a token. Generating further reasoning tokens also creates a route by which previous results can condition subsequent computation, but the intermediate information must pass through the generated sequence. Feedback into a visual encoder would provide yet another route. These arrangements differ in which states remain available and which computations can revisit them. A loop confined to a multimodal backbone could still refine visual representations; returning to the image encoder is not the only possibility.
A particularly informative intervention is Mozer and colleagues’ Recirculation. It mixes deep-layer activations into shallower processing across input steps, supporting state propagation at the same layer instead of adding only depth within each step. The basic method leaves pretrained weights unchanged; an adaptive variant learns mixing coefficients. On their Gemma3 evaluations, recirculation yields more robust improvements than training-free layer repetition. Gains vary across tasks, and processing the prompt sequentially imposes a cost. These are language and reasoning experiments, so they do not establish feedback-driven visual refinement.
The intervention also provides evidence about representations already present in the model. The authors motivate it through the residual stream’s resemblance to a shared blackboard: contributions written by different layers can remain compatible enough for deeper activations to be useful earlier. Successful feedback with unchanged backbone weights supports some such compatibility, while leaving open its extent and mechanism. Separate layers can acquire different roles without their representations becoming mutually unusable. This qualifies the learning-pressure argument: feedforward training can produce resources that benefit from a feedback route absent during training. Architecture shapes both what is learned and how learned resources can subsequently be used.
These distinctions should also govern claims about frontier models. The public GPT-6 Astra system card and GPT-6.1 Sol addendum do not specify the feedback pathways needed for this comparison. Speculation that a model uses looping cannot establish which representations are revisited or what the loop accomplishes. Documented recurrent architectures show possibilities that experiments on particular frontier systems would still need to identify.
The resulting research question concerns what changes when we alter the architecture. Compare otherwise matched models with different routes for task information: connections into visual encoding, into later scene representations, or only into answer generation. In an acting system, add or remove a pathway carrying information about the intended action into sensory processing before its consequences arrive. Keep the visual evidence and task fixed, then measure changes in accuracy, processing cost, interference, and recovery from misleading expectations. Modifying a trained model tests what its existing representations can support; training matched variants with those pathways present tests how they shape learning. These comparisons would connect architectural changes to the efficiency profiles and representational decompositions sought by our earlier experiments.
Thinking with the Environment
The investigation also needs to include the agent’s environment. David Kirsh and Paul Maglio’s study of epistemic actions in Tetris distinguishes moves that directly advance the game’s objective from moves that make information easier to obtain or compute. Rotating a piece can simplify the player’s cognitive task. An action that appears unnecessary when viewed solely as execution of a plan can help determine what the plan should be.
For an AI assistant, marking counted objects, cropping an image, drawing a possible route, or executing a geometric calculation can play a similar role. The resulting deployment has operations unavailable to a model required to answer immediately. Their contribution should be measured explicitly. Did the tool reveal information, preserve intermediate state, perform a transformation, or supply a result the model could not compute? Did the agent select and use it successfully without an experimenter preparing the solution?
This extends the distinction between network and deployment in “Which World Model?” Here the loop runs through action and renewed sensory input: reasoning selects an informative view; new visual evidence changes the working representation; an annotation preserves a decision; the altered artifact becomes something the agent can inspect. The agent’s ability to organize this loop may matter as much as the quality of any one representation within it.
Return to the crowded room. A responsible sequence might begin by recognizing that the photograph does not determine clearance. The agent could request a measurement, locate a supplied scale, or consult a floor plan. It could then construct a simplified geometric arrangement, test candidate movements, and compare its assumptions with the original evidence. Each step has a different informational purpose. Success depends on preserving those purposes through the sequence, including recognizing when an attractive drawing is only a proposal.
There is a connection here to the speculative argument at the end of our concept survey. We suggested that perceptual engagement might help calibrate even abstract reasoning by keeping the processing anchored to stable tokens and correction signals. The new formulation is narrower and more testable: external representations may improve reasoning when they preserve distinctions that internal processing otherwise loses, or provide evidence against an erroneous intermediate state.
A drawing can also make an error durable. An agent that mistakes its own sketch for an observation has gained a persistent source of misinformation. The benefit therefore depends on how artifacts are produced, checked, and used. We should compare revisiting the original evidence with rereading a self-generated description, and distinguish corrections driven by new information from confidence produced by repetition. This turns a broad grounding intuition into questions about particular sources of reliability.
Abstraction introduces a related difficulty: a representation adequate for one task may omit details essential to another. “A bowl of apples” can be adequate for one decision and inadequate for exact counting. A floor plan may preserve connectivity while omitting a threshold that blocks a wheelchair. No summary can be judged sufficient independently of the questions it must support.
An agent could respond by retaining access to the source and revising its abstraction when the task changes. It would need to detect when its present representation is inadequate, locate the missing evidence, and integrate it without losing what was already established. We might call this revisable abstraction. It connects perception, memory, attention, and reasoning through a practical demand: compress enough to think effectively while preserving a route back to the distinctions that can prove the current thought wrong.
Evidence Across Models and Through Training
How much of this organization exists in frontier multimodal systems? The evidence assembled here does not support a single architectural verdict. Some of the strongest mechanistic results concern smaller or older models, chosen partly because their internals are accessible. Behavioral evaluations cover other systems, procedures, and dates. We cannot silently combine them into a description of one hypothetical frontier model possessing every demonstrated mechanism and every reported failure.
Instead, the frontier question calls for linked levels of evidence. Evaluate accessible leading systems on matched tasks with their actual tools and inference budgets. Investigate candidate mechanisms in systems that permit internal intervention. When a laboratory provides mechanistic results for a closed model, keep the conclusions attached to that model and experiment. Then test whether the proposed mechanism predicts behavioral signatures that survive across systems. Agreement would strengthen the hypothesis; it would not erase the difference between prediction and direct inspection.
Manning’s interview raises a related question: how does this representational organization develop during training? Around 17–28 minutes, he discusses Jasper Jian’s study with him of category learning in GPT-2 small. Their measures detect differences between verb classes before differences among individual verbs within a class; they interpret this as evidence for abstraction-first learning. This directly engages the concept survey’s contrast between learning categories and accumulating exemplars. Manning also notes that the model’s small size could influence the ordering.
The interpretation is contested. Zachary Nicholas Houghton and Vsevolod Kapatsinski’s Exemplars in Disguise shows in controlled simulations that learners without shared class representations can reproduce that onset ordering under the same criteria. The observed trajectory therefore needs additional evidence to identify its mechanism. For our proposed visual experiments, tracking training checkpoints could reveal when object tracking, object–property associations, and counterfactual manipulation become available. Interventions and transfer tests would then help determine whether those abilities develop through shared representations. This could help explain how a boundary observed in a finished model developed.
This approach leaves room for meaningful comparison with humans. We need not suppose that human cognition contains a perfectly unified spatial model or that artificial mistakes reveal an absence of understanding. We can compare how different systems respond to forced verbalization, external diagrams, interrupted access to a stimulus, and requirements to maintain several alternatives. Similar final accuracy can coexist with very different costs and patterns of interference. Those differences are evidence about organization.
Crossing in Both Directions
The inquiry leaves several possibilities open. Artificial perception and cognition might be separated by a change of format. They might share formats while differing in control and access. Some distinctions might exist only within particular tasks, while broader coordination is supplied by the deployment.
The earlier world-model essay established that structured knowledge can coexist with unreliable behavior. The concept survey asked what formats make conceptual operations possible. Their continuation concerns how a system moves between tracking what is present, manipulating what could be present, and checking the result against what is actually there.
A useful working hypothesis is that perception and cognition can share representational resources while differing in how those resources answer to evidence and become available for flexible use. Geometric structure can make a transformation tractable. Compositional structure can preserve which object has which property. Control can keep a supposition separate from an observation. In the crowded room, these capacities must cooperate: the assistant needs to estimate clearance, track the furniture through a proposed rearrangement, and keep that proposal distinct from what the photograph shows.
The boundary between seeing and thinking, if we find it, will have to explain these transitions. And the most consequential capacity may be the ability to coordinate them in both directions: to let perceptual evidence correct thinking, and to let the demands of thought shape how that evidence is perceptually extracted and organized.
Sources and scope. This essay continues “What Is a Concept?” and “Which World Model?” The interviews with Andrew Dai and Christopher Manning on The Information Bottleneck were consulted through supplied caption transcripts; approximate timestamps refer to those transcripts. Manning’s discussions of category learning and representation finetuning were checked against the linked papers, with the subsequent methodological critique of category-learning trajectories included. Block’s position was checked against the accessible fifth chapter of The Border Between Seeing and Thinking, rather than assuming access to the entire book. The shared reading conversation supplies a framing question, not independent empirical evidence. Research results are attributed to their tested models and settings; the Ren–Zhang study is discussed at the level of its accessible abstract. The proposed experiments, revisable-abstraction account, architectural comparisons, and synthesis of the perception–cognition question are our argument. The architecture section uses published mechanisms and interventions; it does not infer undisclosed frontier-model designs from speculation about looping. This draft reports no new model experiments.