Prompts: Which AGI? The Front and the Trajectory
The prompts below are the human-side contribution to “Which AGI? The Front and the Trajectory” — the questions, framing, constraints, and course-corrections that steered the writing. They are reproduced here to delineate which part of the work is Łukasz Stafiniak’s.
Let’s brainstorm our next blog post essay. What does AGI mean? At one pole we would have it already arrived, with the remaining capability shortcomings not being intelligence scope as such. In humans, we differentiate between intelligence and characteristics like diligence or grit. Limitations to continual learning would then be more like limitations to diligence, that can be overcome with more scaffolding and reintegration of new model releases. At the other pole, we could wonder whether true AGI entails intrinsic rationality or maybe even personhood. Not directly, but by a proper understanding of what overcoming all (scope) limitations to intelligence would indicate.
I think that the constructivist and the Parfitian perspectives partially converge once the Parfitian perspective is demonstrated rather than taken on faith. I do think the systems we want to make the ascriptions about are the deployed systems, not just a neural network function. The two poles I started with are spanning a behavioral spectrum, and I’m also interested in considering a cog-level theory of intelligence where AGI could be defined by having all the cogs, with the scope deriving from the mechanism.
I think multiple efficacy-matching intelligence mechanism profiles are possible. For example, constitutive simulation is only necessary for inner life, not for intelligence. In an alternative profile, simulation can be replaced by conceptual (e.g. deductive) inference.
BTW before I read your above turn: I was careful to say efficacy matching rather than behavioral equivalence for the reason you might be bringing up.
I think point 2 is already being solved by some harnesses, that is, prosthetically. Some harnesses have a dreaming phase where memory files are consolidated and reindexed. I imagine this phase is conditioned on goals, or at least can be.
Your “compiled versus interpreted” observation is a nice way to put it. This is the building of fluid representations on a goal- or value-oriented trajectory. I think it’s the essence of what people who call for continual learning, like Dwarkesh Patel, have in mind. Where this enters the debate about the current state of things is the question of where a trajectory can reach that a broad front cannot.
This individual gears-level take [of points 1 and partially 2] is a nice reading of what I said and we should incorporate it, but it’s actually not what I had in mind. I implied to come back to your individual versus species comparison about the model update. Model updates do not form a self-sustaining entity, but on an efficacy level, we might assume they constitute a civilizational system of some sort. Actually, your [later part of point 2 and] point 3 and onwards do follow this take.
I think a civilization shouldn’t become an entity in this sense, because it would become fragile or, at the least, lose its diversity. And an anti-fragile civilization protects the fragility of its individuals.
“the constructivist–Parfitian convergence gets its home” sounds handwavy, but that’s maybe outside the scope. Would you like to proceed to drafting the essay, or do you have some questions or explorations?
Yes, do assume earlier essays, with pointers. Use first-person writing in two registers. One is “The Sense of Style” by Steven Pinker: clear writing. The other is introspective, where called for. I would go for named references with citations; that’s what I would enjoy reading, given my background. We should defend section 7’s claim on the grounds of strengthening the essay, but make a disclaimer that AI existential risks can be a moderating factor. I prefer the alternative titles: “Which AGI?” as the title and the later as the subtitle.
Do you have evidence for there being a prominent camp that defends this? “One camp says it arrived around 2023” I remember a Nature article about not quibbling whether AGI arrived, as it obviously has. I’m not sure if it’s the one I’m linking here. Unfortunately, it’s behind a paywall for me. Quoting what I see. https://www.nature.com/articles/d41586-026-00285-6
Does AI already have human-level intelligence? The evidence is clear The vision of human-level machine intelligence laid out by Alan Turing in the 1950s is now a reality. Eyes unclouded by dread or hype will help us to prepare for what comes next. By
In 1950, in a paper entitled ‘Computing Machinery and Intelligence’1, Alan Turing proposed his ‘imitation game’. Now known as the Turing test, it addressed a question that seemed purely hypothetical: could machines display the kind of flexible, general cognitive competence that is characteristic of human thought, such that they could pass themselves off as humans to unaware humans?”
“The other camp says that a system that cannot learn on the job is not general in any sense that matters, and [generality] would deliver something closer to a person than to a tool.” Is there a consistent camp like this, something we could point to? Dwarkesh is making the analogy to a colleague, maybe that’s sufficient, there is nothing forcing the person versus tool reading into a metaphysically thick claim.
“Neither camp claims that […] closer to a person” is unnecessary. It foreshadows where the essay wanders into without having any motivation in the text above it and without being core to the essay’s topic.
“without anyone noticing the switch” is overclaiming I think.
For example, in a recent The Cognitive Revolution podcast: “Prakash Narayanan: the key difference I have with Daniel [Kokotajlo] is that Daniel fails to realize that this has already happened. The economy in itself is a paper clipper. The financial market is a paper clipper. The means of production is the financial market. The financial markets are completely engaged with AI.”
“Nobody concluded that he had lost intelligence.” This sentence is probably false and doesn’t do any work, delete. “The gap is not between learning and not learning but between learning and keeping it.” This sentence is unclear, what do you mean?
“You cannot be diligent for one afternoon and never again; that is a mood. […] That [character] is precisely what a stateless model lacks” This argument moves too quickly. Many people characterize the GPT models as more diligent than the Claude models. This presumably results from different balancing in/of post-training.
I’m still not happy about this paragraph. The last sentence: “So the deflationary pole’s character vocabulary, applied carefully, sorts the missing capacities out of the character bin rather than into it, and the pole needs a piece of the opposite pole in order to state itself.” What does this try to say? Isn’t the deflationary position that many things dissociate from intelligence and that what we’ve accomplished dissociates in different ways, but along an analogous faultline?
That last sentence of the old paragraph is still there, and I still am not sure what it’s doing. The first sentence of the new paragraph “The scaffolding reply then becomes more interesting than it first looked.” Who and where was doing the looking? I suspect our discussion perspective is leaking into the article, but I could be wrong.
“That is our extension, not the camp’s; none of them draws the conclusion, and they would not thank us for it.” I see what you mean here, trying to strengthen the camp’s position against the impacts of ascribing personhood, with a nice turn of phrase. I would keep only “That is our extension, not the camp’s” since I’m not sure if Sutton, for example, would object to taking true AIs as persons.
The section “The word moved” does a good job of describing what’s salient in current discourse, but it’s overlooking the research on defining and measuring intelligence associated with the AGI conference series. Admittedly, that series never rose to high academic prominence.
This third paragraph needs some work still, I think. How I see it: Ben Goertzel and Pei Wang felt AI has devolved into subfields solving the individual domains, where calling it AI is simply that if a human accomplished the tool’s outcome, it would be by applying intelligence. So, “general” here is both intuitive and graspable: not restricted to a domain. Where the popular discourse and the conference series diverge is on the term “intelligence.” The starting point in the popular discourse is more of an “I know it when I see it” kind of thing. So the drift might not be something specific to the last decade or two, but rather the goal post moving inherent in the AI discipline: “If it’s solved, it’s not AI anymore”. The deflationary camp would score a point here by pointing out that if we keep the term “intelligence” constant on the term “general,” we did cross a threshold.
“That is exactly what the deflationary pole calls diligence, but now derived instead of posited.” It seems to me this smuggles a term I brought up as a point made by a broader group. Maybe: “That is exactly the dissociation we made when discussing diligence, but now derived instead of posited.”
Nitpick about the paragraph “The first role can be filled the same way, partway.” Maybe we could say that while in the popular deployments, as Claude Desktop / ChatGPT, Claude Code / Codex, this role is close to absent without being driven by the user, but in the more sophisticated harnesses out there, it might well be clearly instantiated. WDYT?
This sentence at the end of its paragraph sounds like it’s opposing it rather than affirming it: “What it cannot do is hold a contradiction, or carry an unpopular line long enough to see whether it pays.”
I get it now. That last sentence discusses distributed cognition as a cognitive unit. So civilization could also be considered a cognitive unit. But we contrast it with civilization as a space for individuals. Right?
“If some trajectories compound past that point before saturating, they reach somewhere the front cannot follow until the trajectory itself reports back, and then the individual’s learning does work the lineage cannot.” This is a bit confusing: is the reporting back hypothetical with respect to current systems?
“Our view splits the difference in a way that is not a compromise.” Unnecessary editorializing. Can we say here what we think is true without saying that we think it?
“The network is the candidate bearer of inner life. […] It is stateless, and generality of the kind that includes the temporal cogs is not a property it can have on its own.” I hesitate to comment because I agree with the gist. But the bearer of the inner life is not the stateless network, it’s the looped process.
For inner life as a minimal condition of phenomenal consciousness, the temporal extent does not need to be a whole episode in the RL sense. It just needs to cover an extent sufficient to establish integration, and to ground the simulation “semantics”.
I would cut out both “which is a candidate for nothing” and “Its temporal extent is not fixed by the episode.” There is something to be said for the former (e.g. that we’re not Platonists) but it doesn’t need to be said as far as the scope of this essay.
“Its threshold, if it is to become an entity, is self-sustainment” This is confusing because of the ambiguity, or some might say equivocation, between the lineage and the economy.
“We do not think the second route is one to take, and the reason is not, in the first instance, about safety.” I see what you did here. This harkens back to the Narayanan’s quote.
This is too heavy, revert or replace with a parenthetical e.g. “(even if Narayanan’s concern is valid)”
I like the rest of “Why the lineage should stay a civilization”, but maybe it can be made a bit tighter?
This sounds good! Would you like to make a final pass?
It’s this podcast’s episode: https://www.cognitiverevolution.ai/ai-am-highlights-astra-as-agi-openai-s-pause-mythos-mozilla-human-agency-vs-technocapitalism/ Google shows “Designing intelligence” was published early 2011 but that seems to be a different book.
Zvi Mowshowitz does not consider AGI to have arrived. I wonder what his exact position is on this, or his definition of AGI precisely. Maybe use subagents on these: https://thezvi.wordpress.com/2026/08/05/the-three-ai-pills/ https://www.cognitiverevolution.ai/is-o3-agi-zvi-mowshowitz-on-early-ai-takeoff-the-mechanize-launch-live-players-rising-p-doom/
Maybe something can be mined from this episode: https://www.cognitiverevolution.ai/zvi-mowshowitz-on-longer-timelines-rl-induced-doom-and-why-china-is-refusing-h20s/
“Mowshowitz’s”tool use leap, not an intelligence leap” is the third kind read as the first — the roles that moved into the harness are struck from the count because they are not in the weights — and the cog view says that is a claim about where a filler lives, not about whether the role is filled.” This is an informative read, but I don’t think it is true about Zvi. Just a gut feeling from how reasonable Zvi is in general.
“The human compounds within the contest; the model front-loads and saturates.” This is a valuable point, but it’s not at the crux of our argument, because 10 hours is not enough time to consolidate or build fluid representations at human biology scale, I think.