Few terms in technology generate more heat and less light than AGI — artificial general intelligence. Depending on who you ask, it’s either imminent, already here, decades away, or a meaningless concept. What it hasn’t had, until now, is a rigorous, agreed-upon way to measure progress toward it.
Google DeepMind is trying to fix that.
The Problem With AGI as a Concept
The idea of AGI — an AI system that can match the broad, flexible intelligence of a human being, rather than excelling at narrow specific tasks — has driven enormous investment and speculation in recent years.
As large language models have taken on more and more tasks that once seemed exclusively human, claims about crossing the AGI threshold have multiplied.
But those claims have been largely impossible to verify or refute, because there’s been no agreed definition of what AGI actually requires, let alone how to test for it. The result is a debate full of assertion and very little measurement. As DeepMind’s own researchers put it, the ambiguity “fuels subjective claims, makes it difficult to track progress, and risks hindering responsible governance.”
Their new framework is an attempt to put the conversation on firmer ground.
Ten Building Blocks of General Intelligence
Drawing on decades of research across psychology, neuroscience, and cognitive science, the DeepMind team breaks general intelligence down into ten core faculties. Eight are foundational: perception — taking in sensory information and producing outputs like text, speech, or actions — alongside learning, memory, reasoning, attention, and metacognition, the ability to reflect on and regulate your own thinking. Executive functions like planning and impulse control round out the eight.
The remaining two are what the researchers call composite faculties, requiring several of the building blocks to work together simultaneously. These are problem solving and social cognition — the ability to read and respond appropriately to social context, understanding not just what is being communicated but how and why.
The framework deliberately focuses on what a system can do rather than how it does it, keeping it technology-agnostic. Whether a capability is achieved through a large language model, a different architecture, or something not yet invented shouldn’t matter — what matters is the outcome.
How Progress Would Be Measured
Identifying the ten faculties is only half the proposal. The more practically ambitious part is the evaluation methodology. The team suggests subjecting AI systems to a broad battery of cognitive tests targeting each specific ability, then collecting human baselines under identical conditions — using a demographically representative sample of adults with at least a high school education.
The results would be combined into a “cognitive profile” for each system, mapping strengths and weaknesses across all ten dimensions. Comparing those profiles against the human baselines would then provide an empirical answer to the question: has this system matched or surpassed average human general intelligence, and in which areas?
It’s a meaningful shift from current AI benchmarks, which tend to test narrow capabilities in isolation and often produce headline numbers that are impressive-sounding but hard to contextualise.
The Gaps Still to Fill
The researchers are candid about the limitations of their own framework. For some faculties — problem solving, perception — there are already well-established benchmarks. For others, including metacognition, attention, learning, and social cognition, no reliable tests currently exist. Building them is the next challenge, and the team says they’re working with academic partners to develop robust, non-public evaluations that can’t be gamed by training on the test material. That last point matters: many of the best existing benchmarks are publicly available, meaning AI models may have effectively seen the answers during training.
There are also deeper questions the framework can’t fully answer yet. Do these ten traits genuinely capture the essence of human general intelligence, or is something important missing? And even if a system aces every test, does that translate into better performance on real-world problems compared to specialised AI built for specific tasks? Those questions remain open.
Why It Matters Anyway
Even with those caveats, a framework grounded in established cognitive science and designed around measurable, comparable outcomes is a significant improvement on what came before. AGI has been treated for too long as a milestone that different people define differently and claim to have reached — or nearly reached — according to their own standards.
If the DeepMind framework gains traction, it would give researchers, policymakers, and the public something more useful: a shared vocabulary, a set of specific tests, and a way to actually track whether the technology is getting closer to human-level general intelligence — rather than just asserting that it is.
The question of what AGI means may never have a clean answer. But the question of how close we are to it at least deserves a cleaner way of asking.
