New Cognitive Taxonomy Aims to Turn AGI Progress Into Measurable Science
A new paper defines 10 cognitive abilities for evaluating general intelligence in AI, backed by a $200,000 Kaggle hackathon targeting the five hardest-to-measure abilities.

Updated
Why it matters
- The paper "Measuring Progress Toward AGI: A Cognitive Taxonomy" identifies 10 cognitive abilities hypothesized to matter for general intelligence in AI systems.
- A companion Kaggle hackathon offers a $200,000 prize pool, with submissions open March 17 through April 16 and results announced June 1.
- The hackathon focuses on five abilities with the largest evaluation gap: learning, metacognition, attention, executive functions and social cognition.
A newly released paper, "Measuring Progress Toward AGI: A Cognitive Taxonomy," proposes a scientific framework for evaluating the general intelligence of AI systems — and its authors are backing it with $200,000 in hackathon prizes to turn the theory into working benchmarks.
The stakes are straightforward. AGI, the authors write, "has the potential to accelerate scientific discovery and help solve some of humanity's most pressing problems." But nobody can currently say how close any system is to that milestone, because the field lacks empirical tools for evaluating general intelligence. The paper argues that cognitive science offers one piece of the solution.
Ten abilities, drawn from decades of research
The framework pulls from psychology, neuroscience and cognitive science to define a taxonomy of 10 cognitive abilities the authors hypothesize will matter for general intelligence in AI:
- Perception — extracting and processing sensory information from the environment
- Generation — producing outputs such as text, speech and actions
- Attention — focusing cognitive resources on what matters
- Learning — acquiring new knowledge through experience and instruction
- Memory — storing and retrieving information over time
- Reasoning — drawing valid conclusions through logical inference
- Metacognition — knowledge and monitoring of one's own cognitive processes
- Executive functions — planning, inhibition and cognitive flexibility
- Problem solving — finding effective solutions to domain-specific problems
- Social cognition — processing and interpreting social information and responding appropriately in social situations
The list is notable for what it includes beyond raw task performance. Metacognition, executive functions and social cognition are areas where today's benchmark suites are thinnest — and where the paper's authors themselves see the largest evaluation gap.
A three-stage evaluation protocol
Defining abilities is only the taxonomy's first half. The paper also proposes a three-stage protocol for benchmarking AI systems against human capabilities.
First, evaluate AI systems across a broad suite of cognitive tasks covering each ability, using held-out test sets to prevent data contamination. Contamination — test data leaking into training data — has repeatedly inflated reported model performance, which makes held-out sets a practical necessity rather than a methodological nicety.
Second, collect human baselines for the same tasks from a demographically representative sample of adults. That last detail matters: a benchmark anchored to a narrow pool of human respondents measures a narrower band of ability than it appears to.
Third, map each AI system's performance relative to the distribution of human performance in each ability — not against a single averaged score, but against the spread of how humans actually do.
The result, if the protocol works as designed, is a profile. A system could score at the high end of human distribution on reasoning while falling far below it on metacognition or social cognition — a far more informative picture than a single leaderboard number.
From framework to benchmarks, via Kaggle
The authors acknowledge that a paper alone will not measure anything. To move from theory to practice, they are partnering with Kaggle on a hackathon titled "Measuring progress toward AGI: Cognitive abilities."
The competition targets five abilities where the evaluation gap is largest: learning, metacognition, attention, executive functions and social cognition. Participants can build and test their evaluations against a lineup of frontier models using Kaggle's newly launched Community Benchmarks platform.
The prize structure totals $200,000. Each of the five tracks carries $10,000 awards for the top two submissions, and four grand prizes of $25,000 will go to the best overall submissions. The math is deliberate: 5 tracks × 2 awards × $10,000 plus 4 × $25,000.
Submissions opened March 17 and close April 16. Results will be announced June 1.
Why it matters
The announcement lands in a debate that has focused heavily on definitions — when a system counts as AGI, what thresholds matter — and comparatively lightly on measurement. AGI commitments now appear in corporate charters, safety frameworks and policy discussions, yet the empirical footing for tracking progress toward those commitments remains thin.
This framework does not settle the definitional argument. It takes a different route: borrow the construct vocabulary that cognitive science has refined over decades, benchmark systems against representative human distributions, and expose the whole apparatus to outside scrutiny through a public competition.
The hackathon also functions as an admission of limits. By directing prize money at the five least-measured abilities — learning, metacognition, attention, executive functions and social cognition — the authors concede that the hardest parts of the taxonomy are precisely the parts the field cannot yet test well.
Whether community-built evaluations can close that gap will start to become clear on June 1, when the hackathon results are announced and the taxonomy faces its first practical test.
Original: storage.googleapis.com
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
122 articles
Related articles
- OpenAI Predicts AI-Made Discoveries by 2026 as Intelligence Costs Plunge
- Google DeepMind's SIMA 2 Turns AI Into a Gaming Companion
- OpenAI Launches FrontierScience Benchmark for AI Research Skills
- OpenAI Warns Its Own Monitoring Tools Are Failing as AI Nears Self-Improvement
- OpenAI Partners With CodeAI to Train the First AI-Native Students