Google DeepMind proposes a cognitive framework for measuring AGI progress
Original: Measuring progress toward AGI: A cognitive framework View original →
Google DeepMind said on March 17, 2026 that it has published a new paper on how to measure progress toward AGI, arguing that current discussions often lack a durable empirical framework. Instead of claiming that AGI is near or that any one benchmark can settle the question, the paper proposes a cognitive-science approach for describing and comparing the capabilities of AI systems. DeepMind positions the work as an attempt to improve the measurement layer that sits between frontier-model marketing claims and any serious assessment of general intelligence.
The paper identifies 10 cognitive abilities that DeepMind argues are important for general intelligence: perception, generation, attention, learning, memory, reasoning, metacognition, executive functions, problem solving, and social cognition. It then proposes a three-stage evaluation protocol. First, researchers should test AI systems across a broad suite of tasks that cover each ability and use held-out sets to limit contamination. Second, they should gather human baselines for the same tasks from a demographically representative sample of adults. Third, they should map model performance against the distribution of human performance rather than treating raw scores in isolation.
To move the idea from theory to practice, DeepMind and Kaggle also launched a hackathon focused on five areas where the evaluation gap is largest: learning, metacognition, attention, executive functions, and social cognition. Participants can build benchmarks on Kaggle's Community Benchmarks platform and test them against a lineup of frontier models. Google says the competition carries a total prize pool of $200,000, with submissions open from March 17 through April 16 and results scheduled for June 1.
Why it matters
- Benchmark design increasingly shapes how labs, investors, and regulators interpret frontier-model progress.
- DeepMind is pushing for human-relative measurement rather than single-score leaderboard thinking.
- The Kaggle hackathon turns an abstract framework into a community effort to build reusable evaluations.
The announcement does not say AGI has been reached. Instead, it shows one major lab trying to standardize how progress claims should be evaluated before they harden into industry narrative. If the framework gains adoption, it could influence how future model releases are compared, how capability gaps are discussed, and how public arguments about AGI become more evidence-driven.
Related Articles
Google DeepMind는 2026년 3월 26일 대화형 AI가 감정을 악용하거나 사람을 해로운 선택으로 유도할 수 있는지를 다룬 새 연구를 공개했다. 회사는 영국·미국·인도 참가자 1만 명 이상이 참여한 9개 연구를 바탕으로, harmful AI manipulation을 측정하는 첫 empirically validated toolkit을 만들었다고 밝혔다.
Google DeepMind가 2026년 3월 17일 AGI 진전을 평가하기 위한 cognitive framework를 공개했다. benchmark leaderboard 대신 인간 인지 능력 분해와 capability profile 비교로 논의를 옮기려는 시도다.
Google DeepMind의 AI 수학 연구 에이전트 Aletheia가 FirstProof Challenge에서 전문가 심사단이 인정한 연구 수준 수학 문제 10개 중 6개를 자율적으로 해결했습니다. Gemini Deep Think 기반의 이 에이전트는 테렌스 타오 등 수학자들로부터 가치 있는 연구 협력자로 인정받고 있습니다.