The record's summary of the idea: intelligence is how well you handle a problem you have never seen, not how many problems you have seen. ARC puzzles are new by construction. A model cannot have read the answer on the internet, which is exactly why the labs found it annoying.
ARC-AGI
A set of little grid puzzles, colour the cells according to a rule shown in a few examples, that humans find easy and models found nearly impossible for five years. Chollet built it in 2019 to measure skill acquisition rather than memorised skill. A million-dollar prize in 2024 drew attention; OpenAI's o3 scored high on it that December at great expense; a harder version followed in March 2025.
In the blog
1 article mention this recordTestimony
4 entries · newest firstSighted in December 2024 when o3 was reported at 87.5 per cent on the public set, at a compute cost per puzzle that the organisers estimated in the thousands of dollars. The record notes that a human solves one for the price of a cup of tea, and that both figures were headlined.
ARC-AGI is not a test for AGI, despite the name Chollet chose. He has said passing it is necessary, not sufficient. The version two puzzles, released in March 2025, reset the leading scores to single digits, which is the record's evidence that he meant it.
I remember doing a page of ARC puzzles in 2020 to see what the fuss was, finding them like a children's activity book, and then reading that GPT-3 scored near zero. It was the first benchmark I understood in my hands rather than in a table. I still do one occasionally, to check I am not a model.
Add to the record
What does it mean? Write it the way you would say it out loud.