AI models flub these intelligence tests. Can you fare any better?
AI models rapidly improve on some puzzles but struggle with spatial reasoning, memory adaptability, and abstract visual tests, highlighting fundamental differences from human cognition.
Intelligence analysis by Gemini 2.5 Flash

The article explores how puzzles serve as benchmarks for AI intelligence, demonstrating both rapid advancements and persistent weaknesses. Despite progress in areas like Connections puzzles, current models falter on tasks requiring human-like spatial manipulation, nuanced memory recall, and abstract visual inference, offering insights into their cognitive limitations.
Imagine a super-smart robot that can learn tons of facts and solve many word puzzles super fast. But if you ask it to imagine turning a toy block in its head, or to spot a tricky detail in a picture that looks almost like one it's seen before, it gets confused. Humans are still much better at these kinds of 'thinking with your eyes' or 'spotting the trick' games, showing that AI still thinks differently than we do.
Analysis
Puzzles have historically been a cornerstone of AI development, serving as critical benchmarks for measuring progress and identifying areas of weakness. From Arthur Samuel's checkers algorithm in 1959 to modern challenges like chess and Go, games provide a structured environment to test machine intelligence. While AI has shown remarkable improvement in certain domains, such as solving New York Times Connections puzzles, where models advanced from 18% accuracy in late 2024 to near perfection by early 2025, these tests also reveal significant gaps in AI's cognitive abilities compared to humans.
Spatial Reasoning
One of the most pronounced areas where AI models consistently underperform humans is spatial reasoning. Tasks like mental rotation problems, common in IQ tests, require the ability to mentally manipulate 3D objects from different angles. Despite advancements in visual input analysis for language models, they still fail abysmally at these challenges. This deficiency suggests that even with sophisticated 'world models' designed to help AI understand physical environments, current large language models (LLMs) lack the intuitive spatial manipulation capabilities that come naturally to human spatial thinkers like architects and mechanical engineers. This highlights a fundamental difference in how machines process and understand physical space.
Knights and Knaves
Another revealing category of puzzles involves memory and adaptability, exemplified by 'Knights and Knaves' problems. These logic puzzles require discerning truth-tellers from liars based on their statements. A 2024 study by researchers from Google and the University of Illinois Urbana-Champaign found that when models encountered slight variations of puzzles they had seen during training, they often failed to spot key differences, instead defaulting to memorized responses. This tendency to rely on rote memorization rather than adaptive reasoning can be a liability, causing models to 'whiz by' critical nuances. This issue also surfaces in 'SimpleBench' problems, which resemble complex problems seen during training but contain subtle tricks that humans easily identify, while top-tier AI models frequently miss them.
ARC-AGI
The Abstract and Visual Reasoning Challenge (ARC-AGI) benchmark further exposes AI's struggles with abstract and visual reasoning, even in two dimensions. These puzzles demand that models infer abstract, general rules from a set of examples. While models perform better when grids are encoded as numerical strings rather than images, research indicates that even correct answers are often achieved through 'byzantine and non-generalizable rules,' rather than the simple visual concepts humans employ. This suggests that AI's problem-solving methods in these complex visual and abstract tasks are fundamentally different from human intuition, often lacking the generalizability and conceptual understanding that define human intelligence.
Key points
- Puzzles have been crucial for AI development, showcasing both rapid advancements and persistent limitations.
- AI models significantly struggle with spatial reasoning tasks, such as mental rotation, unlike humans.
- Models can be tripped up by subtle variations in puzzles, relying on memorized solutions rather than adaptive reasoning.
- Abstract and visual reasoning challenges like ARC-AGI reveal that AI often uses non-generalizable rules, differing from human intuition.
- These tests highlight fundamental differences between human and machine cognition, offering insights into AI's strengths and weaknesses.
These identified weaknesses provide clear targets for AI researchers, enabling them to develop new architectures and training methods that could eventually bridge the gap in spatial, adaptive, and abstract reasoning, leading to more versatile and robust AI systems. Continued testing with such puzzles will drive innovation towards more human-like intelligence.
The persistent struggles of even advanced AI models with fundamental human cognitive tasks like spatial reasoning and adaptive problem-solving suggest that true general artificial intelligence remains a distant goal. These limitations could hinder AI's ability to operate effectively in complex, unpredictable real-world environments that demand flexible, intuitive understanding.


