AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning
A new benchmark, AhaBench, evaluates whether language agents truly learn from prior experience in long-horizon tasks, moving beyond single-prompt evaluations. It assesses agents' ability to improve behavior after receiving useful experience, even when direct support is re…
Intelligence analysis by Gemini 2.5 Flash

AhaBench is a novel benchmark designed to test the long-term learning capabilities of modern language agents, specifically their ability to leverage past experiences to improve future performance in complex, multi-step scenarios. Unlike traditional evaluations that reset agents or only score final states, AhaBench focuses on whether agents adapt and improve when explicit support is al…
Imagine you have a robot friend who needs to learn new games. Most tests just see if the robot can win one game. But AhaBench is like seeing if your robot friend learns from playing a game, so it can play a similar game better later, even if the rules are slightly changed or hidden. It checks if the robot truly understands and remembers, not just copies.
Analysis
Aha-Puzzle
The Aha-Puzzle component of AhaBench is designed to test an agent's ability to perform no-hint exploration after it has previously solved hidden-state puzzles with explicit guidance. The research indicates that while agents might show improved scores when visible support is present, this often does not translate into a genuine capacity for independent exploration in similar, but unsupported, scenarios. This suggests a fundamental challenge in how current language agents generalize learned strategies beyond the immediate context of their training or initial problem-solving. The benchmark aims to uncover whether agents truly internalize problem-solving methodologies or merely rely on pattern matching within given examples.
Aha-Euler
Aha-Euler focuses on evaluating the transfer of mathematical ideas, drawing inspiration from Project Euler problems. It presents agents with generated taught and held-out tasks, which are then assessed using exact validators. A key finding from this component is the stark difference in performance between agents receiving full teaching (achieving 78.6-100.0% success) versus those attempting answer-only transfer (ranging from 0.0 to 73.9%). This disparity underscores the difficulty for agents to abstract and apply complex mathematical principles without comprehensive instructional support, highlighting a critical area for improvement in their reasoning capabilities.
Aha-Vending
The Aha-Vending component, an open-source implementation inspired by Vending-Bench, simulates a vending agent's operation under real-world conditions. It specifically tests an agent's ability to maintain profitability while navigating delayed feedback and unexpected operational incidents. This task effectively distinguishes between agents that can adapt to dynamic situations and manage unforeseen problems to remain profitable, and those that succumb to bankruptcy or fail to process orders efficiently. It provides a practical measure of an agent's robustness and long-term decision-making skills in an environment with evolving challenges and consequences.
Key points
- AhaBench is a new benchmark for evaluating long-horizon continual learning in language agents.
- It assesses whether agents improve behavior from prior experience, even when explicit support is removed or delayed.
- The benchmark includes three components: Aha-Puzzle, Aha-Euler, and Aha-Vending.
- AhaBench uses a three-part scorecard: Initial Score, Post-Experience Score, and Learning Lift.
- Initial results show leading models like Claude Opus 4.6 and Gemini 3.1 Pro achieve high post-experience scores and learning lift, but struggle with knowledge transfer in specific scenarios.
The development of AhaBench could significantly accelerate the creation of more intelligent and adaptable AI agents capable of continuous learning and improvement over extended periods. By identifying specific learning gaps, researchers can develop models that genuinely leverage past experiences, leading to more robust and autonomous AI systems.
The benchmark's initial findings suggest that even leading models like Claude Opus 4.6 and Gemini 3.1 Pro struggle with transferring learned knowledge effectively, especially in no-hint exploration or answer-only transfer scenarios. This indicates that achieving true long-horizon continual learning remains a significant challenge, potentially slowing the deployment of highly adaptive agents in complex real-world applications.



