BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents
The paper introduces BudgetBench, a new protocol and reference harness for evaluating memory strategies in local large language model (LLM) agents by treating the per-call input-token budget as the independent variable. It aims to expose performance nuances that fixed-bud…
Intelligence analysis by Gemini 2.5 Flash

BudgetBench addresses the scarcity of active context in local LLM agents by proposing a novel evaluation protocol. Instead of fixed memory limits, it systematically tests memory strategies across varying input-token budgets, ranging from 2K to 32K tokens. This approach reveals how different strategies perform under diverse resource constraints, highlighting issues like budget complian…
Imagine you have a super-smart robot helper, but it can only remember a certain amount of information at a time, like a small backpack. This paper introduces a new game called BudgetBench to test how well different robots use their backpack space. Instead of giving them the same size backpack every time, they give them different sizes – sometimes small, sometimes big. They watch if the robot can finish its tasks, how much space it uses, how fast it works, and if it ever tries to put too much in its backpack. This helps scientists figure out the best way for robots to remember things so they can be super helpful without getting confused or running out of space.
Analysis
The article introduces BudgetBench, a novel protocol and reference harness designed to rigorously evaluate memory strategies for local large language model (LLM) agents. The core problem it addresses is the inherent scarcity of "active context" in these local deployments, where factors like memory capacity, prefill latency, cache growth, and service objectives impose strict limits on the number of input tokens an LLM call can process. Traditional evaluation methods often use a fixed context window, which fails to capture the dynamic and budget-constrained realities of real-world local LLM operations.
BudgetBench Protocol
BudgetBench distinguishes itself by treating the per-call input-token budget as the independent variable in its evaluation framework. This means that instead of a single, static budget, the protocol systematically sweeps through a range of budgets—specifically 2K, 4K, 8K, 16K, and 32K tokens. By holding other variables constant, such as the model, task, sampler, and decoding method, BudgetBench can isolate the impact of different memory strategies under varying resource constraints.
The key outcomes measured include quality, budget utilization, latency, and critically, budget-violation rates. This multi-faceted approach provides a more comprehensive understanding of a memory strategy's performance profile, revealing how it adapts or fails when faced with different computational envelopes.
Pilot Studies
To substantiate its protocol, the paper presents several pilot studies rather than definitive rankings of memory strategies. One significant pilot involved a local qwen2.5:1.5b model, tested across 89 items each from the SWE-bench Verified and LongBench v2 datasets. Another study replicated a 50-item Qwen3 30B-A3B LongBench evaluation on a hosted environment, ensuring exact tokenization.
A third, larger oracle study used 500 items from LongMemEval, scored by the official GPT-4o evaluator. These pilot runs were instrumental in exposing critical issues that single-budget evaluations typically obscure, such as budget-compliance failures, unexpected non-monotonic quality curves, and specific operating points where strategies perform optimally or poorly. The authors transparently report on early pilot limitations, including tokenizer approximation issues that led to undercounted prompts and thus diagnostic-only violation rates, emphasizing the protocol's focus on robust failure reporting.
Reusable Contribution
The primary contribution of BudgetBench is its reusable measurement surface, which includes a swappable MemoryStrategy contract, explicit budget enforcement mechanisms, deterministic or versioned graders, and prompt-audit metadata. All these components, along with reproducibility artifacts, are open-sourced, providing a standardized toolkit for the research community. This framework allows other researchers and developers to adopt a consistent methodology for evaluating and comparing memory strategies, fostering more reliable and comparable results across different studies. The emphasis on transparent reporting, including the acknowledgment of early pilot limitations and the diagnostic nature of certain metrics, underscores a commitment to scientific rigor and reproducibility in the rapidly evolving field of local LLM agent development.
Key points
- BudgetBench is a new protocol for evaluating memory strategies in local LLM agents.
- It uses a budget-tiered approach, varying input-token budgets (2K, 4K, 8K, 16K, 32K).
- The protocol measures quality, budget utilization, latency, and budget-violation rates.
- Pilot studies revealed budget-compliance failures and non-monotonic quality curves that fixed-budget evaluations miss.
- The core contribution is a reusable measurement surface, including a swappable MemoryStrategy contract and reproducibility artifacts.
The BudgetBench protocol offers a standardized and robust method for evaluating memory strategies, which could lead to significant advancements in the efficiency and reliability of local LLM agents. By identifying optimal memory management techniques across various budget constraints, developers can create more powerful and resource-aware AI applications, making advanced LLMs more accessible and practical for edge devices and constrained environments.
The paper notes that the "budgeted-versus-full-context direction remains unresolved," with pilot studies showing mixed results (local slice near-null, hosted replication favoring full context). This suggests that while the evaluation protocol is valuable, finding universally superior memory strategies for budget-constrained local LLMs might be more complex than anticipated, potentially leading to trade-offs where efficiency gains come at the cost of quality in certain scenarios.



