Algometrics: Forecasting Under Algorithmic Feedback
A paper argues that forecasts can change the markets they predict, making passive test scores misleading for deployed models.
Intelligence analysis by GPT-5.4 Mini

The paper proposes algometrics, a framework for time series where model outputs feed back into the system they are meant to forecast. It says historical accuracy alone can hide deployment risk, especially when many similar algorithms crowd into the same market.
A weather app that tells people to carry umbrellas could change how many people go outside, and that can change what happens next. This paper says some prediction systems work like that in markets: their guesses can change the market itself.
That means a model can look smart on old records, but act differently once it is actually used. It is like judging a chess player by puzzles, then finding out the puzzles changed because the player’s moves changed the board.
The paper’s big idea is that people should test these systems in a way that checks for feedback, not just past accuracy. Otherwise, the score can be too nice and miss the real risk.
Analysis
What the paper says
Marc Schmitt introduces algometrics, a framework for settings where predictive models are not just observers of a time series but part of the mechanism that generates it. The paper focuses on algorithmic markets, where forecasts are turned into trades, allocations, execution schedules, or risk controls. Once that happens, the act of predicting can alter the future data the model will later be judged against.
The central distinction is between historical risk and deployment risk. Historical risk is what a model looks like under passive evaluation on past data, where predictions do not affect the system. Deployment risk is the error the same forecaster faces after its outputs influence actions in the live environment.
Main results
The paper claims three main results. First, deployment risk cannot be identified from passive historical data alone. In a one-step linear feedback model, many different algorithm-mediated environments can generate the same historical data while implying different deployment outcomes for the same predictor.
Second, model rankings can flip when crowding appears. A predictor with lower passive error can end up with higher deployment error once enough similar algorithms are adopted and their combined actions reshape the market.
Third, the paper says randomized or instrumented actions can identify short-horizon linear feedback, and it derives a finite-sample bound for estimating deployment risk.
Implication
The practical message is straightforward: time-series benchmarks in algorithmic markets should report not only predictive accuracy, but also how sensitive a model is to feedback from its own use. In other words, a strong backtest may still be misleading if the model helps create the very data it is tested on.
Key points
- The paper introduces algometrics for forecasting systems that affect the data they predict.
- It separates passive historical risk from deployment risk in live use.
- It argues deployment risk cannot be recovered from historical data alone.
- It says crowding by similar algorithms can invert model rankings.
- It claims randomized or instrumented actions can identify short-horizon feedback.
- It recommends reporting feedback sensitivity alongside predictive accuracy.



