Forecast Budgets with AI

Where AI can genuinely help budget forecasting and where finance teams still need disciplined modeling, assumptions, and human judgment.

March 16, 2026 · 7 min · Michelle
Cover image

One Forecast, Many Explanations: Why Time-Series Attribution Needs a Horizon Axis

TL;DR for operators A multi-step forecast may produce one trajectory, but the model does not necessarily use the same historical evidence for every point on that trajectory. In the real-world trained-model experiments studied here, an explanation assigned to its own forecast step outperformed explanations borrowed from other steps: the median own-versus-mismatched margin was +0.1418, and 80.6% of runs were positive. ...

September 28, 2026 · 7 min · Zelina
Cover image

When the Test Window Changes the Problem

TL;DR for operators A forecasting benchmark can change the problem being measured without changing the nominal dataset. In the rideshare data examined here, zeros make up 46.9% of the full dataset but only 5.3% of the standard rolling-origin evaluation windows. Under a series-wise split, the evaluation-window zero rate rises to 59.1%, and the interpretation of an autoregressive hurdle model reverses. ...

August 20, 2026 · 7 min · Zelina
Cover image

Important, but Not Direct: When Time-Series Attribution Misstates Model Dependencies

TL;DR for operators A forecasting dashboard can correctly report that an earlier observation influenced a prediction and still give the wrong impression about how that influence enters the model. Amadeo Tunyi’s paper, The Failures of Marginal Influence-Based Attribution Methods for Global Time Series Explanations, argues that familiar scalar attribution methods cannot in general recover the model’s direct temporal dependency structure.1 Marginal methods can assign importance to an earlier variable whose influence is entirely mediated through a later, autocorrelated observation. Gradient methods can report sensitivity that exists only outside the support of the data the model actually sees. ...

August 19, 2026 · 8 min · Zelina
Cover image

The Forecast Can Be Wrong and Still Save the Charge

TL;DR for operators EV charging optimization has a small, rude problem: the most important variable is often the one the operator does not know. A plugged-in car may leave in twenty minutes or three hours. That difference determines whether the controller can wait for cheap electricity or must charge immediately like an anxious intern with a deadline. ...

June 26, 2026 · 16 min · Zelina
Cover image

From Playbooks to Probabilities: When AI Starts Thinking Like a Football Manager

Football is usually explained after the fact. A team “pressed high.” A winger “found space.” A midfield line “lost compactness.” These statements may be accurate, but they arrive with the comforting uselessness of a weather report read after the picnic. The real managerial question is not merely what happened. It is what could have happened if the opponent shifted earlier, if the team protected the half-space, if the attacking line stretched the back four, or if the next pass invited three different futures instead of one. ...

April 14, 2026 · 17 min · Zelina
Cover image

Rationales Before Results: Teaching Multimodal LLMs to Actually Reason About Time Series

Dashboard work has a familiar little ritual. Someone opens a chart, zooms into the last few points, notices a dip, a rebound, or a suspiciously clean trend line, and then says something that sounds analytical: “Looks like it will continue.” Sometimes that is wisdom. Sometimes it is just a human staring confidently at a squiggle. ...

January 7, 2026 · 15 min · Zelina
Cover image

When Physics Remembers What Data Forgets

Data is expensive. Worse, in real scientific and industrial systems, the most useful data is often the data you do not have yet: the failure condition, the rare regime shift, the long-horizon trajectory, the sensor reading after something starts behaving strangely. This is why “just train a larger model” is not always an operating strategy. Sometimes it is only a procurement strategy wearing a lab coat. ...

December 27, 2025 · 12 min · Zelina
Cover image

Crystal Ball, Meet Cron Job: What FutureX Reveals About ‘Live’ Forecasting Agents

TL;DR for operators FutureX is less interesting as a leaderboard and more interesting as an operating model for evaluating AI agents that claim to forecast the future. The benchmark runs a live loop: collect future-facing questions from curated web sources, ask agents to predict before the answer exists, wait for resolution, crawl the answer, and score the prior prediction. That matters because most “forecasting” evaluations are either historical backtests with leakage risk or static datasets quietly ageing into trivia. ...

August 19, 2025 · 13 min · Zelina
Cover image

Forecast: Mostly Context with a Chance of Routing

TL;DR for operators Most forecasting teams already have decent numerical forecasters. Their problem is not that ARIMA, ETS, Lag-Llama, Chronos, or internal demand models suddenly forgot how Tuesdays work. The problem is that many important forecast shocks arrive as text: heat-wave notices, maintenance schedules, holiday effects, price caps, promotions, policy changes, store closures, one-off events, and all the other messy little business facts that refuse to fit politely into a clean covariate table. ...

August 16, 2025 · 17 min · Zelina