TL;DR for operators

A multi-step forecast may produce one trajectory, but the model does not necessarily use the same historical evidence for every point on that trajectory. In the real-world trained-model experiments studied here, an explanation assigned to its own forecast step outperformed explanations borrowed from other steps: the median own-versus-mismatched margin was +0.1418, and 80.6% of runs were positive.

That matters for model review. If an audit report attaches one importance map to a 24-step forecast, it may accurately summarize the forecast in aggregate while obscuring which historical observations mattered to step 3 versus step 20.

The proposed Horizon-Resolved eXplanation framework, or HRX, keeps that forecast-step dimension instead of averaging it away. The broader finding is stronger than one particular attribution method: seven alternative estimators also showed positive step-specific margins. Yet the per-step explanations are not entirely independent. Much of their structure can be represented by a small number of shared directions.

For operators, the result supports horizon-specific explanation and documentation where the attribution matrix shows meaningful variation. It does not establish causal drivers of the underlying time series, and the paper reports no demonstrated improvement in forecast accuracy from explanation-guided input selection or pruning.

One importance map can hide a step-specific mismatch

Consider a forecasting system that consumes the last 96 observations and predicts the next 24. An operator asks which historical values drove the forecast.

The usual answer is an importance vector over those 96 historical positions. That representation quietly assumes that a single ranking can stand in for all 24 future outputs.

The experiments in Explaining Time Series Forecasting with Horizon-Resolved Attribution challenge that assumption.1 Across real-world trained-backbone runs, the authors report a median $g_{\mathrm{margin}}$ of +0.1418, with 80.6% of runs positive. The metric compares how well a forecast step’s own attribution map explains that step against a map taken from another forecast step.

The interpretation is specific: the historical positions most relevant to one output are often not ranked the same way for another output. Averaging those rankings into one vector can therefore discard information before an auditor ever sees the explanation.

HRX changes the representation from one vector to a matrix:

$$ E\in\mathbb{R}^{H\times L}, $$

where $H$ is the number of forecast steps and $L$ is the lookback length. Row $h$ contains the importance scores for forecast step $h$.

This does not modify the forecasting model. HRX is a post-hoc explanation layer attached to a differentiable forecaster.

The shuffled row is what makes the horizon claim testable

Producing $H$ attribution maps is straightforward. Demonstrating that those maps actually belong to different horizons is harder.

The paper therefore evaluates three comparisons. First, $g_{\mathrm{own}}$ asks whether deleting inputs ranked by a step’s own attribution row changes that step more than deleting inputs ranked by the shared, horizon-averaged vector. Second, $g_{\mathrm{shuffled}}$ repeats the test with an attribution row borrowed from another forecast step. The difference,

$$ g_{\mathrm{margin}} = g_{\mathrm{own}}-g_{\mathrm{shuffled}}, $$

tests whether the row is specifically matched to its own forecast step.

That shuffled control is central. A concentrated attribution map could outperform an averaged map simply because averaging blurs strong features. But if a row performs better for its own output than for another output, the experiment is detecting horizon-specific alignment rather than merely sharper attribution.

The synthetic tests provide the cleanest validation because the generators contain known horizon-dependent structure. Across linear, CNN, and Transformer backbones, the horizon-resolved matrices outperform the shared vector on all four reported ground-truth attribution measures. The expected deletion pattern also appears: $g_{\mathrm{own}}>0$, $g_{\mathrm{shuffled}}<0$, and the margin is larger still.

Real data cannot provide true attribution labels, so the evidence there is necessarily faithfulness-based. Across 1,126 runs used for comparisons with prior explanation metrics, all four borrowed metrics are positive in aggregate. The proposed margin gives the clearer horizon-specific test.

The result survives changing the attribution estimator

One possible explanation for these findings would be that raw input gradients happen to generate horizon-varying maps.

The estimator comparison is therefore best read as a robustness test of the horizon axis, not as a contest to identify the best explanation algorithm.

The authors reconstruct the per-horizon explanation using seven alternative estimators, including saliency, gradient-times-input, integrated gradients, occlusion, Dynamask, ExtremalMask, and ContraLSP. Every estimator reports a positive confidence interval for $g_{\mathrm{margin}}$ across the tested real-world configurations.

That substantially changes the interpretation. HRX is not mainly a proposal for a new gradient technique. Its methodological claim is that whatever attribution estimator an organization already trusts, multi-output forecasting may require preserving the output axis rather than collapsing it.

There is a compute boundary. Gradient-based approaches are comparatively inexpensive, while optimization-based masking methods become hundreds of forward-pass equivalents because separate optimization is required across forecast steps. Explanation policy therefore has to consider both fidelity and the number of horizons being analyzed.

Per-step does not mean dozens of unrelated explanations

Keeping one map per forecast step might suggest that a 96-step forecast requires 96 unrelated explanations.

The rank analysis points to a more structured picture.

The authors compute an effective rank from the singular values of the attribution-magnitude matrix. Rank near one means the rows are nearly redundant; higher values imply that several distinct attribution directions are needed.

The observed relationship is systematic. Runs with effective rank between 1.0 and 1.1 have a reported $g_{\mathrm{own}}$ of +0.015. For rank 2.0 to 3.0, it rises to +0.250; above 3.0, to +0.323. Rank-truncated reconstructions reinforce the same point: a rank-1 approximation recovers none of the full horizon-resolution gain, rank 4 recovers about 51%, and rank 8 about 71%.

So the useful mental model is not “every horizon is independent.” It is closer to a small shared basis whose coefficients change as the forecast advances.

For model governance, that distinction matters. Effective rank can serve as a triage signal: forecasts close to rank one are candidates for a simpler shared explanation, while higher-rank cases deserve horizon-specific inspection. This is a Cognaptus inference from the paper’s measurement results, not a deployment rule validated by the study.

Governance gains are clearer than forecasting gains

The strongest practical use is diagnostic.

A model-review process can ask whether the explanation supplied for a long forecast actually remains representative across early, middle, and late horizons. Financial-model documentation could similarly separate the evidence used for near-term and longer-term outputs rather than attaching one attribution chart to the entire trajectory.

The paper also finds that horizon dependence is concentrated more strongly along the time axis than the channel axis in the tested multivariate forecasts. That directs attention toward which historical positions are being used, rather than assuming that horizon variation mainly comes from switching among input variables.

But HRX does not turn attribution into causal evidence. Its deletion test measures what the trained model’s prediction changes when selected inputs are perturbed. It does not establish that those observations causally drive the underlying physical, economic, or financial process.

Nor does better explanation faithfulness imply better forecasts. The authors report no demonstrated forecasting or diagnostic improvement from explanation-guided input selection or horizon-specific pruning. They therefore position HRX as a measurement tool.

The unresolved question is deployment selection. The authors could not predict the size of the horizon-resolution benefit on a new series from simple statistics such as spectral concentration, autocorrelation, target correlation, or naive-forecast difficulty. Effective rank is informative once explanations have been computed, but the paper does not yet supply a cheap pre-screening rule that tells an operator where HRX will matter most.

Preserve the output axis when the decision depends on it

The paper’s contribution is ultimately representational. A multi-step forecast contains multiple output decisions, yet explanation pipelines have often compressed those decisions into one attribution object.

The evidence here indicates that this compression frequently removes real model-use structure. The structure is not arbitrary, and it is not overwhelmingly high-dimensional. A modest number of shared attribution directions can explain much of the variation, but one shared direction is usually insufficient.

For teams using forecast explanations in validation, governance, audit, or risk communication, that makes the design question concrete: before accepting one importance map for an entire forecast, check whether the horizons actually share one.

Cognaptus: Automate the Present, Incubate the Future.


  1. Seunghan Lee and Jun Seo and Jaehoon Lee and Junhyeok Kang and Sangjun Han and Sungdong Yoo and Minjae Kim and Tae Yoon Lim and Dongwan Kang and Hwanil Choi and Soonyoung Lee and Wonbin Ahn (2026). Explaining Time Series Forecasting with Horizon-Resolved Attribution. arXiv:2609.12639. https://arxiv.org/abs/2609.12639 ↩︎