Why a Single Number Loses Important Information
A point forecast of a financial quantity – for example, a future return, volatility or portfolio loss – is usually perceived as “the most probable future”. But from the standpoint of forecasting theory this is wrong. A complete forecast of the quantity \(Y_{T+h}\) given the information set \(\mathcal F_T\) is the conditional distribution \(p(Y_{T+h}\mid\mathcal F_T)\). A single number is only a functional of this distribution: a mean, median, quantile or another summary. This does not make a point estimate useless: to choose an action that minimizes expected losses, one functional of the distribution is sometimes sufficient. But for decisions sensitive to the probability and scale of deviations, to tail losses or to several scenarios, a point forecast is usually insufficient.
Here it is important to separate the choice of a forecast from the assessment of its risk. At finite variance, for a numerical forecast \(a\) the following holds
Consequently, the conditional mean is sufficient to choose the optimal \(a\) under squared error. But assessing the expected size of that error also requires the conditional variance. In this article \(\mathcal F_T\) denotes the information at the time of the forecast, while \(F\) with subscripts denotes a distribution function.
Which particular functional is optimal depends on the loss function. Under squared error the conditional mean is optimal, under absolute error the conditional median, and under an asymmetric piecewise-linear loss function (quantile/pinball loss) the corresponding quantile (Forecasting: principles and practice; Regression Quantiles; Working Paper 11188). Therefore two forecasts with the same point value can imply completely different risks: one distribution may be narrow and nearly symmetric, the other – broad, asymmetric and heavy-tailed. For a financial decision this difference often matters more than the difference between central estimates.
One Mean, Two Distributions and Different Decisions
Let us set up a reproducible teaching example that is not tied to a real asset. Let \(Y=100\log(P_1/P_0)\) be a one-period log-return on the percentage scale. Let us compare two conditional forecasts: A – a normal distribution with mean 0 and standard deviation 1; B – a normal distribution with mean 0 and standard deviation 2. Both have a zero median. They differ only in scale: this is an example of higher tail risk at a fixed threshold, not an example of mathematically heavy tails.
Denote the standard normal distribution function by \(\Phi\) and its quantile function by \(\Phi^{-1}\). For scale \(s\in\{1,2\}\) the central interval with coverage \(c\) equals \([-s\Phi^{-1}((1+c)/2);s\Phi^{-1}((1+c)/2)]\), and the probability \(Y\le-3\) equals \(\Phi(-3/s)\). This is enough to reproduce the table; the values are rounded.
| Measure | A: standard deviation 1 | B: standard deviation 2 |
|---|---|---|
| Mean and median of the log-return | 0 | 0 |
| Central 80% interval for \(Y\) | \([-1{,}282;1{,}282]\) | \([-2{,}563;2{,}563]\) |
| Central 95% interval for \(Y\) | \([-1{,}960;1{,}960]\) | \([-3{,}920;3{,}920]\) |
| Probability \(Y\le-3\) | 0,135% | 6,681% |
| 95% quantile of the logarithmic loss measure \(-Y\) | 1,645 | 3,290 |
Now let us define a specific decision: under a conditional teaching rule a scenario passes the risk filter if the modeled probability of \(Y\le-3\) does not exceed 1%. A passes this filter, B does not. The 1% threshold is chosen purely for the example, is not a standard and gives no recommendation regarding a position. Two identical point forecasts of 0 lead to different results under the same rule, because the rule uses the probability of a deviation.
The logarithmic measure \(-Y\) is not equal to monetary loss. For a long position of initial value \(V\), with no intermediate payments, fees or changes in the position, the monetary loss equals \(V(1-e^{Y/100})\). Hence \(Y=-3\) corresponds to a loss of about 2,955% of initial value, not exactly 3%. In what follows, quantitative measures must be computed for that loss object which is actually used in the decision.
What Uncertainty an Interval Describes
In the example the mean is known by construction of the distributions. Even so, the future observation is random: even a perfectly known mean does not replace a prediction interval. In an applied model the mean and other parameters still have to be estimated from the data, and an additional source of uncertainty appears.
A confidence interval for a parameter or a conditional mean describes the precision of its estimation. A prediction interval relates to a future observation and also accounts for its own randomness. These are different objects, not two names for one range (NIST: forecast uncertainty).
In the frequentist formulation, the coverage level characterizes the procedure over repeated samples and forecasts under its assumptions. In the Bayesian formulation, an interval from the posterior predictive distribution expresses the conditional probability of the future outcome within the chosen model. Neither interpretation makes a wrong model correct: nominal 95% may differ from the actual share of coverages.
An interval at one horizon does not describe the entire future trajectory. For example, the question “what will the price be in a month?” differs from “will the price cross a barrier at least once during the month?”. The second requires joint dynamics: a set of separate 95% intervals by day is not a 95% region for the whole trajectory. Therefore one first defines the event on which the decision depends, and only then chooses the form of the forecast.
Where Uncertainty Comes From
A forecast has at least four distinct sources of uncertainty: future random disturbances; estimation of parameters from a finite sample; choice of model structure; a possible change of regime. The first is present even with known parameters. The second can be partly accounted for through the estimation procedure. The third requires a comparison of substantively different specifications. The fourth cannot be automatically eliminated by more precise estimation of the parameters of the previous regime.
These distinctions are practically important. If a prediction interval is obtained by plugging in a single parameter estimate, it may not reflect the uncertainty of their estimation. If several nearly identical models are compared, their agreement does not mean that model-structure uncertainty has been accounted for. Finally, a wide interval from an old model does not necessarily describe the consequences of an event that its mechanism does not envisage at all.
Randomness of the future under a given model is often called aleatoric uncertainty, while ignorance of parameters and structure is called epistemic. The boundary between them depends on the formulation. For a report, it is more useful to list exactly what is included in the interval calculation than to stop at these two labels.
Financial Specifics: Volatility, Heavy Tails and Asymmetry
Financial returns often display properties that make a symmetric normal interval of the form \(\hat y \pm 1{,}96\sigma\) a risky simplification.
First, volatility is usually not constant: periods of large movements often follow large movements, and quiet periods – quiet ones. This is called volatility clustering. ARCH- and GARCH-class models allow the conditional variance to depend on recent errors and past variance values, so a prediction interval can widen after periods of high volatility and narrow in quiet periods (Engle, 1982; Bollerslev, 1986).
Second, return distributions often have heavy tails. Extreme movements occur more often than a normal distribution with the same variance predicts. Heavy tails often persist even after filtering volatility with a GARCH model, so the choice of the innovation distribution requires empirical checking (Cont, 2001).
Third, asymmetry may be present, but it cannot be treated as a universal property of any asset, period and data frequency. Asymmetry should be checked separately for the particular object and model.
It is also important to record exactly what is being forecast: price, log-return, volatility, cash flow or portfolio loss. Intervals and loss functions for these objects differ. For example, transforming a log-return forecast back into a price level can create asymmetry and bias; one cannot simply exponentiate the point mean log-return without explanation.
How to Build a Forecast: One Route and Its Extensions
Let us start with a one-period model \(Y_{T+1}=\mu_T+\sigma_T Z_{T+1}\), where \(\mu_T\) is the conditional center, \(\sigma_T>0\) is the conditional scale, and the distribution of the standardized disturbance \(Z_{T+1}\) is given by the model. With fixed estimates of center and scale, the quantile equals \(\mu_T+\sigma_T q_p(Z)\), and the probability of the event \(Y_{T+1}\le b\) is \(F_Z((b-\mu_T)/\sigma_T)\).
The A/B table is a special case with \(\mu_T=0\), a standard normal \(Z\) and two scale values. In a working forecast, center and scale are estimated only from the information available at time \(T\). Then the required probability or quantile is obtained from the chosen distribution, the forecast is preserved until the outcome arrives and is compared with the observation. It is this cycle, and not the interval formula itself, that makes it possible to check the model's adequacy.
The normal distribution is convenient for the calculation here, but it is not obliged to fit the data. The distribution of standardized residuals is checked separately: after accounting for changing volatility, asymmetry, heavy tails or dependence may still remain in them. Replacing the innovation distribution changes quantiles and probabilities even with the same center and scale.
The further method depends on which part of the forecast is needed and which assumptions are insufficient:
- Quantile model directly estimates the required conditional quantile. A single quantile does not determine the full distribution and does not provide arbitrary event probabilities.
- Simulation propagates the chosen disturbances through the model. At several horizons it can produce joint trajectories if their dependence is reproduced; independent draws from separate marginal distributions do not replace this.
- Bootstrap uses repeated resampling, and its scheme must preserve the relevant dependence. Simple shuffling of raw residuals with changing variance can destroy the needed structure; re-estimating parameters on resamples allows part of the estimation uncertainty to be included (Bootstrapping and bagging).
- Bayesian approach averages the forecast over the posterior distribution of parameters instead of substituting a single estimate:
Here \(M\) is a fixed model and \(\theta\) its parameters. The averaging accounts for their uncertainty within \(M\), but does not prove the correctness of the structure itself (Posterior Predictive Checks).
Comparing or combining several models helps to examine sensitivity to specification. However, the spread of their point forecasts cannot be presented as a prediction interval without additional justification: it may not include the randomness of the future outcome. For unanticipated regimes, stress scenarios are considered separately, and no precisely measured probability is assigned to them without grounds.
From Distribution to Financial Decision
The A/B example has already shown the transition from model to decision rule. It can be extended to monetary units: with an initial position value of 100 monetary units, the 95% quantile of monetary loss equals \(100(1-e^{-s\Phi^{-1}(0{,}95)/100})\). For A this is about 1,631, for B – 3,236 monetary units. The formula uses a monotone transformation of the log-return, not the approximation “loss equals minus return”.
The teaching decision above used threshold-crossing probability. If the decision instead limits the size of losses, a quantile or the average size of the tail loss is required. Denote by \(L\) the chosen loss object: for example, the monetary loss of the position, rather than automatically \(-Y\). Then
where \(F^{-1}\) is the quantile of the conditional loss distribution and \(\alpha\) is the quantile level; the upper tail share is \(1-\alpha\). For continuous \(L_{t+h}\) this means
Realized coverage on historical forecasts only approximates this nominal level and does not prove it. VaR is a loss threshold that under the model is not exceeded with probability \(\alpha\). It does not say how large the losses are if the threshold is nevertheless exceeded.
For continuous distributions, Expected Shortfall at level \(\alpha\) can be written as the mean of the tail quantiles:
Under continuity this coincides with the conditional expected loss above the VaR threshold:
In the presence of jumps, probability atoms or in the discrete case, the integral definition via quantiles specifies a more careful tail-average convention. ES additionally describes the average size of losses in the tail and is useful when the decision is sensitive to the severity of extreme losses rather than only to the fact of crossing the threshold (Acerbi and Tasche, On the coherence of Expected Shortfall). In the internal models standard for market risk, one-sided Expected Shortfall at level 97,5% is used; this is an example of a practical need to account for the tail, not a universal definition of a risk measure (Basel Framework, MAR33).
If a return or P&L is used instead of a loss, the sign convention and the transformation into a loss must be stated explicitly. Otherwise it is easy to confuse the direction of the tail: for a positive loss quantity, risk is described by the right tail of the loss distribution, not the left one.
In addition to VaR and ES, the predictive distribution can provide:
- the probability of exceeding a loss threshold;
- probabilities of scenario events;
- intervals for several confidence levels;
- multi-period scenario trajectories or simultaneous prediction regions with a separately specified level, if the decision depends on the path rather than on a single horizon; a set of marginal intervals at each horizon is not sufficient for this.
How to Check Probabilistic Forecasts
The quality of a predictive distribution cannot be assessed with a single point-error metric. The check must be out-of-sample and sequential.
Rolling forecasting origin
Forecasts should be built on successive out-of-sample windows: each forecast is made only from the data known at the time of the forecast (Time series cross-validation). This imitates real model use and helps avoid an inflated estimate of quality due to fitting the whole sample. At each forecast origin, the operational protocol must be reproduced: if in the real system parameters are re-estimated daily or at every new window, this must also happen in the backtest; if the model is re-estimated less often, the backtest must preserve that schedule. The window may be expanding or of fixed length, and the choice of window and re-estimation schedule must be fixed in the verification specification.
Interval Coverage
For interval forecasts, the share of actual values falling into forecast intervals is evaluated. But the share alone is not enough. Both the overall hit frequency (unconditional coverage) and the dependence of violations over time are checked. Conditional coverage requires both properties to hold jointly: the correct unconditional coverage probability and the absence of dependence in the sequence of exceedances within the chosen specification of the test (Evaluating Interval Forecasts). A classical test may check, for example, a specific first-order Markov dependence rather than any dependence on all available information. This is only an operational check relative to the chosen information set and test; the absence of obvious clusters does not by itself prove correctness. A formally correct average hit rate can hide systematic underestimation of risk in periods of high volatility if the violations are dependent.
Calibration and sharpness
The quality of a probabilistic forecast has two sides (Probabilistic Forecasts, Calibration and Sharpness).
Calibration – statistical agreement of forecasts with observed outcomes; the exact definition depends on the type of forecast. For a binary event the intuitive example is this: among cases assigned a probability of about 20%, the event should occur in roughly 20% of cases. The overall hit share and the behavior in individual regimes answer different questions: correct average coverage can be combined with systematic errors in periods of high volatility.
Sharpness, or concentration – the concentration of the predictive distribution. Narrow intervals and concentrated distributions are more informative, but only if calibration is preserved.
The goal is maximum sharpness subject to calibration. Therefore a wider interval is not automatically more honest, and a narrower one is not automatically better. Comparison is admissible only with account of actual calibration and appropriate evaluation rules.
Probability integral transform
For a full predictive distribution the standard diagnostic uses the probability integral transform (Evaluating Density Forecasts). For a continuous predictive CDF with a correctly specified true conditional CDF, the one-period PIT value can be written as \(U_{t+1}=F_{t+1\mid t}(Y_{t+1}\mid\mathcal F_t)\); it should then have the distribution \(\operatorname{Uniform}(0,1)\), and under standard conditions the sequence should be independent. With an estimated CDF, uniformity and independence are only checked approximately and depend on model estimation and the information set. Non-uniformity may indicate an error of location, scale or distribution shape, but the finite sample and estimation also affect the picture; a single histogram cannot unambiguously establish the cause. For multi-horizon overlapping forecasts, uniformity remains the target property under a correct model, but independence is usually violated: tests and standard errors that account for dependence are needed, rather than an automatic conclusion about missed dynamics. For discrete or mixed distributions, PIT requires randomization or another special modification.
Metrics: Point Forecast versus Probabilistic Forecast
RMSE and MAE evaluate only the chosen point functional. They are useful, but they cannot compare the quality of the entire predictive distribution. Two models can have the same RMSE of the point forecast yet differ substantially in interval calibration and tail risk.
For probabilistic forecasts, proper scoring rules are used – evaluation rules that elicit the true predictive distribution (Strictly Proper Scoring Rules). For the whole distribution, the following are suitable:
- log score – the logarithm of the forecast density at the realized value for a continuous outcome (for a discrete outcome – the logarithm of the probability assigned to it). Under this convention a larger value is better; for the negative log score as a loss function – smaller;
- CRPS – continuous ranked probability score, comparing the forecast distribution function with the actual outcome: \[\operatorname{CRPS}(F,y)=\int_{-\infty}^{\infty}\left(F(x)-\mathbf{1}\{y\le x\}\right)^2dx.\] Here CRPS is treated as a loss: smaller is better. The literature also contains sign conventions for score that are oriented toward maximization, so the direction of “better” must be stated explicitly;
For individual characteristics of the distribution, other consistent criteria are used:
- quantile loss, or pinball loss, for evaluating individual quantiles;
- interval score, which jointly accounts for interval coverage and width.
Intervals cannot be evaluated by width alone or by coverage alone. A narrow interval with poor coverage is bad; a wide interval with good coverage may be uninformative. A joint criterion is needed. When comparing average scores, sampling uncertainty and temporal dependence of errors should also be taken into account: a small difference in average values does not by itself prove the superiority of one model.
Limitations and Typical Errors
All probabilities and intervals are conditional on the data, model, estimation method and assumptions about the future regime. They are not an unconditional guarantee of the result.
Several important limitations:
- A 95% prediction interval does not mean that a particular realized interval “must” contain the future with an objective probability of 95%. This is nominal model coverage, which is checked on repeated forecasts.
- A historical backtest with a small number of crisis observations gives especially uncertain estimates of coverage for 99% intervals, VaR and Expected Shortfall.
- ARCH/GARCH models do not guarantee an accurate description of tails and may fail under structural breaks.
- “Stylized facts” of financial returns are recurring empirical patterns, not laws operating identically for any asset, period and data frequency.
- Normality, stationarity, independence of residuals and immutability of the regime are not facts but testable working assumptions.
- Uncertainty often grows with the horizon, but the specific dynamics depend on the model and the forecast object. In stationary or mean-reverting processes, forecast variance may stabilize.
- An unanticipated structural break is hard to represent within the predictive distribution of a fixed model. Known regime-switching mechanisms can be modeled probabilistically through regime-switching or change-point models, but this does not eliminate uncertainty about the appearance of a fundamentally new regime.
Practical Template for a Forecast Report
If the decision depends on the probability and scale of deviations, the central estimate must be supplemented with the relevant characteristics of uncertainty. A reasonable report on a probabilistic forecast may include the following elements.
- Forecast object. State clearly what is being forecast: price, log-return, volatility, cash flow or a positive loss quantity.
- Central estimate. State which functional is used: mean, median or another quantile, and why it matches the loss function.
- Several quantiles or intervals. For example, 50%, 80% and 95% intervals, or a set of quantiles showing the shape of the distribution.
- Tail measures. VaR and Expected Shortfall, if the task concerns loss risk, with explicit indication of level and sign convention.
- Model assumptions. Which model was used, which distributions and assumptions about the regime were adopted.
- Calibration backtest. Results of out-of-sample verification: interval coverage, independence of violations, PIT diagnostics or other appropriate checks.
- Quality metrics. Not only RMSE or MAE, but also proper scoring rules if the distribution is evaluated.
- Stress scenarios. Separate scenarios for events that are poorly described by the historical model or have few observations in the tails.
Conclusion
In the A and B example the mean forecast of the log-return is the same, but the probability of crossing one threshold differs, so one rule gives different decisions. This is the reason to move from point to distribution: not for the sake of more indicators, but for the information that the particular decision uses.
The practical sequence is this: define the object and the decision rule, construct the required probabilities or quantiles, state the assumptions explicitly, and check the forecasts against subsequent observations. The mean may be sufficient for choosing a point forecast under squared error; it does not describe the risk of that forecast. An interval, VaR or ES add the necessary information, but its quality remains a property of the model being tested, not a guarantee of a future result.