Value at Risk and backtesting
One-day VaR and expected shortfall by three methods, and an out-of-sample test of whether they held up.
What it does
Value at Risk (VaR) at 95% is the one-day loss that is exceeded on only 5% of days. Expected shortfall (ES) is the average loss on those 5% of days. The Risk section estimates both three ways and then checks each method against what actually happened.
Why it is used
VaR is the most widely used single number for short-term downside risk, and also one of the most criticized. Showing three methods side by side, with a backtest, turns it from a single confident figure into a comparison of assumptions.
Inputs
- Daily portfolio returns (ten years by default).
- For filtered historical simulation, a GARCH(1,1) fit to those returns.
Formulas
Assumptions
- Historical simulation assumes the past window's distribution of daily returns applies tomorrow, with every day weighted equally.
- The parametric method assumes normally distributed returns with zero mean.
- Filtered historical simulation assumes the GARCH model captures how volatility changes, and that the standardized residuals are drawn from a stable distribution whose shape the history reveals.
- Holdings are held for one day at their current weights.
How to read the results
The headline uses filtered historical simulation because it reacts to current volatility while keeping the fat tails of real returns. When the three methods disagree, that disagreement is informative: a high parametric figure relative to historical usually means recent volatility is high; a high historical figure relative to parametric usually means fat tails. In the backtest, a 95% VaR should be breached on about 5% of days, and breaches should not arrive in clusters.
Limitations
- VaR says nothing about how large losses are beyond the threshold. Expected shortfall is shown for that reason.
- One-day figures do not scale simply to longer horizons when volatility clusters.
- The backtest covers about three years, so it contains roughly 38 expected breaches at 95% and fewer than 8 at 99%: enough to catch badly wrong models, not subtle ones.
Where it can fail
- All three methods fail when tomorrow is worse than anything in the estimation window. In February 2020, a VaR estimated on the calm prior year was badly exceeded.
- The parametric method understates tail losses for assets with fat tails or skew.
- Historical simulation is slow to react: a volatile period enters and leaves the window abruptly.
Changes from the original version
DeanOS began as a personal tool. Rebuilding it for the public meant rechecking each model; these are the changes that came out of that.
- The original version's third method drew random numbers from a normal distribution with the sample mean and volatility, which reproduced the parametric result with simulation noise. It is replaced by filtered historical simulation.
- The backtest (Kupiec and Christoffersen tests) is new.
Validation on current data
Out-of-sample 95% VaR backtests for the example portfolios over the last three years of the current snapshot. Each forecast uses only data available before that day; GARCH is refitted every 21 trading days on a trailing 1,000-day window.
References
- Jorion, P. (2007). Value at Risk, 3rd ed. McGraw-Hill.
- Barone-Adesi, G., Giannopoulos, K. and Vosper, L. (1999). VaR without correlations for portfolios of derivative securities. Journal of Futures Markets 19(5).
- Kupiec, P. (1995). Techniques for verifying the accuracy of risk measurement models. Journal of Derivatives 3(2).
- Christoffersen, P. (1998). Evaluating interval forecasts. International Economic Review 39(4).
- Artzner, P., Delbaen, F., Eber, J.-M. and Heath, D. (1999). Coherent measures of risk. Mathematical Finance 9(3).