← All projects

Forecasting Realized Volatility with HAR Models: In-Sample Estimation, Out-of-Sample Forecasting & VaR Backtesting

ACFI233 - Econometrics for Finance II

University of Liverpool Management School · 2025

Overview

An advanced econometrics project examining the accuracy of the Heterogeneous Autoregressive (HAR) family of realized volatility models and their downstream application to Value-at-Risk (VaR) risk management, using two decades of daily realized measures and returns for the SPDR S&P 500 ETF (SPY).

What I did

  • Extracted a defined sample period of daily realized variance, realized quarticity, signed semivariances, VIX², and returns for SPY, and structured the dataset for modelling in Python
  • Tested whether VIX² was an unbiased estimator of realized variance via linear regression, including individual and joint hypothesis tests on the intercept and slope at 5% and 1% significance, repeated under square-root and logarithmic transformations
  • Computed summary statistics and tested the normality of the realized variance and VIX² series (and their transformations) using the Jarque-Bera test
  • Estimated four HAR-family in-sample models (HAR, SHAR, HAR-Q, HAR-IV) across raw, square-root, and logarithmic functional forms, reporting Newey-West robust standard errors (10 lags) and adjusted R² for each
  • Interpreted the statistical significance, sign, and explanatory power of each coefficient, and evaluated the trade-offs between the different functional-form transformations
  • Selected the two best-performing in-sample models and generated out-of-sample forecasts using a 1,500-day rolling window, evaluated via the Mincer-Zarnowitz test
  • Constructed a model-averaging forecast combining the selected models and benchmarked it against the individual models using MSE and QLIKE loss functions, explaining why MAE and HMSE were unsuitable metrics for this setting
  • Tested whether model averaging significantly outperformed the individual forecasts using the Diebold-Mariano test
  • Built Value-at-Risk measures at the 1% and 5% significance levels from both the individual and model-averaged volatility forecasts, and plotted returns against each VaR series
  • Backtested the VaR estimates using the Proportion of Failures (PoF) and Traffic Light tests to assess whether violation rates matched theoretical expectations
  • Delivered a 10-minute individual presentation of the methodology and findings alongside the written report and Python code

Skills & tools

PythonTime-series econometricsHAR volatility modelsRegression & hypothesis testing (Newey-West SEs, Jarque-Bera)Rolling-window forecastingModel averagingForecast evaluation (Mincer-Zarnowitz, Diebold-Mariano, MSE, QLIKE)Value-at-Risk modelling & backtesting (PoF, Traffic Light)Academic report writing & presenting

© 2026 Zach Stow