Bachelor thesis · PCA covariance forecasting

Completed

A reproducible Python pipeline from market data to forecasting and evaluation. I investigate how much a single PCA indicator can predict.

Updated
2026-09-18
pythonpandasnumpypcaarfimareproducibility

The question

How far can covariance forecasting be simplified? In my bachelor thesis, I investigated whether a single dominant PCA indicator can describe and forecast the evolution of a covariance matrix.

I processed one-minute prices for Bitcoin, Ether and BNB from 2024. Log returns are grouped into non-overlapping 30-minute covariance matrices. The PCA basis is estimated on the training period and remains fixed during evaluation.

What I built

I developed a Python pipeline connecting data acquisition, validation, modelling and evaluation.

  • Downloading monthly archives and checking checksums, timestamps, completeness and price values.
  • Computing returns and covariance matrices with pandas and NumPy.
  • Forecasting the dominant indicator with ARFIMA(0, d, 0) and reconstructing covariance matrices.
  • Comparing against two simple persistence forecasts in a chronological holdout.
  • Generating result figures and tables automatically.

Data processing, forecasting and evaluation live in separate modules. The analysis scripts and reproduction notebook use the same implementations.

Interpreting the results

Across 8,784 holdout covariance matrices, the ARFIMA-based approach achieved an aggregate RMSE 22.70% lower than direct covariance persistence. This baseline uses the most recently observed matrix as the next forecast.

The advantage was unevenly distributed over time. The approach had a smaller error than this baseline at only 35.3% of forecast origins. Reducing the representation to a single indicator also caused substantial approximation error. The result is therefore specific to the data and study design. It does not establish general superiority or better investment decisions.

Reproducibility

Dependencies are recorded in uv.lock. Unit tests cover forecasting weights, parameter estimation, covariance reconstruction and error aggregation, among other components. An executed Jupyter notebook walks through the analysis and reconciles its outputs with the reported metrics.

Beyond modelling, this project is about validating data and being precise about what an observed result actually supports.