Repository for the Uncertainty in Machine Learning course (WBAI054-05) at the University of Groningen.
We forecast next-step S&P 500 (SPY) returns and quantify the predictive uncertainty around those forecasts. Two models are compared:
- Variational GP (
variational_gp/vgp): Stochastic variational Gaussian Process that produces calibrated predictive distributions. - LSTM (
lstm): Recurrent baseline with MC Dropout for uncertainty estimation at inference.
Both share a common training, evaluation, and diagnostic-plotting pipeline, and are scored with point metrics (RMSE, directional Hit Rate) and probabilistic metrics (NLL, ECE).
- Maurice Meijer (S5480604)
- Marcus Harald Olof Persson (S5343798)
We use uv for project management.
- Clone the project.
- Synchronise the project.
uv sync- Create a copy of example.config.yaml and rename it to
config.yaml. Update the configuration, if desired.
The project can be run via a CLI, for convenient usage and testing. Every pipeline stage is available as a subcommand of a single entry point:
uv run main.py <command> [options]download: Download the raw dataset from Kaggle.preprocess: Preprocess the data and engineer features.train: Train a model.evaluate: Evaluate a saved model and regenerate all plots.tune: Tune model hyperparameters via cross-validation.
Run uv run main.py <command> --help for command-specific options. Each stage can also be invoked directly as a module (e.g. uv run -m src.training.train ...); the options below apply to both forms.
uv run main.py download [--force]--force: Forces a redownload of the data, in the event of missing or corrupted raw data. Defaults toFalse.
This requires a Kaggle API token, configured either via the kaggle section of config.yaml or a token file: https://www.kaggle.com/settings/api
Dataset: https://www.kaggle.com/datasets/yousefeddin/s-and-p-500-stock-price-end-of-2024
- Extract the archive and place the
.csvfile in the directory<DATA_DIR>/raw/
uv run main.py preprocess [--window-size] [--train-end-year] [--val-end-year]--window-size: The size of the sliding window for sequence generation. Defaults to 21.--train-end-year: The last year of the training split. Defaults to 2016.--val-end-year: The last year of the validation split. Defaults to 2020.
uv run main.py train --model [--epochs] [--lr] [--batch-size] [--num-inducing] [--hidden-size] [--num-layers] [--dropout] [--patience] [--monitor] [--seed]--model: The model:["lstm", "vgp", "variational_gp"]. Required.--epochs: The number of epochs to train. Defaults to 100.--lr: The initial learning rate. Defaults to 0.01.--batch-size: The training batch size. Defaults to 64.--num-inducing: The number of inducing points. Exclusive for Variational GP. Defaults to 100.--hidden-size: The size of the LSTM hidden state. Exclusive for LSTM. Defaults to 64.--num-layers: The number of stacked LSTM layers. Exclusive for LSTM. Defaults to 2.--dropout: The dropout probability between stacked LSTM layers. Exclusive for LSTM. Defaults to 0.2.--mc-samples: The number of MC-Dropout forward passes used to estimate predictive uncertainty at evaluation. Exclusive for LSTM. Set to 1 for a deterministic (homoscedastic) model. Defaults to 50.--patience: The number of epochs to wait for early stopping. Defaults to 10.--monitor: The validation metric to track for early stopping:["RMSE", "NLL", "ECE", "Hit_Rate"]. Defaults to "NLL".--seed: The deterministic seed value to apply. Defaults to theseedset inconfig.yaml.
Training saves the trained model to <MODELS_DIR>/, an evaluation report (<MODEL>_evaluation_report.json) and a set of plots. to <RESULTS_DIR>/.
The following configurations were selected via cross-validated hyperparameter tuning, optimising for validation NLL.
uv run main.py train --model lstm \
--lr 0.0028824672368677117 --batch-size 32 \
--hidden-size 128 --num-layers 3 --dropout 0.24703173405578172 \
--epochs 100 --patience 10 --monitor NLLuv run main.py train --model vgp \
--lr 0.014542214474659556 --batch-size 64 --num-inducing 100 \
--epochs 100 --patience 10 --monitor NLLRegenerate all diagnostic plots and the metrics JSON for a saved model without retraining. The script reconstructs the model architecture from the checkpoint, re-fits the LSTM's aleatoric noise on the validation set (matching the training procedure), and produces every plot the training pipeline would have generated.
uv run main.py evaluate --model <model> [--mc-samples] [--batch-size] [--seed]--model: The model to evaluate:["lstm", "vgp", "variational_gp"]. Required.--mc-samples: [LSTM only] Number of MC-Dropout forward passes. Defaults to 50.--batch-size: Batch size for evaluation. Defaults to 64.--seed: The deterministic seed value to apply. Defaults to theseedset inconfig.yaml.
Tune hyperparameters via chronological (expanding-window) cross-validation and Bayesian Optimisation with Optuna.
uv run main.py tune --model [--n-trials] [--n-splits] [--epochs] [--patience] [--monitor] [--sampler] [--timeout] [--seed]--model: The model to tune:["lstm", "vgp", "variational_gp"]. Required.--n-trials: The number of Bayesian Optimisation trials. Defaults to 30.--n-splits: The number of expanding-window CV folds. Defaults to 5.--epochs: The maximum training epochs per fold. Defaults to 50.--patience: The early-stopping patience per fold. Defaults to 10.--monitor: The validation metric to optimise:["RMSE", "NLL", "ECE", "Hit_Rate"]. Defaults to "NLL".--sampler: The Optuna sampler:["gp", "tpe"]. Defaults to "gp".--timeout: The wall-clock budget in seconds for the whole study. Defaults to no limit.--seed: The deterministic seed value to apply. Defaults to theseedset inconfig.yaml.
The best configuration and metrics are written to <RESULTS_DIR>/<MODEL>_tuning_report.json.
tensorboard --logdir logs/tensorboard