Cross-attention and reconstruction-regularized Transformers for multi-label chest X-ray classification
This repository includes the code of the final project for the class Machine Learning for Human Data at Università degli Studi di Padova, Fall 2025.
The project tackles the problem of multi-label classification of chest X-ray images from the ChestMNIST dataset using 4 baseline models and 3 newly proposed Transformer-based models.
Sample images and associated labels from the ChestMNIST dataset:
Models used: ResNet50, DenseNet121, EfficientNetB0, MedViT.
All models except MedViT are trained and evaluated in ml4hd2026-baseline.ipynb.
A structured cross-attention framework that explicitly models interaction between local convolutional features and global transformer representations.
Model is implemented in dbct.py and evaluated in ml4hd2026-dbct.ipynb.
A Vision Transformer architecture with additional reconstruction-based regularization strategy to improve representation robustness under severe class imbalance.
Model is implemented in rrvit.py and evaluated in rr-vit.ipynb.
A variant of RR-ViT with added architectural scaling and inductive bias.
Model is implemented in hybridrrvit.py and evaluated in hybrid-rr-vit.ipynb.
3 sample images of different dimensions are available in demo/. To run the Streamlit demo:
streamlit run streamlit_demo.py


