| Titre : |
Evaluating vision transformer adaptation for Chest X-Ray disease classification |
| Type de document : |
document multimédia |
| Auteurs : |
Ali Chihab Edine Benbertal, Auteur ; Saad Zenikheri, Auteur ; Ilyes Guellouma, Directeur de thèse |
| Editeur : |
Laghouat : Université Amar Telidji - Département d'informatique |
| Année de publication : |
2026 |
| Importance : |
54 p. |
| Accompagnement : |
1 disque optique numérique (CD-ROM) |
| Note générale : |
Option : Data science and artificial intelligence |
| Langues : |
Anglais (eng) |
| Mots-clés : |
Artificial intelligence Chest X-rays Vision transformers Transfer learning Low-Rank Adaptation (LoRA) Parameter-efficient fine-tuning Medical image classification DenseNet-121 NIH ChestX-ray14 Multi-Label classification |
| Résumé : |
This thesis investigates the optimal strategy for adapting Vision Transformers (ViT) to chest X-ray disease classification under resource-constrained conditions. Starting from models pre-trained on ImageNet, we systematically evaluate three fine-tuning paradigms—Full Fine-Tuning, Partial Fine-Tuning, and Low-Rank Adaptation (LoRA) — against a DenseNet-121 CNN baseline on the NIH ChestX-ray14 dataset covering 14 pathological findings.
All experiments are conducted on a patient-stratified split of 70,000 training, 14,000 validation, and approximately 28,000 test images on an NVIDIA Tesla T4 GPU.
DenseNet-121 achieves the highest Mean AUROC of 0.8322 (Epoch 10, Macro F1 = 0.2542), confirming that CNN inductive biases remain advantageous at moderate data scales. Notably, with extended training of up to 25 epochs, the ViT-based strategies converge to substantially stronger results than previously reported at shorter budgets.
ViT Full Fine-Tuning attains 0.8185 AUROC at Epoch 9 using 2,696 MB VRAM, while ViT LoRA achieves 0.7998 AUROC at Epoch 17 with only 1,412 MB VRAM and approximately 1.0 million trainable parameters. ViT Partial Fine-Tuning reaches 0.7957 AUROC at the lowest GPU memory footprint of 1,168 MB—a 57% reduction compared to Full Fine-Tuning.
A key finding of this study is that ViT LoRA, with only 1.2% of Full Fine-Tuning’s trainable parameters, achieves competitive AUROC (0.7998 vs. 0.8185, a gap of 0.0187) while requiring 47% less GPU memory. Moreover, LoRA’s training curve was the most stable among all ViT strategies, converging progressively across all 20 epochs without premature overfitting. These findings yield a clear, evidence-based guideline: when a Vision Transformer must be adapted for medical imaging under limited computational budgets, LoRA emerges as the preferred ViT adaptation strategy. However, at the scale of approximately 70,000 labeled training images, CNNs remain the overall default choice owing to their translation equivariant inductive biases.
All reported metrics reflect validation performance on a patient-stratified split. As this split was also used for early stopping decisions in certain configurations, reported metrics may be optimistically biased relative to a fully independent test set — a limitation acknowledged throughout this thesis. |
| note de thèses : |
Mémoire de master en informatique |
Evaluating vision transformer adaptation for Chest X-Ray disease classification [document multimédia] / Ali Chihab Edine Benbertal, Auteur ; Saad Zenikheri, Auteur ; Ilyes Guellouma, Directeur de thèse . - Laghouat : Université Amar Telidji - Département d'informatique, 2026 . - 54 p. + 1 disque optique numérique (CD-ROM). Option : Data science and artificial intelligence Langues : Anglais ( eng)
| Mots-clés : |
Artificial intelligence Chest X-rays Vision transformers Transfer learning Low-Rank Adaptation (LoRA) Parameter-efficient fine-tuning Medical image classification DenseNet-121 NIH ChestX-ray14 Multi-Label classification |
| Résumé : |
This thesis investigates the optimal strategy for adapting Vision Transformers (ViT) to chest X-ray disease classification under resource-constrained conditions. Starting from models pre-trained on ImageNet, we systematically evaluate three fine-tuning paradigms—Full Fine-Tuning, Partial Fine-Tuning, and Low-Rank Adaptation (LoRA) — against a DenseNet-121 CNN baseline on the NIH ChestX-ray14 dataset covering 14 pathological findings.
All experiments are conducted on a patient-stratified split of 70,000 training, 14,000 validation, and approximately 28,000 test images on an NVIDIA Tesla T4 GPU.
DenseNet-121 achieves the highest Mean AUROC of 0.8322 (Epoch 10, Macro F1 = 0.2542), confirming that CNN inductive biases remain advantageous at moderate data scales. Notably, with extended training of up to 25 epochs, the ViT-based strategies converge to substantially stronger results than previously reported at shorter budgets.
ViT Full Fine-Tuning attains 0.8185 AUROC at Epoch 9 using 2,696 MB VRAM, while ViT LoRA achieves 0.7998 AUROC at Epoch 17 with only 1,412 MB VRAM and approximately 1.0 million trainable parameters. ViT Partial Fine-Tuning reaches 0.7957 AUROC at the lowest GPU memory footprint of 1,168 MB—a 57% reduction compared to Full Fine-Tuning.
A key finding of this study is that ViT LoRA, with only 1.2% of Full Fine-Tuning’s trainable parameters, achieves competitive AUROC (0.7998 vs. 0.8185, a gap of 0.0187) while requiring 47% less GPU memory. Moreover, LoRA’s training curve was the most stable among all ViT strategies, converging progressively across all 20 epochs without premature overfitting. These findings yield a clear, evidence-based guideline: when a Vision Transformer must be adapted for medical imaging under limited computational budgets, LoRA emerges as the preferred ViT adaptation strategy. However, at the scale of approximately 70,000 labeled training images, CNNs remain the overall default choice owing to their translation equivariant inductive biases.
All reported metrics reflect validation performance on a patient-stratified split. As this split was also used for early stopping decisions in certain configurations, reported metrics may be optimistically biased relative to a fully independent test set — a limitation acknowledged throughout this thesis. |
| note de thèses : |
Mémoire de master en informatique |
|  |