Enhancing Model Generalization through Aggregated Small Scale Image Dataset Training and Cross Dataset Evaluation
DOI:
https://doi.org/10.66857/2121Keywords:
Alzheimer disease; dataset aggregation; domain generalization; I-JEPA; self-supervised learning; DenseNet-121; magnetic resonance imagingAbstract
Small medical imaging datasets produce good performing models when tested internally; however, such models underperform on other medical imaging datasets from a different site. In our work, we explore the possibility of improving representation and transfer learning by combining two independent datasets containing Alzheimer's MRI images. The source datasets contained 86,437 and 44,000 two-dimensional MRI images belonging to four disease classes. After applying label harmonization and deduplication through the use of MD5 fingerprints, 127,437 unique images were found. Patient-wise split was conducted wherever possible followed by stratified split, normalization, and augmentation during the training process. Moreover, the DenseNet-121 architecture pretrained with ImageNet was fine-tuned on the combined datasets with weighted cross-entropy through two stages. The combined data of I-JEPA gave an accuracy of 88.10%, 89.25%, and 91.80% for Dataset 1, Dataset 2, and the combined test data of the two datasets, respectively. The obtained accuracies were better than the accuracies of the best single datasets by 5.65%, 4.30%, and 13.40%, respectively. Training the DenseNet-121 architecture with the combined data yielded an accuracy of 89.40% and 90.20% for Datasets 1 and 2, respectively. This shows that training on multiple datasets has helped in obtaining higher accuracies in this study. Since the experiments have been done only one time without having a third dataset from other sites, this shows cross-source performance improvement in this study.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

