Multimodal Uncertainty Quantification with Evidential Learning and Selective Prediction for Lightweight Vision-Language Classification on LUMA
- Maya He
- EIRA Journal of Multidisciplinary Research and Development (EIRAJMRD)
- https://doi.org/10.5281/zenodo.21609978
Published:
Sunday, 26 July 2026
Volume:
Volume 2, Issue 4 (2026)
Section:
Articles
Abstract
Reliable vision-language classifiers require predictions that are not only accurate but also honest about residual uncertainty. This paper presents a full empirical evaluation of uncertainty quantification on the LUMA image-text branch for lightweight multimodal classification. The analyzed corpus contains 44,000 32 x 32 images, 62,875 text descriptions, and 44 in-distribution labels with both image and text support. The experiment pairs all text rows from image-supported classes with class-consistent images, uses a held-out validation set for temperature scaling, a separate calibration split for conformal prediction, and a test split for clean, corrupted, and semantic-conflict evaluation. Four uncertainty mechanisms are compared: temperature scaling, a three-member deep ensemble, evidential learning with a Dirichlet output layer, and split conformal prediction. On the clean test split, temperature scaling retained 96.21% accuracy while reducing expected calibration error to 0.0052. The deep ensemble achieved the strongest clean top-1 accuracy at 96.78% and the best semantic-conflict AUROC at 0.974. Conformal prediction reached 89.91% clean coverage at a 90% target, while its coverage decreased under distribution shift. The results show that scalar calibration, ensembles, evidential uncertainty, conformal sets, and selective prediction capture different failure modes, and that lightweight multimodal tasks benefit most from combining calibrated probabilities with ensemble-based uncertainty ranking.
Keywords: multimodal uncertainty quantification, LUMA, evidential learning, deep ensembles, conformal prediction, selective prediction, probability calibration
How to cite this work: Maya He. (2026). Multimodal Uncertainty Quantification with Evidential Learning and Selective Prediction for Lightweight Vision-Language Classification on LUMA. EIRA Journal of Multidisciplinary Research and Development (EIRAJMRD), 2(4), 40–54. https://doi.org/10.5281/zenodo.21609978
