Research Article | Open Access | Download PDF
Volume 74 | Issue 7 | Year 2026 | Article Id. IJETT-V74I7P119 | DOI : https://doi.org/10.14445/22315381/IJETT-V74I7P119Reliability and Explainability Analysis of CNN and Vision Transformer Models for Brain Tumor MRI Classification
Anuj Gupta, Anita, Manish Gupta
| Received | Revised | Accepted | Published |
|---|---|---|---|
| 14 Feb 2026 | 10 Jun 2026 | 18 Jun 2026 | 28 Jul 2026 |
Citation :
Anuj Gupta, Anita, Manish Gupta, "Reliability and Explainability Analysis of CNN and Vision Transformer Models for Brain Tumor MRI Classification," International Journal of Engineering Trends and Technology (IJETT), vol. 74, no. 7, pp. 281-309, 2026. Crossref, https://doi.org/10.14445/22315381/IJETT-V74I7P119
Abstract
The clinically feasible explanations and the high performance in classifications are not only strong screening outcomes of the Magnetic Resonance Imaging (MRI) on brain tumors, but also trustworthy confidence estimates and explanations. The present paper gives a single-run benchmark of three convolutional neural networks (VGG16, ResNet50, EfficientNetB0) and two families of vision transformers (ViT-Base and Swin-Base) in a 100-epoch training scheme in four-class brain MRI classification (glioma, meningioma, pituitary, and no-tumor). It has accuracy, macro-precision/recall/F1, micro-average ROC: AUC, per-class, and confusion matrix reports. In order to measure reliability, it determines calibration by on-assessment of reliability diagrams and Estimated Calibration Error (ECE). To evaluate the interpretability, Grad-CAM images are given to all five models using post-processing (thresholding and smoothing) to identify class-discriminative areas. In addition to discrimination, the assessment presents the differences in confidence, reliability and the quality of explanation among architects as noteworthy, indicating that the accuracy level alone is insufficient to use a system in clinical practice. The outcomes of this work bring forward the need to be reliability-conscious and responsive towards explainability validation before the process is transformed into real neuroimaging practices. The modern backbones (EfficientNetB0 and Swin) are more discriminative and better calibrated models than VGG16, and saliency explainability is also more correlated with pathology, as opposed to classical baselines. It puts the findings in context of the more current state-of-the-art, in terms of summarizing those recently advanced variants of transformer that can be tuned to better calibration output (up to 98.9% accuracy and ECE as low as 0.023) when using multi-scale patch embeddings, feature calibration modules and selective attention.
Keywords
Brain tumor classification, MRI, Convolutional neural networks, Vision transformer, Swin transformer, Calibration, Reliability, Grad-CAM, Explainable AI.
References
[1] Ujjwal Baid et al., “The
RSNA-ASNR-MICCAI BraTS 2021 Benchmark on Brain Tumor Segmentation and
Radiogenomic Classification,” arXiv Preprint, pp. 1-19, 2021.
[CrossRef] [Google Scholar] [Publisher Link]
[2] Maria Correia de Verdier et
al., “The 2024 Brain Tumor Segmentation Challenge: Glioma Segmentation on
Post-Treatment MRI,” arXiv, pp. 1-10, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[3] Ashish Vaswani et al.,
“Attention is All You Need,” Advances in Neural Information Processing
Systems, vol. 30, 2017.
[Google Scholar] [Publisher Link]
[4] Ze Liu et al., “Swin
Transformer: Hierarchical Vision Transformer using Shifted Windows,” 2021
IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC,
Canada, pp. 9992-10002, 2021.
[CrossRef] [Google Scholar] [Publisher Link]
[5] Alexey Dosovitskiy et al.,
“An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale,” arXiv,
pp. 1-22, 2021.
[CrossRef] [Google Scholar] [Publisher Link]
[6] Omid Nejati Manzari et al.,
“MedViT: A Robust Vision Transformer for Generalized Medical Image
Classification,” Computers in Biology and Medicine, vol. 157, 2023.
[CrossRef] [Google Scholar] [Publisher Link]
[7] Bas H.M. van der Velden et
al., “Explainable Artificial Intelligence in Deep Learning-based Medical Image
Analysis,” Medical Image Analysis, vol. 79, pp. 1-21, 2022.
[CrossRef] [Google Scholar] [Publisher Link]
[8] Ke Zou et al., “A Review of
Uncertainty Estimation and its Application in Medical Imaging,” Meta-Radiology,
vol. 1, no. 1, pp. 1-10, 2023.
[CrossRef] [Google Scholar] [Publisher Link]
[9] Chuan Guo et al., “On
Calibration of Modern Neural Networks,” Proceedings of the 34th
International Conference on Machine Learning, PMLR, vol. 70, pp. 1321-1330,
2017.
[Google Scholar] [Publisher Link]
[10] Amna Iqbal, Muhammad Arfan
Jaffar, and Rashid Jahangir et al., “Enhancing Brain Tumour
Multi-Classification using Efficient-Net B0-based Intelligent Diagnosis for
Internet of Medical Things (IoMT) Applications,” Information, vol. 15,
no. 8, pp. 1-15, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[11] Katarzyna Borys et al.,
“Explainable AI in Medical Imaging: An Overview for Clinical Practitioners –
Beyond Saliency-based XAI Approaches,” European Journal of Radiology,
vol. 162, pp. 1-11, 2023.
[CrossRef] [Google Scholar] [Publisher Link]
[12] Jieneng Chen et al.,
“TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation,” arXiv,
pp. 1-13, 2021.
[CrossRef] [Google Scholar] [Publisher Link]
[13] Ali Hatamizadeh et al.,
“UNETR: Transformers for 3D Medical Image Segmentation,” Proceedings of the
IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp.
574-584, 2022.
[Google Scholar] [Publisher Link]
[14] Shakif Ahmed et al.,
“Advancing Brain Tumor MRI Classification using SwRD: A Parallel Swin
Transformer-ResNet Approach,” Virtual Reality and Intelligent Hardware,
vol. 7, no. 5, pp. 501-522, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[15] C. Sankari, V. Jamuna, and
A.R. Kavitha et al., “Hierarchical Multi-Scale Vision Transformer Model for
Accurate Detection and Classification of Brain Tumors in MRI-based Medical
Imaging,” Scientific Reports, vol. 15, no. 1, pp. 1-19, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[16] Mohammad Ali Labbaf Khaniki
et al., “Vision Transformer with Feature Calibration and Selective
Cross-Attention for Brain Tumor Classification,” Iran Journal of Computer
Science, vol. 8, no. 2, pp. 335-347, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[17] C. Kishor Kumar Reddy et
al., “A Fine-Tuned Vision Transformer based Enhanced Multi-Class Brain Tumor
Classification using MRI Scan Imagery,” Frontiers in Oncology, vol. 14,
pp. 1-23, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[18] Yigitcan Cakmak, and Ishak
Pacal, “Comparative Analysis of Transformer Architectures for Brain Tumor
Classification,” Exploration of Medicine, vol. 6, pp. 1-14, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[19] Khawla Hussein Ali,
“ViT-BT: Improving MRI Brain Tumor Classification using Vision Transformer with
Transfer Learning,” International Journal of Soft Computing and Engineering,
vol. 14, no. 4, pp. 16-26, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[20] Ramprasaath R. Selvaraju et
al., “Grad-CAM: Visual Explanations from Deep Networks via Gradient-based
Localization,” Proceedings of the IEEE International Conference on Computer
Vision (ICCV), pp. 618-626, 2017.
[Google Scholar] [Publisher Link]
[21] Haomin Chen et al.,
“Explainable Medical Imaging AI needs Human-Centered Design: Guidelines and
Evidence from a Systematic Review,” npj Digital Medicine, vol. 5, no. 1,
2022.
[CrossRef] [Google Scholar] [Publisher Link]
[22] Ling Huang et al., “A
Review of Uncertainty Quantification in Medical Image Analysis: Probabilistic
and Non-Probabilistic Methods,” Medical Image Analysis, vol. 97, 2024.
[CrossRef] [Google Scholar] [Publisher Link]