International Journal of Engineering
Trends and Technology

Research Article | Open Access | Download PDF
Volume 74 | Issue 7 | Year 2026 | Article Id. IJETT-V74I7P119 | DOI : https://doi.org/10.14445/22315381/IJETT-V74I7P119

Reliability and Explainability Analysis of CNN and Vision Transformer Models for Brain Tumor MRI Classification


Anuj Gupta, Anita, Manish Gupta

Received Revised Accepted Published
14 Feb 2026 10 Jun 2026 18 Jun 2026 28 Jul 2026

Citation :

Anuj Gupta, Anita, Manish Gupta, "Reliability and Explainability Analysis of CNN and Vision Transformer Models for Brain Tumor MRI Classification," International Journal of Engineering Trends and Technology (IJETT), vol. 74, no. 7, pp. 281-309, 2026. Crossref, https://doi.org/10.14445/22315381/IJETT-V74I7P119

Abstract

The clinically feasible explanations and the high performance in classifications are not only strong screening outcomes of the Magnetic Resonance Imaging (MRI) on brain tumors, but also trustworthy confidence estimates and explanations. The present paper gives a single-run benchmark of three convolutional neural networks (VGG16, ResNet50, EfficientNetB0) and two families of vision transformers (ViT-Base and Swin-Base) in a 100-epoch training scheme in four-class brain MRI classification (glioma, meningioma, pituitary, and no-tumor). It has accuracy, macro-precision/recall/F1, micro-average ROC: AUC, per-class, and confusion matrix reports. In order to measure reliability, it determines calibration by on-assessment of reliability diagrams and Estimated Calibration Error (ECE). To evaluate the interpretability, Grad-CAM images are given to all five models using post-processing (thresholding and smoothing) to identify class-discriminative areas. In addition to discrimination, the assessment presents the differences in confidence, reliability and the quality of explanation among architects as noteworthy, indicating that the accuracy level alone is insufficient to use a system in clinical practice. The outcomes of this work bring forward the need to be reliability-conscious and responsive towards explainability validation before the process is transformed into real neuroimaging practices. The modern backbones (EfficientNetB0 and Swin) are more discriminative and better calibrated models than VGG16, and saliency explainability is also more correlated with pathology, as opposed to classical baselines. It puts the findings in context of the more current state-of-the-art, in terms of summarizing those recently advanced variants of transformer that can be tuned to better calibration output (up to 98.9% accuracy and ECE as low as 0.023) when using multi-scale patch embeddings, feature calibration modules and selective attention.

Keywords

Brain tumor classification, MRI, Convolutional neural networks, Vision transformer, Swin transformer, Calibration, Reliability, Grad-CAM, Explainable AI.

References

[1] Ujjwal Baid et al., “The RSNA-ASNR-MICCAI BraTS 2021 Benchmark on Brain Tumor Segmentation and Radiogenomic Classification,” arXiv Preprint, pp. 1-19, 2021.
[
CrossRef] [Google Scholar] [Publisher Link]

[2] Maria Correia de Verdier et al., “The 2024 Brain Tumor Segmentation Challenge: Glioma Segmentation on Post-Treatment MRI,” arXiv, pp. 1-10, 2024.
[
CrossRef] [Google Scholar] [Publisher Link]

[3] Ashish Vaswani et al., “Attention is All You Need,” Advances in Neural Information Processing Systems, vol. 30, 2017.
[
Google Scholar] [Publisher Link] 

[4] Ze Liu et al., “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, pp. 9992-10002, 2021.
[
CrossRef] [Google Scholar] [Publisher Link]

[5] Alexey Dosovitskiy et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale,” arXiv, pp. 1-22, 2021.
[
CrossRef] [Google Scholar] [Publisher Link]

[6] Omid Nejati Manzari et al., “MedViT: A Robust Vision Transformer for Generalized Medical Image Classification,” Computers in Biology and Medicine, vol. 157, 2023.
[
CrossRef] [Google Scholar] [Publisher Link]

[7] Bas H.M. van der Velden et al., “Explainable Artificial Intelligence in Deep Learning-based Medical Image Analysis,” Medical Image Analysis, vol. 79, pp. 1-21, 2022.
[
CrossRef] [Google Scholar] [Publisher Link]

[8] Ke Zou et al., “A Review of Uncertainty Estimation and its Application in Medical Imaging,” Meta-Radiology, vol. 1, no. 1, pp. 1-10, 2023.
[
CrossRef] [Google Scholar] [Publisher Link]

[9] Chuan Guo et al., “On Calibration of Modern Neural Networks,” Proceedings of the 34th International Conference on Machine Learning, PMLR, vol. 70, pp. 1321-1330, 2017.
[
Google Scholar] [Publisher Link]

[10] Amna Iqbal, Muhammad Arfan Jaffar, and Rashid Jahangir et al., “Enhancing Brain Tumour Multi-Classification using Efficient-Net B0-based Intelligent Diagnosis for Internet of Medical Things (IoMT) Applications,” Information, vol. 15, no. 8, pp. 1-15, 2024.
[
CrossRef] [Google Scholar] [Publisher Link]

[11] Katarzyna Borys et al., “Explainable AI in Medical Imaging: An Overview for Clinical Practitioners – Beyond Saliency-based XAI Approaches,” European Journal of Radiology, vol. 162, pp. 1-11, 2023.
[
CrossRef] [Google Scholar] [Publisher Link]

[12] Jieneng Chen et al., “TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation,” arXiv, pp. 1-13, 2021.
[
CrossRef] [Google Scholar] [Publisher Link]

[13] Ali Hatamizadeh et al., “UNETR: Transformers for 3D Medical Image Segmentation,” Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 574-584, 2022.
[
Google Scholar] [Publisher Link]

[14] Shakif Ahmed et al., “Advancing Brain Tumor MRI Classification using SwRD: A Parallel Swin Transformer-ResNet Approach,” Virtual Reality and Intelligent Hardware, vol. 7, no. 5, pp. 501-522, 2025.
[
CrossRef] [Google Scholar] [Publisher Link]

[15] C. Sankari, V. Jamuna, and A.R. Kavitha et al., “Hierarchical Multi-Scale Vision Transformer Model for Accurate Detection and Classification of Brain Tumors in MRI-based Medical Imaging,” Scientific Reports, vol. 15, no. 1, pp. 1-19, 2025.
[
CrossRef] [Google Scholar] [Publisher Link]

[16] Mohammad Ali Labbaf Khaniki et al., “Vision Transformer with Feature Calibration and Selective Cross-Attention for Brain Tumor Classification,” Iran Journal of Computer Science, vol. 8, no. 2, pp. 335-347, 2024.
[
CrossRef] [Google Scholar] [Publisher Link]

[17] C. Kishor Kumar Reddy et al., “A Fine-Tuned Vision Transformer based Enhanced Multi-Class Brain Tumor Classification using MRI Scan Imagery,” Frontiers in Oncology, vol. 14, pp. 1-23, 2024.
[
CrossRef] [Google Scholar] [Publisher Link]

[18] Yigitcan Cakmak, and Ishak Pacal, “Comparative Analysis of Transformer Architectures for Brain Tumor Classification,” Exploration of Medicine, vol. 6, pp. 1-14, 2025.
[
CrossRef] [Google Scholar] [Publisher Link]

[19] Khawla Hussein Ali, “ViT-BT: Improving MRI Brain Tumor Classification using Vision Transformer with Transfer Learning,” International Journal of Soft Computing and Engineering, vol. 14, no. 4, pp. 16-26, 2024.
[
CrossRef] [Google Scholar] [Publisher Link]

[20] Ramprasaath R. Selvaraju et al., “Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 618-626, 2017.
[
Google Scholar] [Publisher Link]

[21] Haomin Chen et al., “Explainable Medical Imaging AI needs Human-Centered Design: Guidelines and Evidence from a Systematic Review,” npj Digital Medicine, vol. 5, no. 1, 2022.
[
CrossRef] [Google Scholar] [Publisher Link]

[22] Ling Huang et al., “A Review of Uncertainty Quantification in Medical Image Analysis: Probabilistic and Non-Probabilistic Methods,” Medical Image Analysis, vol. 97, 2024.
[
CrossRef] [Google Scholar] [Publisher Link]