Quantifying Triple-XAI Disagreement and Attribution Reliability across Grad-CAM, SHAP and LIME for a Hybrid CNN–Transformer Brain Tumor Classifier

Authors

  • Muhammad Sheraz Faculty of Computer Science & IT, Superior University Lahore, 54000, Pakistan Author
  • Nasir Ayub Engineering Calrom Limited, M16EG, United Kingdom Author
  • Bisma Tanveer Mirza Faculty of Computing and Information Technology(FCIT), University of the Punjab, Lahore, Pakistan Author
  • Chaudary Umair Mehmood Edge Hill University, Ormskirk L39 4QP, United Kingdom Author
  • Hamayun Khan Department of Computer Science, Faculty of Computer Science & IT, Superior University, Lahore, 54000, Pakistan Author
  • Muhammad Zunnurain Hussain Bahria University, Lahore Campus Author
  • Muhammad Waleed Khawar Innova Network, 184 C, Airline Society, Lahore, 54000, Pakistan Author
  • Muhammad Abdullah IMAB Printing and Tech Solutions, Lahore, 54000, Pakistan Author
  • Syed Muhammad Rizwan Department of Computer Engineering, University of Engineering and Technology, Lahore, 54000, Pakistan Author

Keywords:

Machine Learning; brain tumor classification, magnetic resonance imaging, explainable artificial intelligence, Grad-CAM, SHAP, LIME, attribution reliability, vision transformer

Abstract

Deep-learning models are accurate to a degree that leaves little scope for improvement; brain tumor classification models from MRI are hard to deploy to practice, typically barely worth the decisions that they make. If there is a justification provided, this is generally produced by only one of the pothooking explanation techniques and refers to an overall explanation of the reasoning of the model. This paper proceeds on the basis of that assumption. We use a patient-disjoint protocol and leakage-safe manner to train a Hybrid CNN–Transformer classifier on the Mendeley dataset of four-level brain tumor MRIs; we then use Grad-CAM, SHAP and LIME to visualize and quantify the explanations, and their agreement to the predictions, as well as the trustworthiness of the three explanations. The classifier achieves a 99.88% accuracy (SD 0.37%, 4-fold patient-disjoint cross validated mean 98.85% (95% CI 98.26–99.44%)) and a 99.88% macro F1-score on a held-out client never seen during training. To eliminate the contribution from ResNet50 (which has the same training opportunity), the same ablating procedure is repeated on the same set of images when the two models disagree in their contribution. The ablating is repeated alone on the same set of images where the two models disagree with each other, and the hybrid is correct on 21 images, wrong on none, for which the McNemar p-value is 9.54 × 10⁻⁷. In the process of explanation analysis, the main findings of the paper are produced. There is only a 50% to 33% intersection between any two of the three attribution methods with Dice-DW coefficient ranging between 0.3299 and 0.4383. They are generally within a range of each other (0.7008, 0.6922, 0.6750) and both are significantly higher than a random attribution simulation (0.0240), that is, each method gathers the evidence that is relevant to the class, but that is not enough to definitively outperform the other. When the three measurements are combined together in the same proportions to the degrees of the measured reliability, they yield a combined map of fidelity 0.6895, which is just below, instead of above, the items that individually had the highest level of fidelity. We openly report this negative finding because we believe that reliability-weighted fusion can be viewed as a way to reconcile different explanations in addition to being a means of maximising fidelity.

Downloads

Download data is not yet available.

Downloads

Published

2026-02-28

How to Cite

Quantifying Triple-XAI Disagreement and Attribution Reliability across Grad-CAM, SHAP and LIME for a Hybrid CNN–Transformer Brain Tumor Classifier. (2026). Annual Methodological Archive Research Review, 4(2), 355-384. https://amresearchjournal.com/index.php/Journal/article/view/2510

Similar Articles

1-10 of 800

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)

<< < 1 2 3 4 > >>