4 vials — 10% off · 10 vials — 20% off | Volume discounts applied automatically at checkout
peptides-pro — research peptides

Multi-paradigm Vision Transformer ensemble with regional attention and MLP meta-fusion for explainable dermosc

The automated classification of dermoscopic skin lesions is inherently challenging due to class imbalance, minimal inter-class variance, visual similarity across lesion types, and the requirement for clinically interpret

The present study designed a heterogeneous Vision Transformer ensemble framework for seven-class skin lesion classification using the HAM10000 dataset. The framework integrates three architecturally distinct backbones Swin-Tiny, ViT-Base, and DeiT-Small enhanced with a novel Regional Attention Wrapper (RAW) for spatially selective feature aggregation. The generated outputs are combined via a stacking protocol wherein a trained MLP meta-learner resolves class-aware disagreements among the models. Class imbalance is addressed using class-adaptive augmentation with class-weighted focal loss and MixUp regularisation. The proposed framework achieved 98.37% accuracy, weighted F1-score of 98.39%, and mean AUC of 0.999, surpassing all three individual backbones across all metrics. MEL misclassifications were reduced by 78% compared to the weakest baseline, confirmed by McNemar's test (p < 0.0001). Comprehensive ablation studies validate the contributions of the RAW module, each ensemble component, and each augmentation strategy. Three-fold cross-validation yields a mean accuracy of 96.14 ± 0.36% and bootstrap confidence interval analysis confirms the reliability and reproducibility of the reported results. An exhaustive explainability framework comprising Regional Attention Maps, GradCAM++, SHAP, and t-SNE provides complementary spatial, gradient-based, pixel-level, and embedding-level interpretability, ensuring clinical trust, transparency, and trustworthiness expected from an automated dermoscopy system.

WhatsApp