Adaptive Confidence-Guided Multi-Scale Vision Transformer Framework for Explainable Microscopic Fungal Species Classification
ID:96
Submission ID:507 View Protection:ATTENDEE
Updated Time:2026-07-25 18:00:37
Hits:21
Online
Start Time:2026-07-30 15:55 (Asia/Kolkata)
Duration:15min
Session:[S4] Computer Vision and Pattern Recognition » [S4-2] Computer Vision and Pattern Recognition
Video
No Permission
Presentation File
Tips: The file permissions under this presentation are only for participants. You have not logged in yet and cannot view it temporarily.
Abstract
Pathogenic fungi are difficult to identify in the field. The task of reading microscopic slides is still primarily one that requires training from human eyes — a point of delay in clinical
decisions and an introduction of inter-observer error. This paper proposes a new confidence-guided multi-scale fusion framework called ACMF. The framework runs two Vision Transformer (ViT-B/16) encoders at different input resolutions — 224 and
384 pixels — and merges the outputs in four different ways. These range from simple averaging (S1) and accuracy-weighted blending (S2) to a per-image entropy-based weighting scheme (S3) and an expert MLP fusion head (S4). Experiments on the
DeFungi benchmark — five clinically relevant fungal species, 6,801 microscopic images — demonstrate competitive accuracy. The best-performing variant is S2, with a macro F1 of 91.63% on the held-out test set — a 3.16-point gain over ViT-B/16-224 used alone. Attention Rollout maps are generated at inference time, highlighting hyphae and spore structures without additional annotation. Comparisons with ResNet-50, EfficientNet-B3, and Swin-Tiny confirm that ACMF delivers the most consistent value.
decisions and an introduction of inter-observer error. This paper proposes a new confidence-guided multi-scale fusion framework called ACMF. The framework runs two Vision Transformer (ViT-B/16) encoders at different input resolutions — 224 and
384 pixels — and merges the outputs in four different ways. These range from simple averaging (S1) and accuracy-weighted blending (S2) to a per-image entropy-based weighting scheme (S3) and an expert MLP fusion head (S4). Experiments on the
DeFungi benchmark — five clinically relevant fungal species, 6,801 microscopic images — demonstrate competitive accuracy. The best-performing variant is S2, with a macro F1 of 91.63% on the held-out test set — a 3.16-point gain over ViT-B/16-224 used alone. Attention Rollout maps are generated at inference time, highlighting hyphae and spore structures without additional annotation. Comparisons with ResNet-50, EfficientNet-B3, and Swin-Tiny confirm that ACMF delivers the most consistent value.
Keywords
fungal classification,vision transformer,deep learning,explainable AI,multi-scale fusion,DeFungi,microscopy
Speaker
Comment submit