科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in medicine2026-01-01

An explainable ResNet50-BiLSTM-attention framework with spatial token modeling and imbalance-aware learning for multi-class knee osteoarthritis severity grading.

Ali Raza, Asma Abbas, Fareeha Hanif

原始摘要(英文原文)· Original abstract
Knee osteoarthritis (KOA) severity grading from radiographs remains challenging because adjacent Kellgren-Lawrence stages exhibit subtle and overlapping structural characteristics. This study proposes ResBiAtt-KOA-Net, an explainable image-level framework that introduces spatial-sequence reasoning into convolutional KOA classification. A pre-trained ResNet50 extracts a (2, 048 × 7 × 7) feature map, which is transformed into 49 ordered spatial tokens. A two-layer bidirectional long short-term memory network models contextual dependencies among these tokens, while additive attention identifies the most informative anatomical regions. The attention-guided representation is subsequently fused with global average- and maximum-pooled convolutional descriptors. Weighted random sampling, class-weighted cross-entropy, and label smoothing are incorporated to improve learning from underrepresented severity categories. The framework was evaluated on 9,786 knee radiographs using a patient-level split, direct same-protocol baselines, component-wise ablation, bootstrap confidence intervals, McNemar testing, expert explainability assessment, and external validation. On the internal test set, ResBiAtt-KOA-Net achieved an accuracy of 0.9392, balanced accuracy of 0.9283, macro-F1 of 0.9263, and macro ROC-AUC of 0.9478. It improved macro-F1 by 0.0177 over the strongest baseline, with an adjusted McNemar (p)-value of 0.0181. External evaluation on the MOST cohort produced an accuracy of 0.8615 and macro-F1 of 0.8389, indicating reasonable cross-dataset transfer despite domain differences. Expert assessment identified 82.7% of Grad-CAM maps and 78.0% of spatial-token attention maps as clinically relevant. The model required 32.47 million parameters and an average inference time of 7.6 ms per image. These findings demonstrate that combining convolutional representation learning, spatial-token dependency modeling, attention-guided fusion, and imbalance-aware optimization provides an accurate, interpretable, and reproducible approach to multi-class KOA severity grading.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

An explainable ResNet50-BiLSTM-attention framework with spatial token modeling and imbalance-aware learning for multi-class knee osteoarthritis severity grading. — 科研速览 Science Skim