科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-09-07· health informatics

Rethinking Input Complexity in Transformer-Based Clinical Prediction: Implications for Feature Dimensionality and Sequence Length in Longitudinal EHR Data

W. Chen, B. Zhou, R. S. Zeiger, W. W. Crawford, M. Schatz, E. J. Puttock, S. Xu, M. T. Slaughter, F. Xie

原始摘要(英文原文)· Original abstract
Objective: Transformer-based models for clinical prediction using longitudinal electronic health record (EHR) data are often developed with large feature sets and long patient histories under the assumption that more data improves performance. However, high-dimensional inputs and long sequences increase computational burden, potentially limiting scalability. We evaluated how feature dimensionality and sequence length affect predictive performance, calibration, risk stratification, and computational efficiency in EHR prediction. Methods: Using longitudinal EHR data from adults with mild asthma in an integrated healthcare system, we evaluated input representation design for predicting acute asthma exacerbation. Feature dimensionality was reduced using Integrated Gradients attribution scores, univariate performance-based selection, and clinically guided selection strategies. Sequence length was varied using percentile-based truncation of patient histories. Performance was assessed using discrimination, calibration, high-risk classification, threshold-based event capture, and computational efficiency across regions. Results: Models using fewer features achieved discrimination comparable to the 80-feature reference model, with AUROC values ranging from 0.843 to 0.864 versus 0.870 for the full model. Moderate sequence-length truncation reduced training time by more than 70% with minimal loss in discrimination. Although reduced-dimensional models showed attenuation of predicted risk at the upper tail, they identified similar high-risk populations and captured comparable proportions of asthma exacerbation events at clinically relevant thresholds. Conclusion: Transformer-based prediction models maintained strong performance across reduced feature sets. While dimensionality reduction modestly affected calibration at the highest risk levels, moderate sequence-length reduction substantially reduced computational burden with limited change in overall discrimination. These findings highlight trade-offs between input complexity, predictive performance, and computational efficiency.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Rethinking Input Complexity in Transformer-Based Clinical Prediction: Implications for Feature Dimensionality and Sequence Length in Longitudinal EHR Data — 科研速览 Science Skim