科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ PloS one2026-01-01

Estimating presenteeism from repeated smartphone-based multimodal behavioral responses.

Taiga Noguchi, Shotaro Doki, Masakazu Hirokawa, Soma Nishimura, Katsuya Hotta, Yuya Iwata, Naoko Kouda, Shota Matsumoto, Kenji Suzuki

一句话结论 · In one sentence

Interpretable survival models were developed and internally validated for this clinically heterogeneous population. GBM supported risk stratification, whereas Cox provided calibrated model-based survival estimates within the available clinicopathological framework. The moderate predictive performance and absence of external validation should be considered, and future studies should incorporate additional pathological and molecular predictors.

原始摘要(英文原文)· Original abstract
Presenteeism, attending work despite physical or mental health problems, is a major source of productivity loss worldwide, yet its early detection in everyday settings remains challenging. Conventional self-report instruments are time-consuming and require psychiatrist interpretation, which limits their scalability. We developed a multimodal machine-learning model that estimates work-function impairment as an indicator of presenteeism from brief smartphone-based dialogues. 38 employees recorded short self-report videos (10-30 s) using front-facing smartphone cameras, which were independently rated on an ordinal three-level work-function scale (Healthy / Moderate / Unwell) by three psychiatrists; the consensus label served as supervision. We tested whether integrating multimodal behavioural signals (acoustic, facial, and linguistic) would provide additional cues beyond linguistic content alone, particularly when verbal cues are sparse. Acoustic, facial, and linguistic features were extracted from each clip and integrated using a three-stream Attention-based Multiple Instance Learning (Attention-MIL) framework with learned late fusion, an ordinal-regression head, and class-prior logit adjustment. Generalisation was assessed under a participant-disjoint protocol (25-fold repeated GroupKFold; n = 1,768 out-of-fold clips from 29 participants). On the Full configuration, the proposed multimodal framework achieved macro-F1 = 0.772 (95% CI [0.697, 0.817]), accuracy = 0.840, macro-AUROC = 0.940, and low expected calibration error (ECE ≈0.024). At the aggregate level, the four text-containing configurations (Full / Audio+Text / Face+Text / Text-only) were mutually indistinguishable on macro-F1 (Holm-Bonferroni padj > 0.40), whereas text-free configurations performed substantially worse (Δ<-0.33, padj < 0.001), confirming the necessity of the linguistic channel. In an exploratory subgroup analysis, Audio+Text outperformed Text-only in the clinically ambiguous subgroup-short-speech responses from non-Healthy participants (n = 77; Δ macro-F1 =+0.030, exploratory, requires replication). These findings provide localised but clinically meaningful support for the multimodal hypothesis under low-information conditions and motivate prospective replication.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Estimating presenteeism from repeated smartphone-based multimodal behavioral responses. — 科研速览 Science Skim