Objective: To evaluate linear-time RWKV-based multimodal fusion for classifying mild cognitive impairment (MCI) from synchronized smartphone audio and consumer wearable signals, and to compare alternative fusion strategies for remote screening.
Background: Early cognitive decline is often under-recognized because clinic-based testing is episodic and difficult to scale. Remote longitudinal protocols using everyday devices can capture speech as well as physiologic/behavioral correlates of cognition. Prior clinical validation has shown that voice biomarkers generalize across distinct linguistic and healthcare contexts, motivating multimodal fusion that preserves interpretability while enabling long-sequence modeling.
Method: We analyzed 150 participants (117 MCI; 33 cognitively unimpaired) enrolled in a remote monitoring protocol aligned with established voice-biomarker deployments across two cohorts (Poland and South Korea). Participants completed structured speech tasks (e.g., picture description, memory recall) and scheduled phone interactions, while concurrent wearable sensing provided complementary behavioral and physiologic time series. Models were evaluated using group-stratified 5-fold cross-validation. Performance metrics included AUC, F1, accuracy, precision, and recall.
Results: Hierarchical RWKV fusion achieved the best overall performance (AUC 0.8239±0.0503; F1 0.8416±0.0617; accuracy 0.7467±0.0886; precision 0.8455; recall 0.8493). Interleaved fusion yielded a similar AUC (0.8183) but lower accuracy (0.6967). Cross-modal and parallel strategies underperformed (AUC 0.7626 and 0.7495, respectively). The observed advantage of hierarchical fusion is consistent with staged integration of modality-specific representations, particularly when speech-derived features (e.g., pause/silence statistics and spectral variability) provide a strong cognitive signal and wearables add complementary context.
Conclusion: RWKV-based fusion provides an efficient alternative to quadratic-complexity attention for multimodal MCI screening using mainstream devices, with hierarchical fusion offering the most reliable discrimination. Given prior cross-cultural clinical validation of voice biomarkers, RWKV multimodal modeling is a promising step toward scalable, longitudinal telemedicine screening; external validation across sites and broader clinical spectra is warranted.
To cite this abstract in AMA style:
D. Hemmerling, JH. Yun, J. Krzywdziak, B. Eljasiak, A. Pruszek, L. Lazarski, M. Grzeszczyk, T. Brzózka, W. Szecówka, W. średniawa, H. Lee, W. Baek, S. Yoon, M. Matuszewski. Linear-Time RWKV Multimodal Fusion of Smartphone Audio and Wearable Signals for Mild Cognitive Impairment Screening [abstract]. Mov Disord. 2026; 41 (suppl 1). https://www.mdsabstracts.org/abstract/linear-time-rwkv-multimodal-fusion-of-smartphone-audio-and-wearable-signals-for-mild-cognitive-impairment-screening/. Accessed October 1, 2026.« Back to 2026 International Congress
MDS Abstracts - https://www.mdsabstracts.org/abstract/linear-time-rwkv-multimodal-fusion-of-smartphone-audio-and-wearable-signals-for-mild-cognitive-impairment-screening/
