Objective: To evaluate the performance of deep learning approach based on Mel-spectrogram features in classifying PD using two real-world speech tasks – sustained phonation and articulated speech.
Background: Parkinson’s disease (PD) is a progressive neurodegenerative disorder which speech impairment can occur early, making voice a promising non-invasive digital biomarker for PD detection [1]. However, many machine learning studies rely on small or laboratory-controlled datasets, limiting their generalizability to real-world environments [2]. Deep learning approaches using spectrogram-based representations may better capture complex time–frequency speech patterns than traditional approaches.
Method: Voice recordings were collected using the CheckPD mobile application from individuals with PD and healthy controls in unsupervised real-world environments. Participants performed two tasks: sustained phonation (prolonged “ahh”) and articulated speech (a sentence from the Thai Mini-Mental State Examination). Recordings were segmented and standardized into 5-second segments (Figure 1). Audio signals were converted into Mel-spectrograms – a perceptually scaled time–frequency representation of speech – and used as input for pretrained convolutional neural networks (MobileNet, EfficientNet, and ResNet). Models were fine-tuned and evaluated using subject-level train–validation–test splitting to avoid data leakage Performance was assessed on the test set using area under the receiver operating characteristic curve (AUC), accuracy, sensitivity, and specificity.
Results: A total of 648 participants were included (PD = 281; controls = 367). Participants with PD were older than controls (63.69 ± 11.77 vs 56.79 ± 14.15 years, p < 0.001), and most had mild-to-moderate disease severity. Models trained on articulated speech outperformed sustained phonation. Sustained phonation showed moderate performance (AUC 0.690–0.770, accuracy up to 69.12%), while articulated speech achieved higher performance, with the best model (EfficientNet-B1) reaching 83.82% accuracy and an AUC of 0.897 on the test set (Figure 2).
Conclusion: CNN-based voice classification using Mel-spectrogram features demonstrates the feasibility of real-world mobile detection of PD. Articulated speech consistently outperformed sustained phonation, indicating that connected speech provides more discriminative information for PD screening.
CheckPD Voice Processing Pipeline
ROC Curves for PD Voice Classification Models
References: [1] Bhidayasiri R, Sringean J, Phumphid S, Anan C, Thanawattano C, Deoisres S, Panyakaew P, Phokaewvarangkul O, Maytharakcheep S, Buranasrikul V, Prasertpan T. The rise of Parkinson’s disease is a global challenge, but efforts to tackle this must begin at a national level: a protocol for national digital screening and “eat, move, sleep” lifestyle interventions to prevent or slow the rise of non-communicable diseases in Thailand. Frontiers in neurology. 2024 May 13;15:1386608.
[2] Rusz J, Krack P, Tripoliti E. From prodromal stages to clinical trials: The promise of digital speech biomarkers in Parkinson’s disease. Neuroscience & Biobehavioral Reviews. 2024 Dec 1;167:105922.
To cite this abstract in AMA style:
N. Leabthong, J. Sringean, T. Laosombut, N. Jeh-Voh1, Z. Win, P. Rattanajun, S. Phumphid, C. Anan, J. Meesri, S. Wekhinhiran, O. Phokaewvarangkul, P. Panyakaew, S. Maytharakcheep, P. Jagota, R. Bhidayasiri. Real-World Mobile Voice Classification for Parkinson’s Disease Using Convolutional Neural Networks and Mel-Spectrogram Features [abstract]. Mov Disord. 2026; 41 (suppl 1). https://www.mdsabstracts.org/abstract/real-world-mobile-voice-classification-for-parkinsons-disease-using-convolutional-neural-networks-and-mel-spectrogram-features/. Accessed October 1, 2026.« Back to 2026 International Congress
MDS Abstracts - https://www.mdsabstracts.org/abstract/real-world-mobile-voice-classification-for-parkinsons-disease-using-convolutional-neural-networks-and-mel-spectrogram-features/


