Objective: To generate a list of clinical features representative of Parkinson’s disease from unstructured clinical notes and to evaluate the performance of different large language models (LLMs).
Background: Clinical notes of patients with Parkinson’s disease contain critical information for accurate diagnosis, treatment planning, and prognosis prediction. However, because this information is often complex and individualized, clinical notes are typically written in an unstructured narrative format. As a result, identifying relevant clinical features from clinical notes is challenging and time-consuming.
Method: Admission and discharge notes from inpatients and initial and follow-up visit notes from outpatients at a tertiary hospital were collected between January 2005 and January 2025. From 1,559 patients, 100 patients stratified by the year of their first hospital visit, were randomly selected. Longitudinal clinical notes from these patients were used as model inputs for seven different LLMs: Qwen3-32B, Qwen3-30B-A3B-Instruct-2507, Magistral-Small-2506, NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, EXAONE-4.0-32B, medgemma-27b-text-it, and gpt-oss-20b. After iterative prompt engineering and qualitative evaluation of model outputs by a movement specialist, the number of unique features suggested by each model was counted, along with the hallucination rate. Hallucinations were identified by verifying whether the clinical record referenced by the model for each suggested feature was actually present in the input.
Results: Common clinical features suggested by the LLMs included the Hoehn and Yahr scale, bradykinesia and rigidity of upper limb, constipation, visual hallucination, wearing off, and dyskinesia, suggesting that these features were frequently documented in the clinical notes. Qwen3-32B, Qwen3-30B-A3B-Instruct-2507, Magistral-Small-2506, NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, EXAONE-4.0-32B, medgemma-27b-text-it, and gpt-oss-20b generated 758, 776, 460, 522, 515, 939, and 926 unique features, respectively. The hallucination rates were 9.7%, 6.8%, 16.5%, 25%, 26.4%, 10.3%, and 7.2%, respectively.
Conclusion: LLMs can generate lists of relevant clinical features from clinical records, although hallucinations occur and performance varies across models. The validity and clinical relevance of these features will be evaluated by specialists in future work.
To cite this abstract in AMA style:
GY. Lee, HY. Kwon, S. Jo, M. Choi, N. Kim, SJ. Chung. Large Language Models for Clinical Feature Suggestion from Parkinson’s Disease Clinical Notes [abstract]. Mov Disord. 2026; 41 (suppl 1). https://www.mdsabstracts.org/abstract/large-language-models-for-clinical-feature-suggestion-from-parkinsons-disease-clinical-notes/. Accessed October 1, 2026.« Back to 2026 International Congress
MDS Abstracts - https://www.mdsabstracts.org/abstract/large-language-models-for-clinical-feature-suggestion-from-parkinsons-disease-clinical-notes/
