MOST FREQUENTLY ASKED QUESTIONS BY OLDER ADULTS IN GERIATRIC REHABILITATION: EVALUATING LARGE LANGUAGE MODELS AS A SOURCE OF INFORMATION


Sözlü U., GÜNAY S. M., Türe S. D., Özkazanç G.

Turk Geriatri Dergisi, cilt.29, sa.2, ss.191-202, 2026 (SCI-Expanded, SSCI, Scopus, TRDizin)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 29 Sayı: 2
  • Basım Tarihi: 2026
  • Doi Numarası: 10.29400/tjgeri.2026.491
  • Dergi Adı: Turk Geriatri Dergisi
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Social Sciences Citation Index (SSCI), Scopus, EMBASE, TR DİZİN (ULAKBİM), Academic Search Ultimate (EBSCO), Biomedical Reference Collection: Corporate Edition (EBSCO)
  • Sayfa Sayıları: ss.191-202
  • Anahtar Kelimeler: Artificial Intelligence, Geriatrics, Health Literacy, Patient Education as Topic, Rehabilitation
  • Bursa Uludağ Üniversitesi Adresli: Evet

Özet

Introduction: Older adults in geriatric rehabilitation are increasingly turning to internet-based resources and large language models for health-related information outside of clinical follow-up. The aim of this study was to compare the reliability, clinical accuracy, quality, usefulness, and readability of ChatGPT-5.2, Gemini 3, and DeepSeek V3.2 responses to patient questions regarding geriatric rehabilitation. Materials and Method: In this cross-sectional comparative content analysis, 24 predefined questions on geriatric rehabilitation were developed from YouTube comments, relevant literature, and clinical experience. Each question was submitted to three large language models under standardized conditions. Anonymized responses were independently evaluated by two experienced physiotherapists for reliability, clinical accuracy, quality, and usefulness, with disagreements resolved by consensus with a specialist physician. Readability was assessed using the Flesch Reading Ease scores. Results: No statistically significant difference was found among the three large language models in terms of reliability scores (p = 0.097). However, significant differences were observed for clinical accuracy, quality, usefulness, readability, and text characteristics (all p = 0.001). ChatGPT-5.2 and DeepSeek V3.2 showed the highest clinical accuracy scores, while ChatGPT-5.2 was superior in terms of quality and usefulness. For readability, ChatGPT-5.2 and Gemini 3 outperformed DeepSeek V3.2. Conclusion: Although all three large language models generally produced reliable content, ChatGPT-5.2 and DeepSeek V3.2 showed stronger clinical accuracy performance. Nevertheless, because of the risk of incorrect information being generated, the use of large language models by the older population should preferably be done under expert supervision.