Evaluating the reliability of AI-based chatbots on dental treatment of children with systemic diseases
31. Uluslararası Türk Pedodonti Derneği Kongresi, Antalya, Turkey, 4 - 07 October 2025, pp.81, (Summary Text)
- Publication Type: Conference Paper / Summary Text
- City: Antalya
- Country: Turkey
- Page Numbers: pp.81
- Bursa Uludag University Affiliated: Yes
Abstract
AIM: To evaluate the accuracy and day-to-day consistency of responses given by different AI-based chatbots to questions on dental management in
common pediatric systemic conditions.
METHODS: Fifteen questions were derived from authoritative guidelines and textbooks covering systemic conditions frequently seen in childhood
(asthma, diabetes, epilepsy, congenital heart disease, and hematologic disorders). The questions were asked to four chatbots (ChatGPT-5.0, Chat-
GPT-4.0, Google Gemini, and Microsoft Bing Copilot) on two separate days. Each response was compared with reference answers from the source
materials. For each chatbot, accuracy (%) was calculated, and repeatability (identical answer to the same question across days) was assessed.
Results were summarized descriptively.
RESULTS: Across both days, 80 of 120 responses were correct (66.7%). Day 1 accuracies were: ChatGPT-5.0, 73.3% (11/15); ChatGPT-4.0, 66.7%
(10/15); Copilot, 66.7% (10/15); Gemini, 53.3% (8/15). On Day 2, performance was more even: ChatGPT-4.0 reached 73.3% (11/15), while ChatGPT-5.0,
Gemini, and Copilot each achieved 66.7% (10/15). Answer repetition rates between days were 93.3% (14/15) for ChatGPT-5.0, 86.7% (13/15) for Gemini,
80.0% (12/15) for ChatGPT-4.0, and 73.3% (11/15) for Copilot; the greatest change occurred with Copilot (4/15), and the least with ChatGPT-5.0
(1/15). Pooled accuracies were: ChatGPT-4.0, 70.0% (21/30); ChatGPT-5.0, 70.0% (21/30); Copilot, 66.7% (20/30); Gemini, 60.0% (18/30).
CONCLUSIONS: AI-based chatbots may aid information retrieval for dental care in children with systemic diseases; however, variability in accuracy and
consistency limits their use as stand-alone, reliable sources. Clinical decision-making should account for these limitations.
AMAÇ: Bu çalışmanın amacı, çocuklarda sık görülen sistemik hastalık durumlarında dental yaklaşım ile ilgili sorulara farklı yapay zekâ tabanlı sohbet
botlarının verdiği yanıtların doğruluk ve tutarlılığını değerlendirmektir.
YÖNTEM: Çocukluk çağında sık karşılaşılan astım, diyabet, epilepsi, kardiyolojik problemler ve hematolojik bozukluklar gibi sistemik durumlara ilişkin
çeşitli rehberlerden ve kaynak kitaplardan 15 soru hazırlandı. Sorular dört farklı yapay zekâ tabanlı sohbet botuna (ChatGPT 5, ChatGPT 4.0, Google
Gemini, Microsoft Bing Copilot), iki farklı günde yöneltildi. Verilen yanıtlar, rehber ve kaynak kitaplardan alınan doğru yanıtlarla karşılaştırıldı. Yanıtlar
kaydedilerek her yapay zekâ tabanlı sohbet botu için doğru cevap yüzdesi hesaplandı.
BULGULAR: Farklı iki günde 4 sohbet robotuna yöneltilen toplam 120 sorunun 80’i doğru yanıtlandı (%66,7). 1. günde doğruluk oranları ChatGPT-5.0:
%73,3 (11/15), ChatGPT-4.0: %66,7 (10/15), Copilot: %66,7 (10/15) ve Gemini: %53,3 (8/15) olarak kaydedildi. 2. günde dağılım ilk güne göre daha
dengeliydi; ChatGPT-4.0 %73,3 (11/15) ile en yüksek orana ulaşırken ChatGPT-5.0, Gemini ve Copilot’un her biri %66,7 (10/15) doğrulukta seyretti.
Gün 1–Gün 2 karşılaştırmasında yanıt tekrar oranı (aynı soruya aynı yanıt) ChatGPT-5.0’da %93,3 (14/15), Gemini’de %86,7 (13/15), ChatGPT-4.0’da
%80,0 (12/15) ve Copilot’ta %73,3 (11/15) olarak hesaplandı; en fazla yanıt değişimi Copilot’ta (4/15), en az değişim ChatGPT-5.0’da (1/15) gözlendi.
İki gün bir arada değerlendirildiğinde toplam doğruluk oranlarının ChatGPT-4.0: %70,0 (21/30), ChatGPT-5.0: %70,0 (21/30), Copilot: %66,7 (20/30) ve
Gemini: %60,0 (18/30) şeklinde olduğu görüldü.
SONUÇLAR: Çocuklarda sistemik hastalıklarla ilgili bilgilerin değerlendirilmesinde yapay zekâ uygulamaları potansiyel bir destek aracı olabilir. Ancak
doğruluk ve uyum düzeylerindeki farklılıklar, bu teknolojilerin tek başına güvenilir bir bilgi kaynağı olarak kullanılmasını sınırlamaktadır.