Multimodal Büyük Dil Modellerinin BT’de Bone-RADS Klasifikasyonundaki Tanısal Performansı


Creative Commons License

Kaya H. E., ATAŞ A. E.

Uludağ Üniversitesi Tıp Fakültesi Dergisi, cilt.52, sa.0, 2026 (TRDizin)

Özet

Çalışmamızın amacı multimodal büyük dil modellerinin (MBDM) BT görüntülerinde tespit edilen soliter kemik lezyonlarına Bone-RADS kategorileri atamadaki performanslarını değerlendirmektir. Hastanemizin PACS’ı soliter kemik lezyonu içeren BT tetkikleri için taranmıştır (Ağustos 2024-Ağustos 2025). Her lezyon için bir kas-iskelet radyoloğu tarafından kitleyi en iyi temsil eden bir kesit seçilmiş ve lezyonlara birer referans Bone-RADS skoru atanmıştır. Daha sonra bir abdominal radyolog, ChatGPT 5 ve Gemini 2.5 Pro aynı vakaları kategorilemiştir. Doğruluk, doğru şekilde kategorize edilen Bone-RADS 1 ve 4 vakaları olarak tanımlanmış ve McNemar testi kullanılarak karşılaştırılmıştır. Referansla uyum, ağırlıklı Cohen κ katsayısı kullanılarak değerlendirilmiş ve bootstrap yöntemi ile karşılaştırılmıştır. Referans kategorileri şu şekilde belirlenmiştir: Bone-RADS 1, n=23; 2, n=4; 3, n=0; 4, n=23. Doğruluk, radyolog için %84,8 (39/46), Gemini için %78,3 (36/46) ve ChatGPT için %65,2 (30/46) olarak bulunmuştur. Radyoloğun, ChatGPT'den daha iyi performans gösterdiği (p=0,012); radyolog ile Gemini (p=0,604) ve Gemini ile ChatGPT (p=0,360) arasındaki farkların anlamlı olmadığı görülmüştür. Radyolog, referans standardı ile en yüksek uyumu elde etmiş (κ = 0,715, %95 GA: [0,543-0,887]), bunu Gemini (κ = 0,542, %95 GA: [0,313-0,770]) ve ChatGPT (κ = 0,292, %95 GA: [0,104-0,479]) izlemiştir. Bootstrap ile yapılan karşılaştırmalar, radyoloğun κ değerinin ChatGPT'den daha yüksek olduğunu göstermiş (%95 GA: 0,140-0,675), ancak radyolog ile Gemini (%95 GA: −0,113-0,434) ve Gemini ile ChatGPT (%95 GA: −0,041-0,522) arasındaki fark anlamlı bulunmamıştır. Sonuç olarak genel amaçlı MBDM'ler henüz Bone-RADS kategorizasyonu için eğitimli radyologların yerini tutabilecek durumda görünmemekle beraber bu modellerin günlük pratikte radyologlara yardımcı olabileceği düşünülmektedir.
The aim of the study is to assess the performance of multimodal large language models (MLLMs) in assigning Bone-RADS categories to bone lesions identified on CT images. An MSK radiologist selected one representative slice for 50 bone lesions seen on CT studies and assigned reference Bone-RADS categories using clinical records. Three raters categorized each case: an abdominal radiologist, OpenAI ChatGPT 5, and Google Gemini 2.5 Pro. Accuracy was defined as the correctly labeled Bone-RADS 1 and 4 cases and compared using McNemar test. Agreement with the reference was assessed using weighted Cohen’s κ with 95% CIs; pairwise κ differences were tested via bootstrap. Reference categories were Bone-RADS 1, n=23; 2, n=4; 3, n=0; 4, n=23. Accuracy was 84.8% (39/46) for the radiologist, 78.3% (36/46) for Gemini, and 65.2% (30/46) for ChatGPT. The radiologist outperformed ChatGPT (p=0.012); differences between the radiologist vs Gemini (p=0.604) and Gemini vs ChatGPT (p=0.360) were not significant. The radiologist achieved the highest agreement with the reference standard (κ = 0.715, 95% CI: [0.543-0.887]), followed by Gemini (κ = 0.542, 95% CI: [0.313-0.770]) and ChatGPT (κ = 0.292, 95% CI: [0.104-0.479]). Bootstrap comparisons showed that the radiologist’s κ was higher than ChatGPT’s (95% CI for difference, 0.140-0.675), while radiologist vs Gemini (−0.113-0.434) and Gemini vs ChatGPT (−0.041-0.522) were not significant. In conclusion, general-purpose MLLMs cannot yet replace trained radiologists for Bone-RADS classification, though they may still aid routine clinical practice.