Information matters, but cues are dangerous: a comparative evaluation of three multimodal AI chatbots in oral and maxillofacial radiology

Clicks: 1
ID: 325366
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #43 of 46 articles by views in dentomaxillofacial radiology

Most read Least read

Bar heights use a square-root scale.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Objectives To compare the diagnostic performance of three multimodal AI chatbots on oral and maxillofacial radiographic images and to examine how additional information, delivery mode, and user-suggested diagnoses affect accuracy. Methods Three AI chatbots (GPT-5.1, Gemini 3 Flash, Claude Opus 4.7) were tested on 90 cases comprising normal controls, osteomyelitis, and benign jaw lesions. Inputs were combinations of a panoramic image, a cropped panoramic image, an axial CBCT image or a text-based cue. Inputs were delivered all at once or sequentially. Accuracy was scored at category and specific-diagnosis levels using non-parametric tests with false-discovery-rate correction. Results With the panoramic image alone, accuracy for diseased cases was low (0–60%) but rose to as high as 38–92% in each model's best condition with added information. A cropped image was the most consistently beneficial additional visual input, whereas an axial CBCT image provided less improvement. GPT-5.1 recognised normal cases well but missed lesions, Gemini 3 was sensitive but less specific, and Claude 4.7 defaulted to benign diagnoses. Correct verbal cues increased accuracy, whereas a misleading cue caused decline. Gemini 3 accepted a false benign suggestion in 96% of cases it had initially classified as normal. For benign lesions, specific-diagnosis accuracy was almost half of category-level accuracy. Conclusions Accuracy was strongly affected by model choice, information type and user-suggested diagnoses, but no model reached a level that would be acceptable for clinical use. Advances in knowledge This multi-model comparison isolates the effects of information type, delivery mode, and sycophancy in oral and maxillofacial radiology.
Reference Key
openalex_W7203835141 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Chang‐Ki Min
Journal dentomaxillofacial radiology
Year 2026
DOI
10.1093/dmfr/twag061
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.