Information matters, but cues are dangerous: a comparative evaluation of three multimodal AI chatbots in oral and maxillofacial radiology
Clicks: 1
ID: 325366
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
0.0
/100
1 views
0 readers
AI Quality Assessment
Not analyzed
Readership in this journal
Ranked #43 of 46 articles by views in dentomaxillofacial radiology
Most read
Least read
Bar heights use a square-root scale.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Abstract Objectives To compare the diagnostic performance of three multimodal AI chatbots on oral and maxillofacial radiographic images and to examine how additional information, delivery mode, and user-suggested diagnoses affect accuracy. Methods Three AI chatbots (GPT-5.1, Gemini 3 Flash, Claude Opus 4.7) were tested on 90 cases comprising normal controls, osteomyelitis, and benign jaw lesions. Inputs were combinations of a panoramic image, a cropped panoramic image, an axial CBCT image or a text-based cue. Inputs were delivered all at once or sequentially. Accuracy was scored at category and specific-diagnosis levels using non-parametric tests with false-discovery-rate correction. Results With the panoramic image alone, accuracy for diseased cases was low (0–60%) but rose to as high as 38–92% in each model's best condition with added information. A cropped image was the most consistently beneficial additional visual input, whereas an axial CBCT image provided less improvement. GPT-5.1 recognised normal cases well but missed lesions, Gemini 3 was sensitive but less specific, and Claude 4.7 defaulted to benign diagnoses. Correct verbal cues increased accuracy, whereas a misleading cue caused decline. Gemini 3 accepted a false benign suggestion in 96% of cases it had initially classified as normal. For benign lesions, specific-diagnosis accuracy was almost half of category-level accuracy. Conclusions Accuracy was strongly affected by model choice, information type and user-suggested diagnoses, but no model reached a level that would be acceptable for clinical use. Advances in knowledge This multi-model comparison isolates the effects of information type, delivery mode, and sycophancy in oral and maxillofacial radiology.
| Reference Key |
openalex_W7203835141
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Chang‐Ki Min |
| Journal | dentomaxillofacial radiology |
| Year | 2026 |
| DOI |
10.1093/dmfr/twag061
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.