Multimodal Large Language Models and Dental Students in Radiographic Landmark Identification: A Novel Grid-Based Assessment

Clicks: 3
ID: 325425
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Emerging

Ranked #35 of 46 articles by views in dentomaxillofacial radiology

Most read Least read

Bar heights use a square-root scale.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
OBJECTIVES: To introduce a novel grid-based coordinate framework for evaluating spatial landmark localisation competence in multimodal large language models (MLLMs) applied to dental radiography, and to benchmark two MLLMs against dental student consensus. METHODS: GPT-5.4 and Gemini 3.1 Pro were evaluated on 900 queries spanning 200 dental radiographs (100 panoramic, 50 periapical, 50 cephalometric) and 12 landmarks (9 point, 3 area) under zero-shot and guided prompting across three repetitions (10,800 API calls). Performance was scored against a two-rater adjudicated ground truth and a team-adjudicated fourth-year dental student consensus (n = 40). Metrics included Euclidean distance, the successful detection rate, and the Jaccard index. RESULTS: GPT-5.4 matched student-level performance on cephalometric landmarks but was substantially outperformed on periapical and panoramic tasks. Gemini 3.1 Pro outperformed GPT-5.4 on panoramic and periapical landmarks, while GPT-5.4 retained an advantage on cephalometric points. Guided prompting produced heterogeneous effects, improving localisation of some landmarks while causing severe regression on others; most notably, GPT-5.4 directed the majority of lower-left canine apex predictions to the incorrect lower-right molar region, a prompt-resistant mislocalisation pattern absent in Gemini 3.1 Pro. Both models remained inferior to the student consensus on all periapical and panoramic comparisons, with Gemini 3.1 Pro approaching parity only on cephalometric points under guided prompting. CONCLUSION: Current MLLMs demonstrate modality-specific spatial competence insufficient for autonomous clinical deployment. The grid-based framework provides a reproducible, modality-agnostic benchmark for tracking MLLM spatial reasoning across future model generations.
Reference Key
openalex_W7203735537 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Mehmet Egemen Aydemir, Burcu Sayin, Mehmet Oğuz Borahan, Can Günel, Kadir Cem, Mehmet Ali Gül
Journal dentomaxillofacial radiology
Year 2026
DOI
10.1093/dmfr/twag060
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.