Evaluation of Aesthetic Outcomes Following Botulinum Toxin Treatment Using Multimodal Large Language Models: A Paired Before-and-After Analysis

Clicks: 1
ID: 317265
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #59 of 81 articles by views in Aesthetic Surgery Journal Open Forum

Most read Least read

Bar heights use a square-root scale.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Background Multimodal large language models (MLLMs) are increasingly applied to visual assessment tasks, yet their ability to evaluate aesthetic treatment outcomes remains unclear. Objectives To assess whether contemporary MLLMs can identify treatment state and detect region-specific aesthetic improvement following botulinum toxin (BoNT) treatment. Methods In this observational study, 23 paired facial image cases (46 images; 460 model evaluations) were analyzed. Four MLLMs (GPT-5.4 Pro, Grok 4.1, Gemini 3.1 Pro, Claude Opus 4.6) performed five independent inference runs per case. Models identified the post-treatment image and assessed regional improvement (forehead, glabella, periorbital). Accuracy, sensitivity, specificity, balanced accuracy, Matthews correlation coefficient (MCC), and Fleiss’ κ were calculated descriptively. Performance was compared to majority-class baselines. Exploratory outputs included aesthetic scores and apparent age estimates. Results All models identified the post-treatment image (100% accuracy). Region-specific improvement detection frequently failed to exceed majority-class baselines (65.2–91.3%). Gemini 3.1 Pro showed the highest performance for forehead (74.8%) and glabella (63.5%), while no model reached the periorbital baseline. Inter-run reliability varied widely (κ −0.113 to 0.719). High reliability did not imply correctness. All models systematically overestimated improvement (62.6–94.8% of predictions exceeded ground truth). False positives exceeded false negatives in all 12 model–task combinations. Exploratory outputs indicated perceived rejuvenation. Conclusions MLLMs recognize the format of aesthetic change but not its clinical nuance. Bridging this gap requires more than improved accuracy: a coordinated agenda of domain-specific fine-tuning, expert-rater benchmarking, and structured outcome frameworks, supported by governance of training-data provenance and clear safeguards before any clinical or research deployment.
Reference Key
openalex_W7164664785 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Ibrahim Güler, Armin Kraus, Gerrit Grieb, Henrik Stelling
Journal Aesthetic Surgery Journal Open Forum
Year 2026
DOI
10.1093/asjof/ojag112
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.