Evaluation of Aesthetic Outcomes Following Botulinum Toxin Treatment Using Multimodal Large Language Models: A Paired Before-and-After Analysis
Clicks: 1
ID: 317265
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
0.0
/100
1 views
0 readers
AI Quality Assessment
Not analyzed
Readership in this journal
Ranked #59 of 81 articles by views in Aesthetic Surgery Journal Open Forum
Most read
Least read
Bar heights use a square-root scale.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Abstract Background Multimodal large language models (MLLMs) are increasingly applied to visual assessment tasks, yet their ability to evaluate aesthetic treatment outcomes remains unclear. Objectives To assess whether contemporary MLLMs can identify treatment state and detect region-specific aesthetic improvement following botulinum toxin (BoNT) treatment. Methods In this observational study, 23 paired facial image cases (46 images; 460 model evaluations) were analyzed. Four MLLMs (GPT-5.4 Pro, Grok 4.1, Gemini 3.1 Pro, Claude Opus 4.6) performed five independent inference runs per case. Models identified the post-treatment image and assessed regional improvement (forehead, glabella, periorbital). Accuracy, sensitivity, specificity, balanced accuracy, Matthews correlation coefficient (MCC), and Fleiss’ κ were calculated descriptively. Performance was compared to majority-class baselines. Exploratory outputs included aesthetic scores and apparent age estimates. Results All models identified the post-treatment image (100% accuracy). Region-specific improvement detection frequently failed to exceed majority-class baselines (65.2–91.3%). Gemini 3.1 Pro showed the highest performance for forehead (74.8%) and glabella (63.5%), while no model reached the periorbital baseline. Inter-run reliability varied widely (κ −0.113 to 0.719). High reliability did not imply correctness. All models systematically overestimated improvement (62.6–94.8% of predictions exceeded ground truth). False positives exceeded false negatives in all 12 model–task combinations. Exploratory outputs indicated perceived rejuvenation. Conclusions MLLMs recognize the format of aesthetic change but not its clinical nuance. Bridging this gap requires more than improved accuracy: a coordinated agenda of domain-specific fine-tuning, expert-rater benchmarking, and structured outcome frameworks, supported by governance of training-data provenance and clear safeguards before any clinical or research deployment.
| Reference Key |
openalex_W7164664785
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Ibrahim Güler, Armin Kraus, Gerrit Grieb, Henrik Stelling |
| Journal | Aesthetic Surgery Journal Open Forum |
| Year | 2026 |
| DOI |
10.1093/asjof/ojag112
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.