Guideline-Based Evaluation of Large Language Models in Psoriasis Treatment
Clicks: 9
ID: 329423
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
Emerging Content
2.4
/100
9 views
8 readers
AI Quality Assessment
Not analyzed
Readership in this journal
EmergingRanked #184 of 216 articles by views in clinical and experimental dermatology
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 216 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Abstract Background The therapeutic landscape for moderate-to-severe psoriasis has expanded rapidly with biologics and small molecules, increasing the complexity of clinical decision-making. Although large language models (LLMs) may support medical information synthesis, their reliability in high-stakes dermatologic management remains uncertain. Objectives To compare the guideline adherence, safety, and clinical decision quality of multiple LLMs with dermatology experts in systemic psoriasis management. Methods A 40-scenario question bank was developed from the Living EuroGuiDerm Guideline for the Systemic Treatment of Psoriasis Vulgaris and stratified by risk level. Five LLMs (ChatGPT-5.2, Claude Sonnet 4.5, Gemini 3.0 Pro, Grok 4.1 Thinking, and DeepSeek V3.2) were benchmarked against five board-certified dermatologists. Responses were scored for completeness and safety. Results Normalized completeness scores differed significantly among study groups (χ2 = 20.55, df = 5, P < 0.001, ε2 = 0.066). Gemini 3.0 Pro, DeepSeek V3.2, and Grok 4.1 Thinking scored significantly higher than the expert panel. This advantage was driven primarily by must-know items, for which hit rates differed significantly among groups (P = 0.003; 68% for the human benchmark vs 79–90% for LLMs). In low-risk scenarios, AI models outperformed experts (χ2 = 17.54, P = 0.004), whereas high-risk performance converged (χ2 = 6.14, P = 0.293). However, safety rankings reversed on critical error analysis (Q = 17.10, df = 5, P = 0.004): experts made 1 critical error, while LLMs made 5–9 each. Experts made no critical errors in high-risk scenarios, whereas all LLMs made at least one. Conclusions Current LLMs can match or exceed dermatologists in completeness of guideline-based systemic psoriasis recommendations, particularly in low-risk contexts. However, they remain more prone to critical safety errors in complex, high-stakes scenarios.
| Reference Key |
openalex_W7214057664
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Kuanyu Xia, Lang Min, Dan Jian |
| Journal | clinical and experimental dermatology |
| Year | 2026 |
| DOI |
10.1093/ced/llag417
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.