Guideline-Based Evaluation of Large Language Models in Psoriasis Treatment

Clicks: 9
ID: 329423
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Emerging

Ranked #184 of 216 articles by views in clinical and experimental dermatology

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 216 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Background The therapeutic landscape for moderate-to-severe psoriasis has expanded rapidly with biologics and small molecules, increasing the complexity of clinical decision-making. Although large language models (LLMs) may support medical information synthesis, their reliability in high-stakes dermatologic management remains uncertain. Objectives To compare the guideline adherence, safety, and clinical decision quality of multiple LLMs with dermatology experts in systemic psoriasis management. Methods A 40-scenario question bank was developed from the Living EuroGuiDerm Guideline for the Systemic Treatment of Psoriasis Vulgaris and stratified by risk level. Five LLMs (ChatGPT-5.2, Claude Sonnet 4.5, Gemini 3.0 Pro, Grok 4.1 Thinking, and DeepSeek V3.2) were benchmarked against five board-certified dermatologists. Responses were scored for completeness and safety. Results Normalized completeness scores differed significantly among study groups (χ2 = 20.55, df = 5, P < 0.001, ε2 = 0.066). Gemini 3.0 Pro, DeepSeek V3.2, and Grok 4.1 Thinking scored significantly higher than the expert panel. This advantage was driven primarily by must-know items, for which hit rates differed significantly among groups (P = 0.003; 68% for the human benchmark vs 79–90% for LLMs). In low-risk scenarios, AI models outperformed experts (χ2 = 17.54, P = 0.004), whereas high-risk performance converged (χ2 = 6.14, P = 0.293). However, safety rankings reversed on critical error analysis (Q = 17.10, df = 5, P = 0.004): experts made 1 critical error, while LLMs made 5–9 each. Experts made no critical errors in high-risk scenarios, whereas all LLMs made at least one. Conclusions Current LLMs can match or exceed dermatologists in completeness of guideline-based systemic psoriasis recommendations, particularly in low-risk contexts. However, they remain more prone to critical safety errors in complex, high-stakes scenarios.
Reference Key
openalex_W7214057664 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Kuanyu Xia, Lang Min, Dan Jian
Journal clinical and experimental dermatology
Year 2026
DOI
10.1093/ced/llag417
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.