Automating the Comprehensive Complication Index With Artificial Intelligence: Evaluation of Quality and Clinical Potential
Clicks: 1
ID: 315723
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
0.0
/100
1 views
0 readers
AI Quality Assessment
Not analyzed
Readership in this journal
Ranked #289 of 432 articles by views in the british journal of surgery
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 432 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Abstract Background The Comprehensive Complication Index (CCI) is a validated, clinically established, and sensitive measure of postoperative morbidity but is underutilised due to the labor-intensive and error-prone nature of manual complication extraction and grading. Large language models (LLMs) can analyse free-text surgical discharge summaries and accurately apply the Clavien-Dindo Classification (CDC). However, their ability to reliably compute the more complex CCI – requiring identification of multiple events, correct severity grading and weighted aggregation – remains unclear. Aims To evaluate the feasibility and accuracy of contemporary LLMs in extracting postoperative complications and computing the CCI from unstructured surgical discharge summaries. Methods Six LLMs (ChatGPT-5.2, Claude Sonnet 4.5, DeepSeek-V3.2, Mistral AI (2024), Gemini 3 Flash and Llama 4) were assessed using a tiered validation framework. After conceptual testing, each model analysed 20 de-identified real-world surgical discharge summaries. A three-layered prompting framework (preprocessing, complication extraction and computation, and consistency checking) was compared with naïve end-to-end analysis. Agreement with expert reference CCI values was assessed using intraclass correlation coefficients (ICC) and Bland-Altman analysis. Results All models correctly defined the CCI concept. Naïve analysis showed moderate agreement with reference CCI values (ICC(3,2) = 0.87), with the observed disagreement being driven by missed complications and incorrect aggregation. Structured prompting improved agreement to excellent (ICC(3,1)=0.93) with minimal systematic bias. ChatGPT and Claude achieved full agreement with reference CCI values across all cases. Discrepancies in other LLMs were attributable to ambiguity in clinical documentation or inherent interpretive boundaries of the CCI framework, rather than computational errors. Conclusion LLMs can accurately extract complications, assign CDC grades, and compute CCI from unstructured free-text surgical discharge summaries when guided by structured prompting. This demonstrates that AI can reliably perform complex morbidity aggregation and may reduce clinical workload while improving standardization of surgical outcome reporting. Further large-scale validation is warranted.
| Reference Key |
openalex_W7163325641
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | P Borgas, C Nebiker, S Staubli |
| Journal | the british journal of surgery |
| Year | 2026 |
| DOI |
10.1093/bjs/znag055.007
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.