Automating the Comprehensive Complication Index With Artificial Intelligence: Evaluation of Quality and Clinical Potential

Clicks: 1
ID: 315723
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #289 of 432 articles by views in the british journal of surgery

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 432 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Background The Comprehensive Complication Index (CCI) is a validated, clinically established, and sensitive measure of postoperative morbidity but is underutilised due to the labor-intensive and error-prone nature of manual complication extraction and grading. Large language models (LLMs) can analyse free-text surgical discharge summaries and accurately apply the Clavien-Dindo Classification (CDC). However, their ability to reliably compute the more complex CCI – requiring identification of multiple events, correct severity grading and weighted aggregation – remains unclear. Aims To evaluate the feasibility and accuracy of contemporary LLMs in extracting postoperative complications and computing the CCI from unstructured surgical discharge summaries. Methods Six LLMs (ChatGPT-5.2, Claude Sonnet 4.5, DeepSeek-V3.2, Mistral AI (2024), Gemini 3 Flash and Llama 4) were assessed using a tiered validation framework. After conceptual testing, each model analysed 20 de-identified real-world surgical discharge summaries. A three-layered prompting framework (preprocessing, complication extraction and computation, and consistency checking) was compared with naïve end-to-end analysis. Agreement with expert reference CCI values was assessed using intraclass correlation coefficients (ICC) and Bland-Altman analysis. Results All models correctly defined the CCI concept. Naïve analysis showed moderate agreement with reference CCI values (ICC(3,2) = 0.87), with the observed disagreement being driven by missed complications and incorrect aggregation. Structured prompting improved agreement to excellent (ICC(3,1)=0.93) with minimal systematic bias. ChatGPT and Claude achieved full agreement with reference CCI values across all cases. Discrepancies in other LLMs were attributable to ambiguity in clinical documentation or inherent interpretive boundaries of the CCI framework, rather than computational errors. Conclusion LLMs can accurately extract complications, assign CDC grades, and compute CCI from unstructured free-text surgical discharge summaries when guided by structured prompting. This demonstrates that AI can reliably perform complex morbidity aggregation and may reduce clinical workload while improving standardization of surgical outcome reporting. Further large-scale validation is warranted.
Reference Key
openalex_W7163325641 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors P Borgas, C Nebiker, S Staubli
Journal the british journal of surgery
Year 2026
DOI
10.1093/bjs/znag055.007
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.