UMLS-Augmented PubMedBERT for Clinical Note Diagnosis Classification: A Comparative Study

Clicks: 1
ID: 312684
2025
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #136 of 705 articles by views in Journal of Computing & Biomedical Informatics

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 705 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Transformer-based language models have already demonstrated a good performance in the domain of biomedical text classification, but, the majority of the existing methodology uses unstructured textual data only, thereby, evident deficiency of explicit incorporation of curated biomedical knowledge. The present research evaluates the ways in which varying schemes of integration of Unified Medical Language System (UMLS) concepts can be applied to enhance PubMedBERT functioning in terms of the task to classify clinical note diagnosis. Three model configurations are compared with the publicly available PMC-Patients corpus, (i) a baseline PubMedBERT model that is trained just using the raw clinical notes, (ii) a PubMedBERT model that is trained using the raw clinical notes, but additionally augmented with unfiltered UMLS Concept Unique Identifiers (CUIs), and (iii) a PubMedBERT model that is trained using the raw clinical notes, but further augmented with NER-filtered UMLS concepts. They are trained using the weighted cross-entropy loss to address the issue of class imbalance and evaluated using the assistance of accuracy, macro and weighted F1-scores, per-class models, multiple-seed experiments, and bootstrap confidence intervals as well as paired samples t-tests. The findings imply that naive injection of the complete system of UMLS concepts has a negative impact on the performance that implies that injection of ontologies by unedits injects undesired noise. Active selection of concepts filtered by NER to the UMLS, on the other hand, leads to a gradual, but steady increase in performance in concept classification, with an accuracy of 98.8 percent and a massive improvement in minority classes, such as asthma and heart disease. The latter demonstrate that biomedical knowledge may be handy in enriching performance of transformer-based clinical text classification as it is presented in an intentional and systematic way, and because the enhancement of the performance is accompanied by the increase in computational costs.
Reference Key
imported_1777056009_69ebb909acd7f Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Ghulam Mustafa
Journal Journal of Computing & Biomedical Informatics
Year 2025
DOI
10.56979/1001/2025/1175
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.