Audit-Ready Healthcare Fraud Screening: Split-Safe Provider Aggregation and Explainable Boosted Risk Triage

Clicks: 1
ID: 312424
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #12 of 15 articles by views in Southern Journal of Computer Science

Most read Least read

Bar heights use a square-root scale.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Medical fraud and abnormal billing are often not clearly reflected in individual claim records, but rather in the cumulative abnormal behavior of the same service provider across multiple visits. Based on this characteristic, this paper defines the service provider, rather than a single claim, as the basic unit of risk screening and constructs a provider-level fraud screening process for auditing scenarios. Specifically, we first perform aggregation before data partitioning to minimize the risk of information leakage from the same provider across training and validation sets. Then, we train the LightGBM risk scoring model around audit-significant features such as claims volume, reimbursement, and out-of-pocket intensity, hospitalization duration statistics, duplicate claim characteristics, coding diversity, and beneficiary structure. To make the model output more suitable for actual review processes, further combine TreeSHAP interpretation, threshold scanning, and isotonic calibration, enabling the risk score to simultaneously serve priority ranking, manual review under capacity constraints, and clearer result interpretation. On the publicly available Healthcare Provider Fraud Detection dataset, based on provider-centric out-of-fold evaluation, the proposed method achieves good ranking performance, with an AUC of 0.939, an AUPRC of 0.699, and an F1 score of 0.666 at the selected threshold. The results also show that maximum hospitalization duration, reimbursement intensity, total claims volume, total out-of-pocket expenses, and beneficiary age structure are key risk signals. This provides a dense, interpretable, and auditable enactment method for beneficiary risk showing in health claims settings.
Reference Key
imported_1776990035_69eab753d58ed Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Basit Raza
Journal Southern Journal of Computer Science
Year 2026
DOI
DOI not found
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.