Cross-Dataset Unified Vision Transformer Model for Diabetic Retinopathy Detection
Clicks: 1
ID: 312665
2025
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
0.0
/100
1 views
0 readers
AI Quality Assessment
Not analyzed
Readership in this journal
Ranked #303 of 705 articles by views in Journal of Computing & Biomedical Informatics
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 705 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Diabetic retinopathy (DR) is a major cause of blindness around the world that can be prevented; it is caused by prolonged hyperglycemia that leads to damage to retinal vasculature. Early detection of DR, and good DR grading are important to ensure timely clinical intervention and improved outcomes for patients. Machine learning and convolutional neural network (CNN)-based approaches have held promise for DR screening, but they are often limited in their ability to capture fine-grained lesions or long-range dependencies because of limited receptive fields and insufficient modeling of global context.Vision Transformers, or ViTs, lever self-attention mechanisms to capture global relationships across retinal structures. This is a Review of ViT based frameworks on grading DR. The review concentrates on studies using two of the most widely applied and diverse benchmarks between EyePACS and APTOS; expert annotated fundus images assessment. The review covers a variety of recent advances such as hybrid CNN-ViT architectures, lesion-aware transformer modules, multi-scale feature aggregation, and federated learning strategies for privacy-preserving medical image analysis. It then highlights the role of interpretable attention maps in improving clinical trust and decision transparency. This is also the review of the remaining challenges in DR grading; it includes extreme class imbalance in DR severity levels, high computational costs of transformer models, and the demand for a powerful and robust explainability technique to favor clinical adoption. In connecting current achievements, unsolved issues, and new directions of research, this review endeavors to orient researchers and practitioners towards designing efficient, generalizable, and clinically relevant ViT-based DR detection and grading systems. In general, it shows how much transformer-driven approaches have the potential to revolutionize automated ophthalmic diagnosis and enhance the global diabetic eye care workflow.
| Reference Key |
imported_1777055893_69ebb8953e420
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Shruti Yagnik |
| Journal | Journal of Computing & Biomedical Informatics |
| Year | 2025 |
| DOI |
10.56979/1001/2025/1148
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.