On-chip large-scale all-optical interconnect for ultra-low-latency deep neural network inference

Clicks: 1
ID: 314091
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #196 of 280 articles by views in national science review

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 280 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract As artificial intelligence (AI) models continue to expand rapidly and evolve at an accelerating pace, the sustainability of scaling law faces challenges as computational resources cannot expand indefinitely. A more promising strategy is to use high-performance interconnections to compose multiple modest-capacity computing chips into a high-efficiency system, enabling comparable capability to that of much larger resource-intensive systems. However, an effective on-chip communication and network that provides high bandwidth, low latency, and large-scale parallelism to accelerate distributed computation is still lacking. Here, we address this challenge by proposing an on-chip, all-optical-interconnect-based hybrid optoelectronic distributed computing system, achieving two orders of magnitude lower inference latency than a single graphics processing unit (GPU) while using only one-ninth of its computational resources. The 400-gigabits per second (Gbps) silicon photonic transceiver chips can provide high-speed and error-free optical input/output (I/O) for processing chips, and an ultra-low-loss (≤ 5 dB at 1300 nm), non-blocking 16×16 optical switch chip can further scale out the all-optical network for massive and flexible interconnection. As a validation, we configure the system to realize optical pipeline parallelism (OPP) and execute a five-layer convolutional neural network (CNN) denoising model, which processes a total of 1,000 images of 32,768 bits each in only 105.16 µs. These results convincingly illustrate a transformative pathway toward future high-efficiency computing architectures.
Reference Key
openalex_W7161719461 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Zihan Tao, Yan Zhou, Weizhen Yu, Huajin Chang, Hao Wu, Ailian Cheng, Qian Hong, Yonglin Tang, Xiaolin Bian, Lingwei Meng, Li Li, Gaoang Shen, Cheng Zhang, Yongguang Huang, Wenhua Wang, Haowen Shu, Xingjun Wang
Journal national science review
Year 2026
DOI
10.1093/nsr/nwag282
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.