Chongyi Yang

dblp:238/1979 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 DOME: A Domain-Orchestrated Multi-GPU Optical Network for Rack-Scale Systems
abstract
Modern data centers increasingly use multi-GPU systems for AI and high-performance computing, where growing data transfer demands lead to high energy consumption and performance bottlenecks in electrical networks. Optical interconnects offer compelling advantages to address these challenges, including high bandwidth, distance-independent latency, and better energy efficiency. This paper presents DOME, a rack-scale optical interconnection network that connects multiple GPUs using high-radix optical switches and extends optical interfaces into GPU packages close to memories and multiprocessors, forming distinct in-GPU and GPU-to-GPU network domains. To efficiently manage the paths across switches and domains, we develop a multi-switch arbitration scheme and a time-slotted path reservation scheme that quickly identifies the earliest time when the path is available in all network domains, reducing unnecessary reservation retries. Evaluations reveal that DOME achieves 14% speedup while maintaining comparable energy consumption compared to the state-of-the-art preemptive chain feedback control scheme.
Chongyi Yang, Bohan Hu, Yinyi Liu, Wei Zhang 0012, Jiang Xu 0001
ASP-DAC1
2026 FSR-GeMM: A Scalable FSR-Parallel Photonic Accelerator for Real-Valued GeMM Computing
abstract
Photonic computing is poised to revolutionize artificial intelligence (AI) acceleration by offering exceptional speed and energy efficiency for General Matrix Multiplication (GeMM). However, existing works on photonic tensor core architectures face significant challenges in managing real-valued and dynamic operands. Specifically, Mach-Zehnder interferometer (MZI) meshes require computationally intensive singular value decomposition (SVD) for matrix preprocessing, while microring resonator (MRR) weight banks are limited to non-negative operands, complicating operations with dual negative values. Additionally, coherent interference crossbars, although theoretically capable of supporting real-valued multiplication, struggle with fabrication complexities and sensitivity to environmental variations.To address these limitations, we propose FSR-GeMM schema, a scalable photonic accelerator that leverages free-spectral range (FSR) multiplexing. This architecture eliminates the need for SVD preprocessing, supports direct multiplication of two dynamic real-valued operands, and enhances reliability and scalability. Experimental results from a photonic-electronic prototype demonstrate that FSR-GeMM achieves up to 57× improvements in area efficiency and 13.8× gains in energy efficiency compared to existing photonic GeMM accelerators. Furthermore, it reduces energy consumption by 70% relative to MRR-based systems and achieves 21× speedup against leading photonic GeMM accelerator designs, highlighting its potential to advance practical and scalable AI acceleration.
Yinyi Liu, Minhang Xu, Chongyi Yang, Wei Zhang 0012, Jiang Xu 0001
DATE4
2025 BEAM: A Multi-Channel Optical Interconnect for Multi-GPU Systems
abstract
High-performance computing and AI applications necessitate high-bandwidth communication between GPUs. Traditional electrical interconnects for GPU-to-GPU communication face challenges over longer distances, including high power consumption, crosstalk noise, and signal loss. In contrast, optical interconnects excel in this domain, offering high bandwidth and consistent power dissipation over long distance. This paper proposes BEAM, a Bandwidth-Enhanced optical interconnect Architecture for Multi-GPU systems. BEAM extends electrical-optical interfaces into the GPU package, positioning them close to GPU compute logic and memory. Unlike existing single-channel approaches, each BEAM optical interface incorporates multiple parallel optical channels, further enhancing bandwidth. An arbitration scheme manages channel usage among data transfers. Evaluation on Rodinia benchmarks and LLM training kernels demonstrates that BEAM achieves a speedup of 1.14 – 1.9× and reduces energy consumption by 29 – 44% compared to the electrical-interconnected system and state-of-the-art schemes, while maintaining comparable chip area consumption.
Chongyi Yang, Bohan Hu, Yinyi Liu, Jiang Xu 0001
DATE1
2023 Adaptive Caching Policies for Chiplet Systems Based on Reinforcement Learning
abstract
Chiplet packaging becomes a popular solution to integrate more hardware components. However, shared memory access across chiplets suffers from high miss penalty due to long route latency and low bandwidth of inter-chiplet interconnects. We observe that the aggregated last-level cache (LLC) miss penalty takes approximately 35% of time on data access, and that the miss is dominated by coherence miss as a result of shared reads and writes from other LLCs. To address this problem, we propose a caching manager which speculatively enforces (or discards) LLC caching via online reinforcement learning. On every invalidated cacheline, the caching manager receives the cacheline access features, evaluates the current caching policy, and makes the next caching policy adaptively. Experimental evaluation justifies that the caching manager can reduce more than 10% coherence miss and offers a 3% speedup against a state-of-the-art cache coherence protocol.
Chongyi Yang, Xiaohang Wang 0001, Peng Liu 0016
ISCAS1