VLDB 2026 Research / reviewers in the wild / expert
Yinyi Liu
dblp:316/6614
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0001-9144-4376ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 4 first-author · 15 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Full-Stack System Design and Prototyping for Fully Programmable Electronic-Photonic Neurocomputing
Yinyi Liu, Bohan Hu, Wei Zhang 0012, Jiang Xu 0001 |
ASP-DAC | 1 |
| 2026 | DOME: A Domain-Orchestrated Multi-GPU Optical Network for Rack-Scale SystemsabstractModern data centers increasingly use multi-GPU systems for AI and high-performance computing, where growing data transfer demands lead to high energy consumption and performance bottlenecks in electrical networks. Optical interconnects offer compelling advantages to address these challenges, including high bandwidth, distance-independent latency, and better energy efficiency. This paper presents DOME, a rack-scale optical interconnection network that connects multiple GPUs using high-radix optical switches and extends optical interfaces into GPU packages close to memories and multiprocessors, forming distinct in-GPU and GPU-to-GPU network domains. To efficiently manage the paths across switches and domains, we develop a multi-switch arbitration scheme and a time-slotted path reservation scheme that quickly identifies the earliest time when the path is available in all network domains, reducing unnecessary reservation retries. Evaluations reveal that DOME achieves 14% speedup while maintaining comparable energy consumption compared to the state-of-the-art preemptive chain feedback control scheme. Chongyi Yang, Bohan Hu, Yinyi Liu, Wei Zhang 0012, Jiang Xu 0001 |
ASP-DAC | 4 |
| 2026 | FSR-GeMM: A Scalable FSR-Parallel Photonic Accelerator for Real-Valued GeMM ComputingabstractPhotonic computing is poised to revolutionize artificial intelligence (AI) acceleration by offering exceptional speed and energy efficiency for General Matrix Multiplication (GeMM). However, existing works on photonic tensor core architectures face significant challenges in managing real-valued and dynamic operands. Specifically, Mach-Zehnder interferometer (MZI) meshes require computationally intensive singular value decomposition (SVD) for matrix preprocessing, while microring resonator (MRR) weight banks are limited to non-negative operands, complicating operations with dual negative values. Additionally, coherent interference crossbars, although theoretically capable of supporting real-valued multiplication, struggle with fabrication complexities and sensitivity to environmental variations.To address these limitations, we propose FSR-GeMM schema, a scalable photonic accelerator that leverages free-spectral range (FSR) multiplexing. This architecture eliminates the need for SVD preprocessing, supports direct multiplication of two dynamic real-valued operands, and enhances reliability and scalability. Experimental results from a photonic-electronic prototype demonstrate that FSR-GeMM achieves up to 57× improvements in area efficiency and 13.8× gains in energy efficiency compared to existing photonic GeMM accelerators. Furthermore, it reduces energy consumption by 70% relative to MRR-based systems and achieves 21× speedup against leading photonic GeMM accelerator designs, highlighting its potential to advance practical and scalable AI acceleration. Yinyi Liu, Minhang Xu, Chongyi Yang, Wei Zhang 0012, Jiang Xu 0001 |
DATE | 2 |
| 2025 | Invited paper: SPICE-Compatible Modeling and Design for Electronic-Photonic Integrated CircuitsabstractElectronic-photonic integrated circuit (EPIC) technologies are revolutionizing computing systems by improving their performance and energy efficiency. However, simulating EPIC is challenging and time consuming. In this paper, we propose the physics-informed neural network (PINN) based SPICE-compatible modeling method for EPICs. Experimental results show our method can speed up EPIC simulation by more than 100 times on average compared to FDTD method. Yinyi Liu, Ngai Wong 0001, Jiang Xu 0001 |
ASP-DAC | 2 |
| 2025 | PICELF: An Automatic Electronic Layer Layout Generation Framework for Photonic Integrated CircuitsabstractIn recent years, the advent of photonic integrated circuits (PICs) has demonstrated great prospects and applications to address critical issues such as limited bandwidth, high latency, and high power consumption in data-intensive systems. However, the field of physical design automation for PICs remains in its infancy, with a notable gap in electronic layer layout design tools. Current research on PIC physical design automation primarily focuses on optical layer layouts, often overlooking the equally crucial electronic layer layouts. Although well-established for conventional integrated circuits (ICs), existing EDA tools are inadequately adapted for PICs due to their unique characteristics and constraints. As PICs grow in integration density and size, traditional manual-based design methods become increasingly inefficient and sub-optimal, potentially compromising overall PIC performance. To address this challenge, we propose PICELF, the first framework in the literature for automatic PIC electronic layer layout generation. Our framework comprises a nonlinear binary programming (NBP)-based netlist generator with scalability optimization and a two-stage router featuring initial parallel routing followed by post-routing optimization. We validate our framework's effectiveness and efficiency using a real PIC chip benchmark established by us. Experimental results demonstrate that our method can efficiently generate high-quality PIC electronic layer layouts and satisfy all design rules, within reasonable CPU times, while related existing methods are not applicable. Yinyi Liu, Jiang Xu 0001 |
DATE | 2 |
| 2025 | BEAM: A Multi-Channel Optical Interconnect for Multi-GPU SystemsabstractHigh-performance computing and AI applications necessitate high-bandwidth communication between GPUs. Traditional electrical interconnects for GPU-to-GPU communication face challenges over longer distances, including high power consumption, crosstalk noise, and signal loss. In contrast, optical interconnects excel in this domain, offering high bandwidth and consistent power dissipation over long distance. This paper proposes BEAM, a Bandwidth-Enhanced optical interconnect Architecture for Multi-GPU systems. BEAM extends electrical-optical interfaces into the GPU package, positioning them close to GPU compute logic and memory. Unlike existing single-channel approaches, each BEAM optical interface incorporates multiple parallel optical channels, further enhancing bandwidth. An arbitration scheme manages channel usage among data transfers. Evaluation on Rodinia benchmarks and LLM training kernels demonstrates that BEAM achieves a speedup of 1.14 – 1.9× and reduces energy consumption by 29 – 44% compared to the electrical-interconnected system and state-of-the-art schemes, while maintaining comparable chip area consumption. Chongyi Yang, Bohan Hu, Yinyi Liu, Jiang Xu 0001 |
DATE | 4 |
| 2024 | PhotonNTT: Energy-Efficient Parallel Photonic Number Theoretic Transform AcceleratorabstractFully homomorphic encryption (FHE) presents a promising opportunity to remove privacy barriers in various scenarios including cloud computing and secure database search, by enabling computation on encrypted data. However, integrating FHE with real-world applications remains challenging due to its significant computational overhead. In the FHE scheme, Number Theoretic Transform (NTT) consumes the primary computing resources and has great potential for acceleration. For the first time, we present a photonic NTT accelerator, PhotonNTT, with high energy efficiency and parallelism to address the above challenge. Our approach involves formulating the NTT into matrix-vector multiplication (MVM) operations and mapping the data flow into parallel photonic MVM units. A dedicated data mapping scheme is proposed to introduce free spectral range (FSR) and distributed RAM design into the system, which enables a high bit-wise parallelism level. The system's reliability is validated through the Monte-Carlo BER analysis. The experimen-tal evaluation shows that the proposed architecture outperforms SOTA CiM-based NTT accelerators with an improvement of 50x in throughput and 63x improvement in energy efficiency. Yinyi Liu, Chengeng Li, Shixi Chen, Fengshi Tian, Wei Zhang 0012, Jiang Xu 0001 |
DATE | 4 |
| 2024 | PCC: An End-to-End Compilation Framework for Neural Networks on Photonic-Electronic AcceleratorsabstractPhotonic computing, known for its high bandwidth and energy efficiency, harnesses physical phenomena in the optical domain to accelerate a wide range of computational operations such as dot product, matrix multiplication, Fourier transform, 1D convolution, and more. However, the multitude of computational operations mentioned above poses challenges in mapping realistic neural network workloads onto underlying photonic hardware. This complexity requires extensive expertise and laborious programming, impeding the practical adoption and deployment of photonic acceleration. To address this gap, we propose an end-to-end compilation framework comprising a Photonic Compiler Collection (PCC). This framework automates the mapping of high-level deep neural network (DNN) specifications onto target architectures of photonic-electronic accelerators. Additionally, we present a method to streamline neural network workloads by leveraging the multilevel intermediate representation (MLIR) and compiler optimization techniques, targeting photonic-specific patterns. Moreover, we conduct a comprehensive case study illustrating the integration of a typical computational operator, the Mach-Zehnder Interferometer (MZI) mesh, into PCC. Our experimental results demonstrate that PCC achieves up to a 4x speedup on DNN workloads compared to handcrafted implementations. In summary, our proposed framework offers a practical and automated solution for compiling, optimizing, and flexibly sup-porting newer operators of photonic devices. We anticipate that our framework will significantly accelerate the development and deployment of photonic applications in real-world AI scenarios. Bohan Hu, Yinyi Liu, Wei Zhang 0012, Jiang Xu 0001 |
ICCD | 2 |
| 2024 | NEOCNN: NTT-Enabled Optical Convolution Neural Network AcceleratorabstractIn the realm of neural network computation, optical neural network accelerators (ONNs) have emerged as a promising solution, leveraging the inherent speed and parallelism of optical systems. Despite their potential, current ONN designs often fall short due to inefficient data movement and reliance on traditional electronics-based dataflows. Yinyi Liu, Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
ICS | 2 |
| 2023 | RONet: Scaling GPU System with Silicon Photonic ChipletabstractModern GPU systems integrate hundreds of SMs on a single die, and future scaling envisions even more SMs being incorporated. However, the limited number of transistors per die constrains this growth. While current chiplet technology shows promise, its performance is limited by the bandwidth and energy efficiency of existing chiplet interconnect technologies. In contrast, optical interconnects offer ultra-high bandwidth and energy efficiency, making them ideal for high-performance chiplet-based GPUs. This work proposes a novel region-based optical network, called RONet, that divides a chiplet-based GPU system with a 2D Mesh layout into multiple row and column regions, where each region is connected by a separate optical link. Additionally, RONet employs a tuning-free transmission mechanism to further enhance inter-chiplet bandwidth. Experimental results show that RONet achieves 43% improvement on performance and 25.4% reduction on system energy consumption over the baseline. Chengeng Li, Fan Jiang 0015, Shixi Chen, Yinyi Liu, Lin Chen 0029, Xiao Li 0038, Jiang Xu 0001 |
ICCAD | 5 |
| 2023 | FIONA: Photonic-Electronic CoSimulation Framework and Transferable Prototyping for Photonic AcceleratorabstractRecent advances in the architecture design for photonic accelerators have demonstrated great promise to accelerate deep neural network (DNN) applications, and also allude to the essential collaboration of the electronic subsystems for efficient logic arithmetic and memory access. However, available tools to design and evaluate photonic accelerators usually neglect the cross-stack effects or low-level details in real-world scenarios, ranging from programming-stack inefficiency to electronic peripheral implementation complexity. This frustrating fact makes it difficult to holistically estimate the performance metrics of a practical photonic-electronic collaborative computing system. In addition, until now, no toolchain can provide programmable, hardware-reconfigurable, and end-to-end rapid verification for photonic accelerators. Here we present FIONA, a Full-stack Infrastructure for Optical Neural Accelerator, which comprises a photonic-electronic co-simulation framework for multilevel design space exploration (DSE), and a transferable hardware prototyping template for physical verification. Specifically, the co-simulation framework consists of a functional simulator at the instruction set architecture (ISA) level to agilely verify the programming software stack and a register-transfer level (RTL) cycle-accurate simulator to precisely profile the overall system. We also demonstrate LightRocket as a case study of the FIONA toolchain to show the full workflow of designing a Turing-complete photonic accelerator system that supports arbitrary DNN workloads and on-chip training. The toolchain is open-sourced and available at https://github.com/hkust-fiona/. Yinyi Liu, Bohan Hu, Linfeng Du, Wei Zhang 0012, Jiang Xu 0001 |
ICCAD | 1 |
| 2022 | PHANES: ReRAM-based photonic accelerator for deep neural networksabstractResistive random access memory (ReRAM) has demonstrated great promises of in-situ matrix-vector multiplications to accelerate deep neural networks. However, subject to the intrinsic properties of analog processing, most of the proposed ReRAM-based accelerators require excessive costly ADC/DAC to avoid distortion of electronic analog signals during inter-tile transmission. Moreover, due to bit-shifting before addition, prior works require longer cycles to serially calculate partial sum compared to multiplications, which dramatically restricts the throughput and is more likely to stall the pipeline between layers of deep neural networks. Yinyi Liu, Shixi Chen, Jiang Xu 0001 |
DAC | 1 |
| 2022 | A Reliability Concern on Photonic Neural NetworksabstractEmerging integrated photonic neural networks have experimentally proved to achieve an ultra-high speedup of deep neural network training and inference in the optical domain. However, photonic devices suffer from the inherent crosstalk noise and loss, inevitably leading to reliability concerns. This paper systematically analyzes the impacts of crosstalk and loss on photonic computing systems. We propose a crosstalk-aware model for reliability estimation and find out the worst-case bounds as we increase the footprints and scales of the photonic chips. Our evaluations show that −30dB crosstalk noise can cause maximal photonic chip integration to a sharp drop by 109x. To facilitate very-large-scale photonic integration for future computing, we further propose multiple heterogeneous bijou photonic-cores to address the crosstalk-aware reliability concern. Yinyi Liu, Jun Feng 0008, Shixi Chen, Jiang Xu 0001 |
DATE | 1 |
| 2022 | Accelerating Cache Coherence in Manycore Processor through Silicon Photonic ChipletabstractCache coherence overhead in manycore systems is becoming prominent with the increase of system scale. However, traditional electrical networks restrict the efficiency of cache coherence transactions in the system due to the limited bandwidth and long latency. Optical network promises high bandwidth and low latency, and supports both efficient unicast and multicast transmission, which can potentially accelerate cache coherence in manycore systems. This work proposes a novel photonic cache coherence network with a physically centralized logically distributed directory called PCCN for chiplet-based manycore systems. PCCN adopts a channel sharing method with a contention solving mechanism for efficient long-distance coherence-related packet transmission. Experiment results show that compared to state-of-the-art proposals, PCCN can speed up application execution time by 1.32x, reduce memory access latency by 26%, and improve energy efficiency by 1.26x, on average, in a 128-core system. Chengeng Li, Fan Jiang 0015, Shixi Chen, Yinyi Liu, Jiang Xu 0001 |
ICCAD | 5 |
| 2022 | HERO: Pbit High-Radix Optical Switch Based on Integrated Silicon Photonics for Data CenterabstractTo establish flatten networks and accomplish rapid and efficient communications in the future hyper-scale data centers, HERO, a high-radix optical switch based on integrated silicon photonics, is proposed in this work. The architecture of HERO, including the switch fabric, switch interface, and switch controller, is described in detail. Two new switch control approaches: 1) split-transaction predictive control and 2) wavelength-group switching, are developed. The efficient control together with the optimized high-radix integrated optical switch fabrics help HERO achieve over 1 Pbps switching capacity. Even for small packets, such as 64–256 B Ethernet packets, the maximal utilization of the switch can reach up to 83%, and the throughput can approximate 1 Pbps. Further design explorations on the packet length and some key configurations, including the number of wavelengths and wavelength groups, are also conducted in this work, paving the way to the design of high-performance flatten data-center networks in the future. Jun Feng 0008, Jiang Xu 0001, Xuanqi Chen, Shixi Chen, Yinyi Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |