VLDB 2026 Research / reviewers in the wild / expert
Xiaoyu Sun 0001
dblp:23/7764-1
· DBLP profile ↗
12ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-5337-5680ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 6 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ultrafast Generative AI by Ultradense 3D Integration: A Case Study on LLM-based Edge InferenceabstractGenerative AI (GenAI) is one of the most critical applications today, continually challenging the limits of semiconductor technology. We introduce a very fine-grained 3D memory-on-logic architecture along with a novel data mapping strategy to support Large Language Model (LLM)-based GenAI, including both prefill and generation stages. Our conceptual analysis shows how ultradense 3D connectivity can enhance text generation speed and energy-efficiency well-beyond current limits. Preliminary findings from a basic analytical model indicate that the single batch autoregressive generation rate for Llama 3.2 1B could surpass 5K tokens/sec by maximizing weight locality and enhancing memory bandwidth through massively parallel 3D links between Multiply-Accumulate (MAC) units in the logic tier and their dedicated memory partitions in the 3D stack. We also explore the impact of advanced logic nodes and quantify their benefits in reducing prefill latency. Finally, we examine the challenges associated with memory access power and power density under extreme bandwidth conditions and present pipelined access strategies to address them. Kerem Akarvardar, Xiaoyu Sun 0001, Brian Crafton, Xiaochen Peng, Haruki Mori, Abhiroop Bhattacharjee, Hidehiro Fujiwara, H.-S. Philip Wong |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | Finding the Pareto Frontier of Low-Precision Data Formats and MAC Architecture for LLM InferenceabstractTo accelerate AI applications, numerous data formats and physical implementations of matrix multiplication have been proposed, creating a complex design space. This paper studies the efficient MAC implementation of the integer, floating-point, posit, and logarithmic number system (LNS) data formats and Microscaling (MX) and VectorScaled Quantization (VSQ) block data formats. We evaluate the area, power, and numerical accuracy (evaluated as signal-to-quantization noise ratio) of $\mathbf{3 5, 0 0 0}$ MAC designs spanning each data format and several key design parameters such as the inner product size and accumulation width. We find that for the same numerical accuracy, pareto optimal MAC designs with emerging data formats (LNS16, MXINT8, VSQINT4) achieve $1.8 \times 2.2 \times$, and $1.9 \times$ TOPs/W improvement compared to FP16, FP8, and FP4 dot product implementations. Brian Crafton, Xiaochen Peng, Xiaoyu Sun 0001, Ashwin Sanjay Lele, Win-San Khwa, Kerem Akarvardar |
DAC | 3 |
| 2024 | Non-volatile Memory Technologies for Edge AI ApplicationsabstractEmbedded non-volatile memories (NVMs) hold significant promise for advancing edge AI applications by offering unique advantages in near/in-memory computing-based accelerators. This invited paper delivers an in-depth review of the use cases for NVMs, emphasizing their potential to enhance area- and energy-efficiency in edge devices. We begin by outlining the latest advancements in TSMC NVM technologies and examining several NVM-based accelerator test chips enabled by the TSMC University Shuttle Program. Additionally, we delve into the tradeoffs involved in optimizing NVM devices, explore their potential for approximate computing applications, and assess the impact of NVM non-idealities on inference accuracy. Xiaoyu Sun 0001, Win-San Khwa, Xiaochen Peng, Meng-Fan Chang, Kerem Akarvardar |
ICCAD | 1 |
| 2024 | Efficient Processing of MLPerf Mobile Workloads Using Digital Compute-In-Memory MacrosabstractCompute-in-memory (CIM) has recently emerged as a promising design paradigm to accelerate deep neural network (DNN) processing. Continuously better energy and area efficiency at the macrolevel had been reported through many testchips over the last few years. However, in those macro design-oriented studies, accelerator-level considerations, such as memory accesses and processing of entire DNN workloads have not been investigated in-depth. In this article, we aim to fill this gap starting with the characteristics of our latest CIM macro fabricated with cutting-edge FinFET CMOS technology at 4-nm node. We then study, through an accelerator simulator developed in-house, three key items that would determine the efficiency of our CIM macro in the accelerator context while running MLPerf Mobile suite: 1) dataflow optimization; 2) optimal selection of CIM macro dimensions to further improve macro utilization; and 3) optimal combination of multiple CIM macros. Although there is typically a stark contrast between macro-level peak and accelerator-level average throughput and energy efficiency, the aforementioned optimizations are shown to improve the macro utilization by$3.04\times $and reduce the energy-delay product (EDP) to$0.34\times $compared to the original macro on MLPerf Mobile inference workloads. While we exploit a digital CIM macro in this study, the findings and proposed methods remain valid for other types of CIM (such as analog CIM and analog–digital–hybrid CIM) as well. Xiaoyu Sun 0001, Weidong Cao 0001, Brian Crafton, Kerem Akarvardar, Haruki Mori, Hidehiro Fujiwara, Hiroki Noguchi, Yu-Der Chih, Meng-Fan Chang, Yih Wang, Tsung-Yung Jonathan Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Estimating Power, Performance, and Area for On-Sensor Deployment of AR/VR Workloads Using an Analytical FrameworkabstractAugmented Reality and Virtual Reality have emerged as the next frontier of intelligent image sensors and computer systems. In these systems, 3D die stacking stands out as a compelling solution, enabling in situ processing capability of the sensory data for tasks such as image classification and object detection at low power, low latency, and a small form factor. These intelligent 3D CMOS Image Sensor (CIS) systems present a wide design space, encompassing multiple domains (e.g., computer vision algorithms, circuit design, system architecture, and semiconductor technology, including 3D stacking) that have not been explored in-depth so far. This article aims to fill this gap. We first present an analytical evaluation framework, STAR-3DSim, dedicated to rapid pre-RTL evaluation of 3D-CIS systems capturing the entire stack from the pixel layer to the on-sensor processor layer. With STAR-3DSim, we then propose several knobs for PPA (power, performance, area) improvement of the Deep Neural Network (DNN) accelerator that can provide up to 53%, 41%, and 63% reduction in energy, latency, and area, respectively, across a broad set of relevant AR/VR workloads. Last, we present full-system evaluation results by taking image sensing, cross-tier data transfer, and off-sensor communication into consideration. Xiaoyu Sun 0001, Xiaochen Peng, Sai Qian Zhang, Jorge Gomez 0002, Win-San Khwa, Syed Shakib Sarwar, Ziyun Li 0001, Weidong Cao 0001, Chiao Liu, Meng-Fan Chang, Barbara De Salvo, Kerem Akarvardar, H.-S. Philip Wong |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2022 | Achieving High In Situ Training Accuracy and Energy Efficiency with Analog Non-Volatile Synaptic DevicesabstractOn-device embedded artificial intelligence prefers the adaptive learning capability when deployed in the field, and thus in situ training is required. The compute-in-memory approach, which exploits the analog computation within the memory array, is a promising solution for deep neural network (DNN) on-chip acceleration. Emerging non-volatile memories are of great interest, serving as analog synapses due to their multilevel programmability. However, the asymmetry and nonlinearity in the conductance tuning remain grand challenges for achieving high in situ training accuracy. In addition, analog-to-digital converters at the edge of the memory array introduce quantization errors. In this work, we present an algorithm-hardware co-optimization to overcome these challenges. We incorporate the device/circuit non-ideal effects into the DNN propagation and weight update steps. By introducing the adaptive “momentum” in the weight update rule, in situ training accuracy on CIFAR-10 could approach its software baseline even under severe asymmetry/nonlinearity and analog-to-digital converter quantization error. The hardware performance of the on-chip training architecture and the overhead for adding “momentum” are also evaluated. By optimizing the backpropagation dataflow, 23.59 TOPS/W training energy efficiency (12× improvement compared to naïve dataflow) is achieved. The circuits that handle “momentum” introduce only 4.2% energy overhead. Our results show great potential and more relaxed requirements that enable emerging non-volatile memories for DNN acceleration on the embedded artificial intelligence platforms. Shanshi Huang, Xiaoyu Sun 0001, Xiaochen Peng, Hongwu Jiang, Shimeng Yu |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2021 | Heterogeneous Mixed-Signal Monolithic 3-D In-Memory Computing Using Resistive RAMabstractResistive random access memory (RRAM)-based compute-in-memory architecture helps overcome the bottleneck caused by large memory transactions in the convolutional neural network (CNN) accelerators. However, their deployment using 2-D IC technology faces challenges, as today's RRAM cells remain at legacy nodes above 20 nm due to high programming voltages. Besides, power-hungry analog-to-digital converter (ADC) units limit the throughput of RRAM accelerators. In this article, we present the first-ever heterogeneous (multiple technology nodes) mixed-signal monolithic 3-D IC designs of the RRAM CNN accelerator. Our RRAM remains at legacy 40-nm nodes in one tier, but CMOS periphery scales toward advanced 28/16 nm in another tier. Our 3-D designs overcome the bottleneck caused by ADCs and offer up to 4.9× improvement in energy efficiency in TOPS/W and up to 50% reduction in footprint area over 40-nm 2-D IC designs. Compared with existing 2-D works, our 3-D architecture offers up to 28.6× improvement in energy efficiency. Gauthaman Murali, Xiaoyu Sun 0001, Shimeng Yu, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Overcoming Challenges for Achieving High in-situ Training Accuracy with Emerging MemoriesabstractEmbedded artificial intelligence (AI) prefers the adaptive learning capability when deployed in the field, thus in- situ training on-chip is required. Emerging non-volatile memories (eNVMs) are of great interests serving as analog synapses in deep neural network (DNN) on-chip acceleration due to its multilevel programmability. However, the asymmetry/nonlinearity in the conductance tuning remains a grand challenge for achieving high in-situ training accuracy. In addition, analog-to-digital converter (ADC) at the edge of the memory array introduces an additional challenge - quantization error for in-memory computing. In this work, we gain new insights and overcome these challenges through an algorithm-hardware co-optimization. We incorporate these hardware non-ideal effects into the DNN propagation and weight update steps. We evaluate on a VGG-like network for CIFAR-10 dataset, and we show that the asymmetry of the conductance tuning is no longer a limiting factor of in-situ training accuracy if exploiting adaptive "momentum" in the weight update rule. Even considering ADC quantization error, in-situ training accuracy could approach software baseline. Our results show much relaxed requirements that enable a variety of eNVMs for DNN acceleration on the embedded AI platforms. Shanshi Huang, Xiaoyu Sun 0001, Xiaochen Peng, Hongwu Jiang, Shimeng Yu |
DATE | 2 |
| 2018 | Fully parallel RRAM synaptic array for implementing binary neural network with (+1, -1) weights and (+1, 0) neuronsabstractBinary Neural Networks (BNNs) have been recently proposed to improve the area-/energy-efficiency of the machine/deep learning hardware accelerators, which opens an opportunity to use the technologically more mature binary RRAM devices to effectively implement the binary synaptic weights. In addition, the binary neuron activation enables using the sense amplifier instead of the analog-to-digital converter to allow bitwise communication between layers of the neural networks. However, the sense amplifier has intrinsic offset that affects the threshold of binary neuron, thus it may degrade the classification accuracy. In this work, we analyze a fully parallel RRAM synaptic array architecture that implements the fully connected layers in a convolutional neural network with (+1, -1) weights and (+1, 0) neurons. The simulation results with TSMC 65 nm PDK show that the offset of current mode sense amplifier introduces a slight accuracy loss from ~98.5% to ~97.6% for MNIST dataset. Nevertheless, the proposed fully parallel BNN architecture (P-BNN) can achieve 137.35 TOPS/W energy efficiency for the inference, improved by ~20X compared to the sequential BNN architecture (S-BNN) with row-by-row read-out scheme. Moreover, the proposed P-BNN architecture can save the chip area by ~16% as it eliminates the area overhead of MAC peripheral units in the S-BNN architecture. Xiaoyu Sun 0001, Xiaochen Peng, Pai-Yu Chen, Rui Liu 0005, Jae-sun Seo, Shimeng Yu |
ASP-DAC | 1 |
| 2018 | Parallelizing SRAM arrays with customized bit-cell for binary neural networksabstractRecent advances in deep neural networks (DNNs) have shown Binary Neural Networks (BNNs) are able to provide a reasonable accuracy on various image datasets with a significant reduction in computation and memory cost. In this paper, we explore two BNNs: hybrid BNN (HBNN) and XNOR-BNN, where the weights are binarized to +1/-1 while the neuron activations are binarized to 1/0 and +1/-1 respectively. Two SRAM bit cell designs are proposed, namely, 6T SRAM for HBNN and customized 8T SRAM for XNOR-BNN. In our design, the high-precision multiply-and-accumulate (MAC) is replaced by bitwise multiplication for HBNN or XNOR for XNOR-BNN plus bit-counting operations. To parallelize the weighted sum operation, we activate multiple word lines in the SRAM array simultaneously and digitize the analog voltage developed along the bit line by a multi-level sense amplifier (MLSA). In order to partition the large matrices in DNNs, we investigate the impact of sensing bit-levels of MLSA on the accuracy degradation for different sub-array sizes and propose using the nonlinear quantization technique to mitigate the accuracy degradation. With 64×64 sub-array size and 3-bit MLSA, HBNN and XNOR-BNN architectures can minimize the accuracy degradation to 2.37% and 0.88%, respectively, for an inspired VGG-16 network on the CIFAR-10 dataset. Design space exploration of SRAM based synaptic architectures with the conventional row-by-row access scheme and our proposed parallel access scheme are also performed, showing significant benefits in the area, latency and energy-efficiency. Finally, we have successfully taped-out and validated the proposed HBNN and XNOR-BNN designs in TSMC 65 nm process with measured silicon data, achieving energy-efficiency >100 TOPS/W for HBNN and >50 TOPS/W for XNOR-BNN. Rui Liu 0005, Xiaochen Peng, Xiaoyu Sun 0001, Win-San Khwa, Xin Si, Jia-Jing Chen, Jia-Fang Li, Meng-Fan Chang, Shimeng Yu |
DAC | 3 |
| 2018 | XNOR-RRAM: A scalable and parallel resistive synaptic architecture for binary neural networksabstractRecent advances in deep learning have shown that Binary Neural Networks (BNNs) are capable of providing a satisfying accuracy on various image datasets with significant reduction in computation and memory cost. With both weights and activations binarized to +1 or -1 in BNNs, the high-precision multiply-and-accumulate (MAC) operations can be replaced by XNOR and bit-counting operations. In this work, we propose a RRAM synaptic architecture (XNOR-RRAM) with a bit-cell design of complementary word lines that implements equivalent XNOR and bit-counting operation in a parallel fashion. For large-scale matrices in fully connected layers or when the convolution kernels are unrolled in multiple channels, the array partition is necessary. Multi-level sense amplifiers (MLSAs) are employed as the intermediate interface for accumulating partial weighted sum. However, a low bit-level MLSA and intrinsic offset of MLSA may degrade the classification accuracy. We investigate the impact of sensing offsets on classification accuracy and analyze various design options with different sub-array sizes and sensing bit-levels. Experimental results with RRAM models and 65nm CMOS PDK show that the system with 128×128 sub-array size and 3-bit MLSA can achieve accuracies of 98.43% for MLP on MNIST and 86.08% for CNN on CIFAR-10, showing 0.34% and 2.39% degradation respectively compared to the accuracies of ideal BNN algorithms. The projected energy-efficiency of XNOR-RRAM is 141.18 TOPS/W, showing ~33X improvement compared to the conventional RRAM synaptic architecture with sequential row-by-row read-out. Xiaoyu Sun 0001, Shihui Yin, Xiaochen Peng, Rui Liu 0005, Jae-sun Seo, Shimeng Yu |
DATE | 1 |
| 2017 | Low-VDD Operation of SRAM Synaptic Array for Implementing Ternary Neural NetworkabstractFor Internet of Things (IoT) edge devices, it is very attractive to have the local sensemaking capability instead of sending all the data back to the cloud for information processing. For image pattern recognition, neuro-inspired machine learning algorithms have demonstrated enormous powerfulness. To effectively implement learning algorithms on-chip for IoT edge devices, on-chip synaptic memory architectures have been proposed to implement the key operations such as weighted-sum or matrix-vector multiplication. In this paper, we proposed a low-power design of static random access memory (SRAM) synaptic array for implementing a low-precision ternary neural network. We experimentally demonstrated that the supply voltage (VDD) of the SRAM array could be aggressively reduced to a level, where the SRAM cell is susceptible to bit failures. The testing results from 65-nm SRAM chips indicate that VDD could be reduced from the nominal 1-0.55 V (or 0.5 V) with a bit error rate ~0.23% (or ~1.56%), which only introduced ~0.08% (or ~1.68%) degradation in the classification accuracy. As a result, the power consumption could be reduced by more than 8× (or 10×). Xiaoyu Sun 0001, Rui Liu 0005, Yi-Ju Chen, Hsiao-Yun Chiu, Wei-Hao Chen, Meng-Fan Chang, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |