VLDB 2026 Research / reviewers in the wild / expert
Xiaoxuan Yang 0001
dblp:131/7086-1
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-2553-2631ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpiKint: Native Full-Integer Spiking Neural Networks Training with an Efficient CIM-based AcceleratorabstractSpiking neural networks (SNNs) are commonly regarded as more energy-efficient than artificial neural networks (ANNs) due to their event-driven and sparse computations. However, the training overhead of SNNs with backpropagation through time (BPTT) is significantly higher than ANNs, while prior studies mainly focus on accelerating SNN inference rather than training. This work proposes SpiKint, an algorithm-hardware co-design to enable efficient BPTT-based SNN training. At the algorithm level, for the first time, we propose a native full-integer SNN training framework without introducing any quantization scheme. At the architecture level, we design the dedicated computing-in-memory-based (CIM) training hardware that introduces the efficient CIM macro to achieve high utilization via interference-free weight mapping and in-situ weight transpose computations in the backward pass. At the system level, we introduce a novel incremental weight update strategy for SNN training to reduce off-chip memory access while ensuring accurate backpropagation. Experiments show that SpiKint training framework achieves negligible accuracy drop compared to floating-point SNN training. SpiKint accelerator achieves an average 22.06 × energy-delay-product reduction and 50.88 × energy efficiency compared to Nvidia A100 GPU and systolic-based SATA, respectively. The code for SpiKint is available at https://github.com/peilin-chen/SpiKint/. Xiaoxuan Yang 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | Area-Efficient In-Memory Computing for Mixture-of-Experts via Multiplexing and Caching
Hanyuan Gao, Xiaoxuan Yang 0001 |
ISCAS | 2 |
| 2026 | SpikON: A Dual-Parallel and Efficient Accelerator for Online Spiking Neural Networks LearningabstractSpiking neural networks (SNNs) have emerged as a promising paradigm for energy-efficient brain-inspired computing. However, existing online unsupervised SNN learning suffers from low training accuracy and poor scalability. Although current online supervised learning algorithms perform well on large-scale datasets and networks, the non-hardware-friendly operations hinder efficient edge deployment. In this work, we propose SpikON, the first algorithm-hardware co-design framework for efficient and scalable end-to-end online supervised SNN learning. We first propose the learnable threshold through time and scaled weight centralization through time techniques to address the inefficiency of traditional algorithms. Moreover, to reduce latency and energy consumption, we introduce the novel training dataflow and cascade computation reuse scheme for SNNs that allows concurrent forward-backward computation and temporal reuse across timesteps. We further design the dedicated SNN accelerator with a dual-parallel engine and customized SIMD-based SNN core for efficient end-to-end online learning. Experiments show that the SpikON algorithm achieves 32.2% and 35.0% reductions in training latency and energy consumption over the baseline, without sacrificing accuracy. Moreover, the SpikON co-design achieves 7.2× (11.5×) and 26.8× (15.8×) training throughput (energy efficiency) compared with the edge Apple M4 GPU and TPU-like accelerator, respectively. The code is available at GitHub. Xiaoxuan Yang 0001 |
ISLPED | 2 |
| 2025 | Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
Xiaoxuan Yang 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
Tunhou Zhang, Junyao Zhang 0003, Jonathan Hao-Cheng Ku, Yitu Wang, Xiaoxuan Yang 0001, Hai Li 0001, Yiran Chen 0001 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2024 | Block-Wise Mixed-Precision Quantization: Enabling High Efficiency for Practical ReRAM-Based DNN AcceleratorsabstractResistive random access memory (ReRAM)-based processing-in-memory (PIM) architectures have demonstrated great potential to accelerate Deep Neural Network (DNN) training/ inference. However, the computational accuracy of analog PIM is compromised due to the non-idealities, such as the conductance variation of ReRAM cells. The impact of these non-idealities worsens as the number of concurrently activated wordlines and bitlines increases. To guarantee computational accuracy, only a limited number of wordlines and bitlines of the crossbar array can be turned on concurrently, significantly reducing the achievable parallelism of the architecture. While the constraints on parallelism limit the efficiency of the accelerators, they also provide a new opportunity for finegrained mixed-precision quantization. To enable efficient DNN inference on practical ReRAM-based accelerators, we propose an algorithm-architecture co-design framework called Block-Wise mixed-precision Quantization (BWQ). At the algorithm level, BWQ-A introduces a mixed-precision quantization scheme at the block level, which achieves a high weight and activation compression ratio with negligible accuracy degradation. We also present the hardware architecture design BWQ-H, which leverages the low-bit-width models achieved by BWQ-A to perform high-efficiency DNN inference on ReRAM devices. BWQ-H also adopts a novel precision-aware weight mapping method to increase the ReRAM crossbars throughput. Our evaluation demonstrates the effectiveness of BWQ, which achieves a 6.08× speedup and a 17.47× energy saving on average compared to existing ReRAM-based architectures. Xueying Wu, Edward Hanson, Nansu Wang, Qilin Zheng, Xiaoxuan Yang 0001, Huanrui Yang, Shiyu Li 0001, Partha Pratim Pande, Janardhan Rao Doppa, Krishnendu Chakrabarty, Hai Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Improving the Robustness and Efficiency of PIM-Based Architecture by SW/HW Co-DesignabstractProcessing-in-memory (PIM) based architecture shows great potential to process several emerging artificial intelligence workloads, including vision and language models. Cross-layer optimizations could bridge the gap between computing density and the available resources by reducing the computation and memory cost of the model and improving the model's robustness against non-ideal hardware effects. We first introduce several hardware-aware training methods to improve the model robustness to the PIM device's non-ideal effects, including stuck-at-fault, process variation, and thermal noise. Then, we further demonstrate a software/hardware (SW/HW) co-design methodology to efficiently process the state-of-the-art attention-based model on PIM-based architecture by performing sparsity exploration for the attention-based model and circuit-architecture co-design to support the sparse processing. Xiaoxuan Yang 0001, Shiyu Li 0001, Qilin Zheng, Yiran Chen 0001 |
ASP-DAC | 1 |
| 2023 | ESSENCE: Exploiting Structured Stochastic Gradient Pruning for Endurance-Aware ReRAM-Based In-Memory Training SystemsabstractProcessing-in-memory (PIM) enables energy-efficient deployment of convolutional neural networks (CNNs) from edge to cloud. Resistive random-access memory (ReRAM) is one of the most commonly used technologies for PIM architectures. One of the primary limitations of ReRAM-based PIM in neural network training arises from the limited write endurance due to the frequent weight updates. To make ReRAM-based architectures viable for CNN training, the write endurance issue needs to be addressed. This work aims to reduce the number of weight reprogrammings without compromising the final model accuracy. We propose the ESSENCE framework with an endurance-aware structured stochastic gradient pruning method, which dynamically adjusts the probability of gradient update based on the current update counts. Experimental results with multiple CNNs and datasets demonstrate that the proposed method can extend ReRAM’s life time for training. For instance, with the ResNet20 network and CIFAR-10 dataset, ESSENCE can save the mean update counts of up to$10.29\times $compared to the stochastic gradient descent method and effectively reduce the maximum update counts compared with the No Endurance method. Furthermore, an aggressive tuning method based on ESSENCE can boost the mean update count savings by up to$14.41\times $. Xiaoxuan Yang 0001, Huanrui Yang, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty, Hai Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | HERO: hessian-enhanced robust optimization for unifying and improving generalization and quantization performanceabstractWith the recent demand of deploying neural network models on mobile and edge devices, it is desired to improve the model's generalizability on unseen testing data, as well as enhance the model's robustness under fixed-point quantization for efficient deployment. Minimizing the training loss, however, provides few guarantees on the generalization and quantization performance. In this work, we fulfill the need of improving generalization and quantization performance simultaneously by theoretically unifying them under the framework of improving the model's robustness against bounded weight perturbation and minimizing the eigenvalues of the Hessian matrix with respect to model weights. We therefore propose HERO, a Hessian-enhanced robust optimization method, to minimize the Hessian eigenvalues through a gradient-based training process, simultaneously improving the generalization and quantization performance. HERO enables up to a 3.8% gain on test accuracy, up to 30% higher accuracy under 80% training label perturbation, and the best post-training quantization accuracy across a wide range of precision, including a > 10% accuracy improvement over SGD-trained models for common model architectures on various datasets. Huanrui Yang, Xiaoxuan Yang 0001, Neil Zhenqiang Gong, Yiran Chen 0001 |
DAC | 2 |
| 2022 | Approximate Computing and the Efficient Machine Learning ExpeditionabstractApproximate computing (AxC) has been long accepted as a design alternative for efficient system implementation at the cost of relaxed accuracy requirements. Despite the AxC research activities in various application domains, AxC thrived the past decade when it was applied in Machine Learning (ML). The by definition approximate notion of ML models but also the increased computational overheads associated with ML applications-that were effectively mitigated by corresponding approximations-led to a perfect matching and a fruitful synergy. AxC for AI/ML has transcended beyond academic prototypes. In this work, we enlighten the synergistic nature of AxC and ML and elucidate the impact of AxC in designing efficient ML systems. To that end, we present an overview and taxonomy of AxC for ML and use two descriptive application scenarios to demonstrate how AxC boosts the efficiency of ML systems. Jörg Henkel, Hai Li 0001, Anand Raghunathan, Mehdi Baradaran Tahoori, Swagath Venkataramani, Xiaoxuan Yang 0001, Georgios Zervakis 0001 |
ICCAD | 6 |
| 2022 | Research Progress on Memristor: From Synapses to Computing SystemsabstractAs the limits of transistor technology are approached, feature size in integrated circuit transistors has been reduced very near to the minimum physically-realizable channel length, and it has become increasingly difficult to meet expectations outlined by Moore’s law. As one of the most promising devices to replace transistors, memristors have many excellent properties that can be leveraged to develop new types of neural and non-von Neumann computing systems, which are expected to revolutionize information-processing technology. This survey provides a comparative overview of research progress on memristors. Different memristor synaptic devices are classified according to stimulation patterns and the working mechanisms of these various synaptic devices are analyzed in detail. Crossbar-based memristors have demonstrated advantages in physically executing vector-matrix multiplication and enabling highly power-efficient and area-efficient neuromorphic system designs. The extensive uses of crossbar-based memristors cover in-memory logic, vector-matrix multiplication, and many other fundamental computing operations. Furthermore, memristor-based architectures for efficient neural network training and inference have been studied. However, memristors have non-ideal properties due to programming inaccuracies and device imperfections from fabrication, which lead to error or mismatch in computed results. To build reliable memristor-based designs, circuit-level, algorithm-level, and system-level solutions to memristor reliability issues are being studied. To this end, state-of-the-art realizations of memristor crossbars, crossbar-based designs, and peripheral circuitry are presented, which show both promising full-system inference accuracy and excellent power efficiency in multiple tasks. Memristor in-situ learning benefits from high energy efficiency and biologically-imitative characteristics, which are conducive to further realizing hardware acceleration of cognitive learning. At present, the learning and training processes of brain-like networks are complex, presenting great challenges for network design and implementation. Xiaoxuan Yang 0001, Brady Taylor, Ailong Wu, Yiran Chen 0001, Leon O. Chua |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Multi-Objective Optimization of ReRAM Crossbars for Robust DNN Inferencing under Stochastic NoiseabstractResistive random-access memory (ReRAM) is a promising technology for designing hardware accelerators for deep neural network (DNN) inferencing. However, stochastic noise in ReRAM crossbars can degrade the DNN inferencing accuracy. We propose the design and optimization of a high-performance, area-and energy-efficient ReRAM-based hardware accelerator to achieve robust DNN inferencing in the presence of stochastic noise. We make two key technical contributions. First, we propose a stochastic-noise-aware training method, referred to as ReSNA, to improve the accuracy of DNN inferencing on ReRAM crossbars with stochastic noise. Second, we propose an information-theoretic algorithm, referred to as CF-MESMO, to identify the Pareto set of solutions to trade-off multiple objectives, including inferencing accuracy, area overhead, execution time, and energy consumption. The main challenge in this context is that executing the ReSNA method to evaluate each candidate ReRAM design is prohibitive. To address this challenge, we utilize the continuous-fidelity evaluation of ReRAM designs associated with prohibitive high computation cost by varying the number of training epochs to trade-off accuracy and cost. CF-MESMO iteratively selects the candidate ReRAM design and fidelity pair that maximizes the information gained per unit computation cost about the optimal Pareto front. Our experiments on benchmark DNNs show that the proposed algorithms efficiently uncover high-quality Pareto fronts. On average, ReSNA achieves 2.57% inferencing accuracy improvement for ResNet20 on the CIFAR-10 dataset with respect to the baseline configuration. Moreover, CF-MESMO algorithm achieves 90.91% reduction in computation cost compared to the popular multi-objective optimization algorithm NSGA-II to reach the best solution from NSGA-II. Xiaoxuan Yang 0001, Syrine Belakaria, Biresh Kumar Joardar, Huanrui Yang, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty, Hai Li 0001 |
ICCAD | 1 |
| 2020 | ReTransformer: ReRAM-based Processing-in-Memory Architecture for Transformer AccelerationabstractTransformer has emerged as a popular deep neural network (DNN) model for Neural Language Processing (NLP) applications and demonstrated excellent performance in neural machine translation, entity recognition, etc. However, its scaled dot-product attention mechanism in auto-regressive decoder brings a performance bottleneck during inference. Transformer is also computationally and memory intensive and demands for a hardware acceleration solution. Although researchers have successfully applied ReRAM-based Processing-in-Memory (PIM) to accelerate convolutional neural networks (CNNs) and recurrent neural networks (RNNs), the unique computation process of the scaled dot-product attention in Transformer makes it difficult to directly apply these designs. Besides, how to handle intermediate results in Matrix-matrix Multiplication (MatMul) and how to design a pipeline at a finer granularity of Transformer remain unsolved. In this work, we propose ReTransformer - a ReRAM-based PIM architecture for Transformer acceleration. ReTransformer can not only accelerate the scaled dot-product attention of Transformer using ReRAM-based PIM but also eliminate some data dependency by avoiding writing the intermediate results using the proposed matrix decomposition technique. Moreover, we propose a new sub-matrix pipeline design for multi-head self-attention. Experimental results show that compared to GPU and Pipelayer, ReTransformer improves computing efficiency by 23.21× and 3.25×, respectively. The corresponding overall power is reduced by 1086× and 2.82×, respectively. Xiaoxuan Yang 0001, Bonan Yan, Hai Li 0001, Yiran Chen 0001 |
ICCAD | 1 |