Xiaoxin Xu

dblp:14/5613 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PSVS: A Parallel-Series Voltage-Sensing 1T4R RRAM Macro Achieving 15-Level Storage and 17.88 Mb/mm2 Density for LLM Inference
Kaijun Zhang, Wencong Wu, Jiapei Zheng, Jinru Lai, Xiaoxin Xu, Qi Liu 0010, Chixiao Chen
ISCAS6
2025 Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy Efficiency
abstract
Recent advancements in deep learning have significantly increased AI processors' energy consumption, which is becoming a critical factor limiting AI development. Dynamic Voltage and Frequency Scaling (DVFS) stands as a key method in power optimization. However, due to the latency of DVFS control in AI processors, previous works typically apply DVFS control at the granularity of a program's entire duration or sub-phases, rather than at the level of AI operators.
Yijia Zhang 0002, Fuchun Wei, Bingqiang Wang, Yanlin Liu, Zhiheng Hu, Xiaoxin Xu, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
ASPLOS (1)8
2025 A PulseWidth-IN-PulseWidth-Out Universal Nonlinear Processing Element for Time-Domain In-Memory Computing Systems
abstract
Time-Domain In-Memory Computing (TD-IMC) has emerged as a promising analog computing architecture for edge AI applications. However, the lack of developed hardware operators, especially general nonlinear operators, necessitates frequent cross-domain data transmission in practical TD-IMC systems, significantly reducing energy efficiency. In this work, we propose a PulseWidth-IN-PulseWidth-OUT Universal Nonlinear Processing Element (PIPO-UNPE) to address the challenges of nonlinear processing in analog computing. By implementing an RRAM-based two-layer ReLU network, the PIPO-UNPE performs universal nonlinear operations entirely in the time domain. Algorithmically, we introduce Dynamic Loss-Responsive Subset Enhancement (DLRSE) to boost the performance of this low-cost network in function approximation tasks. From a hardware perspective, we design an RRAM-based pulse-driven programmable current source and a low-latency dispersion comparator-based voltage-to-time converter (VTC) to enhance both the energy efficiency and precision of the PIPO-UNPE. Hybrid simulations reveal that the PIPO-UNPE consumes 912 uW of power while delivering a throughput of $\mathbf{1 0 M}$ NOPS (Nonlinear Operations Per Second). Incorporating the PIPO-UNPE into the TD-IMC accelerator can increase energy efficiency by a factor of 9.5 to 25, keeping the accuracy loss below 0.1%.
Pengcheng Feng, Rongxuan Shen, Huaxiang Lu, Xiaoxin Xu
DAC7
2025 Re4PUF: A Reliable, Reconfigurable ReRAM-based PUF Resilient to DNN and Side Channel Attacks
abstract
Resistive random-access memory (ReRAM) based Physical Unclonable Functions (PUFs) have emerged as an attractive hardware security primitive due to their low energy consumption and compact footprint. However, the reliability of existing ReRAM-based PUFs is challenged by read noise and temperature variations, as well as their resistance to Deep Neural Network (DNN) modeling attacks and Side Channel Attacks (SCAs). In this paper, we propose a novel 3T2R ReRAM-based reconfigurable PUF to address these challenges. By adopting the digital 3T2R voltage division cell design, we improve its reliability against ReRAM read noise and temperature variations, while the adjustable analog supply voltage of inverters enables quick, low-cost reconfigurability without reprogramming ReRAMs, effectively mitigating DNN modeling and SCA vulnerabilities. Our $\mathrm{Re}^{4}$ PUF chip has been experimentally validated, achieving a low Bit Error Rate (BER) of $1 \%$ at $85^{\circ} \mathrm{C}$, a 7.59 -fold reduction compared to existing ReRAM-based PUFs. It also demonstrates robust resistance to both DNN modeling attacks (MLP and Transformer) and SCAs, with success rates of approximately 50% and less than $70 \%$, respectively.
Ning Lin, Yangu He, Songqi Wang, Hegan Chen, Kwunhang Wong, Chuxin Li, Jichang Yang, Yongkang Han, Xiaoxin Xu, Dashan Shang
DAC14
2025 An Energy-Efficient High-Utilization Hardware Architecture for Attention Mechanism in Transformer using Balanced Systolic Array and Multi-Row Interleaved Operation Ordering
abstract
Transformer-based neural networks have achieved remarkable performance. Designing energy-efficient and high-speed accelerators for the attention mechanism, which dominates the energy and latency in Transformers, has become increasingly significant. Existing attention accelerators commonly use algorithm-hardware co-design to achieve higher energy efficiency and speed. However, deeply customized algorithms make these accelerators dependent on a particular application. Therefore, optimizing hardware architecture is crucial for achieving general-purpose acceleration. We observe two limitations in the hardware architecture of existing attention accelerators. First, the widely used input stationary, weight stationary, and output stationary systolic arrays (SAs) can’t balance data reuse, register saving, and utilization, which hinders to build more energy-efficient and faster SA-based accelerators. Second, layer-by-layer operation ordering introduces high SRAM access overhead of intermediate results. To address the first limitation, we propose the “Balanced Systolic Array”, which improves energy efficiency by 40% compared to conventional systolic arrays and achieves a utilization rate of 99.5%. To address the second limitation, we propose “Multi-Row Interleaved” operation ordering, which reduces the SRAM energy by 31.7% By integrating two techniques, the proposed attention accelerator achieves a 39% improvement in energy efficiency and a 38% enhancement in throughput×energy efficiency compared to previous works.
Haiyang Zhou, Hongyang Hu, Jinshan Yue, Hanghang Gao, Yuanlu Xie, Xiaoxin Xu, Chunmeng Dou, Ming Liu 0022
DAC6
2025 A monolithic 3D IGZO-RRAM-SRAM-integrated architecture for robust and efficient compute-in-memory enabling equivalent-ideal device metrics
Shengzhe Yan, Zhaori Cong, Zhuoyu Dai, Zeyu Guo 0002, Zhihang Qian, Xufan Li, Chuanke Chen, Nianduan Lu, Chunmeng Dou, Guanhua Yang, Xiaoxin Xu, Di Geng, Jinshan Yue, Ling Li 0013, Ming Liu 0022
Sci. China Inf. Sci.13
2025 An RRAM Digital Computing-in-Memory Macro With Dual-Mode Multiplication and Maximum Value Rounding Adder Tree
abstract
Implementing digital computing-in-memory (DCIM) based on resistive memory (RRAM) faces several critical challenges due to the small signal margin, large device variations, and large energy- and area-overhead induced by the digital adder tree (AT). To address these issues, we propose an RRAM DCIM macro based on the standard foundry one-transistor-one-resistor (1T1R) cell array featuring: 1) dual-mode MAC operation for efficiency- or accuracy-oriented optimization; 2) margin-enhanced digitized unit (MEDU) to amplify the signal ratio; and 3) maximum value rounding AT (MVR-AT) to reduce its power- and area-overhead. A test chip is demonstrated using a 180 nm CMOS process to verify the concept. It achieves a peak energy efficiency (EF) of 63.08 TOPS/W in the efficiency-oriented mode and a minimum error rate of 1.58% in the accuracy-oriented mode. Their combination can meet the requirements of different workloads in AI computing tasks to optimize the overall power consumption with negligible accuracy loss.
Wang Ye, Hanghang Gao, Zhidao Zhou, Linfang Wang, Weizeng Li, Zhi Li 0062, Jinshan Yue, Xiaoxin Xu, Hongyang Hu, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.8
2024 CMN: a co-designed neural architecture search for efficient computing-in-memory-based mixture-of-experts
abstract
Abstract Artificial intelligence (AI) has experienced substantial advancements recently, notably with the advent of large-scale language models (LLMs) employing mixture-of-experts (MoE) techniques, exhibiting human-like cognitive skills. As a promising hardware solution for edge MoE implementations, the computing-in-memory (CIM) architecture collocates memory and computing within a single device, significantly reducing the data movement and the associated energy consumption. However, due to diverse edge application scenarios and constraints, determining the optimal network structures for MoE, such as the expert’s location, quantity, and dimension on CIM systems remains elusive. To this end, we introduce a software-hardware co-designed neural architecture search (NAS) framework, C IM-based M oE N AS (CMN), focusing on identifying a high-performing MoE structure under specific hardware constraints. The results of the NYUD-v2 dataset segmentation on the RRAM (SRAM) CIM system reveal that CMN can discover optimized MoE configurations under energy, latency, and performance constraints, achieving 29.67 × ( 43.10 ×) energy savings, 175.44 ×( 109.89 ×) speedup, and 12.24 × smaller model size compared to the baseline MoE-enabled Visual Transformer, respectively. This co-design opens up an avenue toward high-performance MoE deployments in edge CIM systems.
Shihao Han, Sishuo Liu, Shucheng Du, Mingzi Li, Zijian Ye, Xiaoxin Xu, Dashan Shang
Sci. China Inf. Sci.6
2024 Erratum to: CMN: a co-designed neural architecture search for efficient computing-in-memory-based mixture-of-experts
Shihao Han, Sishuo Liu, Shucheng Du, Mingzi Li, Zijian Ye, Xiaoxin Xu, Dashan Shang
Sci. China Inf. Sci.6
2023 Transport mechanism in Hf0.5Zr0.5O2-based ferroelectric diodes
Tiancheng Gong, Zhaomeng Gao, Xiaoxin Xu, Jianfeng Gao 0005
Sci. China Inf. Sci.6
2023 A 40-nm SONOS Digital CIM Using Simplified LUT Multiplier and Continuous Sample-Hold Sense Amplifier for AI Edge Inference
abstract
Digital computing in memory (CIM) exhibits high precision as well as high energy efficiency (EE) yet still lacks discussion in nonvolatile memory (NVM). In this article, we propose a 40-nm silicon-oxide-nitride-oxide-silicon (SONOS)-based digital NVM CIM macro (DNV-CIM) featuring: 1) a simplified lookup table multiplier (SLUTM) combined with a lookup table (LUT) mapping scheme to improve area and EE and 2) a continuous sample-hold sense amplifier (CSH-SA) with an optimized voltage clamper and comparator for continuous read to reduce overall energy and time consumption for deep neural network (DNN) inference tasks. Performance evaluations indicate that the proposed DNV-CIM can achieve 93.04% accuracy and an EE up to 39.9 TOPS/W when running a 4-bit quantized ResNet18 trained on the CIFAR-10 dataset. This work presents a highly efficient digital CIM solution that can be readily implemented with commodity NVM.
Hongyang Hu, Haiyang Zhou, Danian Dong, Jinshan Yue, Wan Pang, Xiaoxin Xu, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.8
2023 A Security-Enhanced, Charge-Pump-Free, ISO14443-A-/ISO10373-6-Compliant RFID Tag With 16.2-μW Embedded RRAM and Reconfigurable Strong PUF
abstract
Radio frequency identification technology (RFID) has empowered a wide variety of automation industries, such as logistics and freight transportation. To further promote RFID tags adoption, security, power consumption, and cost have always been issues of general concern. This article presents the first synergy of the RFID tag with embedded resistive RAM (RRAM) array and RRAM-based reconfigurable strong physical unclonable function (R-SPUF). The RRAM not only meets the mass storage and technology downscaling but also renders the ultralow-cost “1-cent RFID tag” more feasible. Moreover, the R-SPUF facilitates multiple initializations until a satisfactory distribution and has strong secure keys benefiting from its reconfigurability that improves both safety and reliability. The complete system operates at 13.56 MHz and is compliant with the ISO14443-A and ISO10373-6 (test) protocols. The RFID tag was fabricated on a 1.1-mm2 die based on the 0.18-$\mu \text{m}$CMOS process. Without resorting to the charge pumps for RRAM read–write operations, the total power consumption is as low as 52.3$\mu \text{W}$, of which the RRAM dissipates$16.2~\mu \text{W}$under a wireless power supply.
Qirui Ren, Qiang Huo, Hao Wu 0084, Xiangqu Fu, Xiaoxin Xu, Jianfeng Gao 0005, Xiaojin Zhao, Dengyun Lei, Xinghua Wang 0005, Feng Zhang 0014, Yong Chen 0005, Pui-In Mak
IEEE Trans. Very Large Scale Integr. Syst.9
2022 A Novel Hybrid CNN-LSTM Compensation Model Against DoS Attacks in Power System State Estimation
Xiaoxin Xu, Jian Sun 0014
Neural Process. Lett.1
2021 Investigation of weight updating modes on oxide-based resistive switching memory synapse towards neuromorphic computing applications
Qingting Ding, Tiancheng Gong, Jie Yu 0027, Xiaoxin Xu, Hangbing Lv, Feng Zhang 0014, Ming Liu 0022
Sci. China Inf. Sci.4