VLDB 2026 Research / reviewers in the wild / expert
Zihan Xia 0002
dblp:244/0846-2
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-5409-321XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRIM: Acceleration of Multiplication-Less Neural Networks via Versatile SparsitiesabstractRecently, multiplication-less neural networks (L1NNs), replacing multiplication-intensive dot products with addition-only$\ell _{1}$-distance kernels have emerged to enhance energy efficiency and speed with negligible degradation of accuracy. Despite these gains, such models suffer from pruning challenges that can negate their benefits. In this work, we identify the root cause of these challenges and introduce a novel method called synapse pruning, which is explicitly designed for L1NNs to overcome them for the first time. Building upon this algorithmic innovation, we present TRIM, an algorithm-hardware co-design framework that sparsifies and accelerates L1NNs to achieve ultra-high energy efficiency and speed. On the algorithmic side, we propose structured synapse pruning tailored for hardware-friendly$\mathbb {N:M}$sparsity patterns for L1NNs. On the hardware side, we introduce a sparse processor architecture that efficiently exploits both the proposed$\mathbb {N:M}$structured synapse sparsity and the intrinsic unstructured weight and activation sparsities in L1NNs by skipping redundant operations. Additionally, an$\mathbb {N:M}$sparsity-aware elastic mapping technique is introduced to maximize hardware utilization and data reuse. Evaluations on seven benchmarks demonstrate that TRIM achieves up to 87.5% unstructured sparsity and up to 81.3% structured sparsity with less than 1% accuracy loss. Implemented in a 65nm technology, our processor achieves 5.22 TOPS/W and 1.17 TOPS/mm2, surpassing the existing SOTA L1NN accelerator by$1.7\times $in energy efficiency and$14.9\times $in area efficiency. The source code is available at:https://github.com/XZH28/TRIM-Sparse-L1-Distance-Net.git Zihan Xia 0002, Mingu Kang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | FeNOMS: Enhancing Open Modification Spectral Library Search with In-Storage Processing on Ferroelectric NAND (FeNAND) FlashabstractThe rapid expansion of mass spectrometry (MS) data, now exceeding hundreds of terabytes, poses significant challenges for efficient, large-scale library search — a critical component for drug discovery. Traditional processors struggle to handle this data volume efficiently, making in-storage computing (ISP) a promising alternative. This work introduces an ISP architecture leveraging a 3D Ferroelectric NAND (FeNAND) structure, providing significantly higher density, faster speeds, and lower voltage requirements compared to traditional NAND flash. Despite its superior density, the NAND structure has not been widely utilized in ISP applications due to limited throughput associated with row-by-row reads from serially connected cells. To overcome these limitations, we integrate hyperdimensional computing (HDC), a brain-inspired paradigm that enables highly parallel processing with simple operations and strong error tolerance. By combining HDC with the proposed dual-bound approximate matching (D-BAM) distance metric, tailored to the FeNAND structure, we parallelize vector computations to enable efficient MS spectral library search, achieving 43× speedup and 21× higher energy efficiency over state-of-the-art 3D NAND methods, while maintaining comparable accuracy. Sumukh Pinge, Ashkan Moradifirouzabadi, Keming Fan, Prasanna Venkatesan Ravindran, Tanvir H. Pantha, Po-Kai Hsu, Zihan Xia 0002, Flavio Ponzina, Winston Chern, Taeyoung Song, Priyankka Gundlapudi Ravikumar, Mengkun Tian, Lance Fernandes, Hari Jayasankar, Chinsung Park, Amrit Garlapati, Kijoon Kim, Jongho Woo, Suhwan Lim, Wanki Kim, Daewon Ha, Duygu Kuzum, Shimeng Yu, Tajana Rosing, Mingu Kang |
ICCAD | 9 |
| 2025 | Hybrid SLC-MLC RRAM Mixed-Signal Processing-in-Memory Architecture for Transformer Acceleration via Gradient RedistributionabstractTransformers, while revolutionary, face challenges due to their demanding computational cost and large data movement.To address this, we propose HyFlexPIM, a novel mixed-signal processingin-memory (PIM) accelerator for inference that flexibly utilizes both single-level cell (SLC) and multi-level cell (MLC) RRAM technologies to trade-off accuracy and efficiency.HyFlexPIM achieves efficient dual-mode operation by utilizing digital PIM for highprecision and write-intensive operations while analog PIM for high parallel and low-precision computations.The analog PIM further distributes tasks between SLC and MLC PIM operations, where a single analog PIM module can be reconfigured to switch between two operations (SLC/MLC) with minimal overhead (<1% for area & energy).Critical weights are allocated to SLC RRAM for high accuracy, while less critical weights are assigned to MLC RRAM to maximize capacity, power, and latency efficiency.However, despite employing such a hybrid mechanism, brute-force mapping on hardware fails to deliver significant benefits due to the limited proportion of weights accelerated by the MLC and the noticeable degradation in accuracy.To maximize the potential of our hybrid hardware architecture, we propose an algorithm co-optimization technique, called gradient redistribution, which uses Singular Value Decomposition (SVD) to decompose and truncate matrices based on their importance, then fine-tune them to concentrate significance into a small subset of weights.By doing so, only 5-10% of the weights have dominantly large gradients, making it favorable for HyFlexPIM by minimizing the use of expensive SLC RRAM while maximizing the efficient MLC RRAM.Our evaluation shows that HyFlexPIM significantly enhances computational throughput and energy efficiency, achieving maximum 1.86× and 1.45× higher than state-of-the-art methods. Chang Eun Song, Priyansh Bhatnagar, Zihan Xia 0002, Nam Sung Kim, Tajana Rosing, Mingu Kang |
ISCA | 3 |
| 2025 | Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE ServingabstractAs Large Language Models (LLMs) continue to evolve, Mixture of Experts (MoE) architecture has emerged as a prevailing design for achieving state-of-the-art performance across a wide range of tasks.MoE models use sparse gating to activate only a handful of expert sub-networks per input, achieving billion-parameter capacity with inference costs akin to much smaller models.However, such models often pose challenges for hardware deployment due to the massive data volume introduced by the MoE layers.To address the challenges of serving MoE models, we propose Stratum, a system-hardware co-design approach that combines the novel memory technology Monolithic 3D-Stackable DRAM (Mono3D DRAM), near-memory processing (NMP), and GPU acceleration.The logic and Mono3D DRAM dies are connected through hybrid bonding, whereas the Mono3D DRAM stack and GPU are interconnected via silicon interposer.Mono3D DRAM offers higher internal bandwidth than HBM thanks to the dense vertical interconnect pitch enabled by its monolithic structure, which supports implementations of higher-performance near-memory processing.Furthermore, we tackle the latency differences introduced by aggressive vertical scaling of Mono3D DRAM along the 𝑧-dimension by constructing internal memory tiers and assigning data across layers based on * Equal contribution Yue Pan 0009, Zihan Xia 0002, Po-Kai Hsu, Lanxiang Hu, Hyungyo Kim, Janak Sharda, Minxuan Zhou, Nam Sung Kim, Shimeng Yu, Tajana Rosing, Mingu Kang |
MICRO | 2 |
| 2025 | Exploiting Chiplet Integration Technology for Fast High-Capacity DRAM ModulesabstractAs the end of Moore’s law approaches, chiplet integration technology (or chiplet technology) has emerged to revolutionize future semiconductor chip design. Chiplet technology provides unique advantages over 3-D-stacking technology, including a more cost-efficient and thermal-friendly integration of heterogeneous technologies. Although chiplet technologies have already begun to be used by the latest commercial chips, they have not been explored for commodity dynamic random access memory (DRAM) design yet. Harnessing its advantages for DRAM for the first time, this article evaluates the feasibility of chiplet-based DRAM architecture, considering various physical and electrical constraints imposed by a standard chiplet interface [i.e., universal chiplet interconnect express (UCIe)]. We further explore the DIMM architectures that simplify module packaging and assembly, leading to reductions in total die size and overall costs. The comprehensive cross-level analysis (i.e., device, circuit, chip, and system levels) shows that chiplet-based DRAM reducest_RCD+t_CAS, latency-critical DRAM timing parameters, by$1.32\times $–$1.39\times $, at the same energy consumption. In addition, a$1.39\times $–$2.28\times $improvement int_RRDis obtained. The reduced DRAM timing parameters improve the overall system performance by up to 8.8%–24.7% (geomean 3.4%–8.4%) in real-life benchmarks. The chiplet-based heterogeneous integration achieves a$1.27\times $higher chip-level yield compared with the monolithic chip, along with up to 10% reduction in overall cost compared with traditional DIMMs at emerging process technologies. Zihan Xia 0002, Chihun Song, Ram Krishna, Ashita Victor, Srujan Penta, Muhannad S. Bakir, Elyse Rosenbaum, Nam Sung Kim, Mingu Kang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | A 4.86-pJ/b Energy-Efficient Fully Parallel Stochastic LDPC Decoder With Two-Stage Shared MemoryabstractThe complex calculations of the low-density parity-check (LDPC) decoder result in significant energy and hardware consumption. To solve the challenge, this brief describes a fully parallel stochastic LDPC decoder with a two-stage shared memory (TSM) variable node (VN). To enhance cost efficiency, our design incorporates a shared low-cost random number generator (RNG) for all 2160 channels. We introduce a TSM VN function, which demonstrates faster convergence and reduced hardware overhead in comparison with the existing methods. We have taped out the (2160, 1760) stochastic LDPC decoder in the 55-nm process. The measure results exhibit that the proposed design achieves a throughput of 57.6 Gb/s, an efficiency of 33.68 Gb/s/mm2, and a power efficiency of 4.86 pJ/bit, underlining superior performance in terms of decoding throughput, hardware efficiency, and energy conservation. Yakun Zhou, Jienan Chen, Yizhuo Zhou, Zihan Xia 0002, Chuan Zhang 0001, Runsheng Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | Neural Synaptic Plasticity-Inspired Computing: A High Computing Efficient Deep Convolutional Neural Network AcceleratorabstractDeep convolutional neural networks (DCNNs) have achieved state-of-the-art performance in classification, natural language processing (NLP), and regression tasks. However, there is still a great gap between DCNNs and the human brain in terms of computation efficiency. Inspired by neural synaptic plasticity and stochastic computing (SC), we propose neural synaptic plasticity-inspired computing (NSPC) to simulate the human brain's neural network activity for inference tasks with simple logic gates. The multiplication and accumulation (MAC) is transformed by the wire connectivity in NSPC, which only requires bundles of wires and small width adders. To this end, the NSPC imitates the structure of neural synaptic plasticity from a circuit wires connection perspective. Furthermore, from the principle of NSPC, we use a data mapping method to convert the convolution operations to matrix multiplications. Based on the methodology of NSPC, fully-pipelined and low latency architecture is designed. The proposed NSPC accelerator exhibits high hardware efficiency while maintaining a comparable network accuracy level. The NSPC based DCNN accelerator (NSPC-CNN) processes DCNN at 1.5625M images/s with a power dissipation of 15.42 W and an area of 36.4 mm2. The NSPC based deep neural network (DNN) accelerator (NSPC-DNN) that implements three fully connected layers DNN consumes only 6.6 mm2 area and 2.93 W power, and achieves a throughput of 400M images/s. Compared with conventional fixed-point implementations, the NSPC-CNN achieves 2.77× area efficiency, 2.25× power efficiency; the proposed NSPC-DNN exhibits 2.31× area efficiency and 2.09× power efficiency. Zihan Xia 0002, Jienan Chen, Qiu Huang, Jinting Luo, Jianhao Hu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | Neural Synaptic Plasticity-Like Computing: An Ultra-Low Cost Approach for Artificial Neural Networks ImplementationabstractArtificial neural networks (ANNs) have gained state-of-the-art results in classification and regression tasks. However, there is still great gap between ANNs and human brain in terms of computation efficiency. In this work, we proposed the neural synaptic plasticity-like computing (NSPC) to simulate the neural network activity for inference task with ultra-simple logic gates. The multiplication of weight in traditional ANNs is transformed by the wire connectivity in NSPC, which requires only bundle of wires without any logics. To this end, the NSPC imitates the structure of neural synaptic plasticity from a circuit wires connection perspective. The proposed NSPC exhibits comparable inference accuracy with low hardware cost. According to the implementation results, the NSPC requires only 28% logic gate resources of conventional ANNs scheme, 114% throughput improvement and 8.454 times better hardware efficiency on the average. Zihan Xia 0002, Jienan Chen, Shaoxia He, Shuai Li 0016 |
ISCAS | 1 |