VLDB 2026 Research / reviewers in the wild / expert
Chen Nie
dblp:295/3534
· DBLP profile ↗
14ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 6 first-author · 12 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ELSA: An Elastic Snn Inference Architecture for Efficient Neuromorphic Computing
Kang You, Chen Nie, Lee Jun Yan, Ziling Wei, Yu Feng 0007, Honglan Jiang, Zhezhi He |
ISCA | 2 |
| 2026 | NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data Processing
Chen Nie, Limin Xiao 0001, Weifeng Zhang 0003, Zhezhi He |
ISCA | 3 |
| 2026 | APU: Accelerate Point Cloud Neural Networks via Unified Processing-in-SRAM ArchitectureabstractRecent advances in deep learning have expanded point cloud applications by point-based neural networks (PNNs). However, the escalating complexity and computational demands of PNNs overwhelm conventional computers. Specialized PNN accelerators have emerged, significantly outperforming modern CPUs and GPUs. Nevertheless, existing designs remain inefficient when handling performance-critical mapping kernels of PNNs, involving diverse arithmetic functions (e.g., add, multiply, sort) across separate hardware modules. This fragmentation restricts hardware sharing and data locality, leading to area overhead, redundant data movements, and under-utilization. Therefore, a unified and efficient micro-architecture for mapping kernels is needed to enhance performance and reduce data transfers. This paper presents APU, an efficient processing-in-memory (PIM) architecture for PNN acceleration. We introduce the first unified SRAM-PIM micro-architecture that supports all mapping kernels in mainstream PNNs. Data movement is reduced through extensive on-chip memory and maximized data locality viain-situcomputing approach. At the algorithmic level, we introduce mask grouping and aggregation to eliminate costly sorting operations, enabled by hardware support for in-memory vector max-search. This refined strategy reduces computational overhead and data transfers while improving inference accuracy.We further enhance performance by exploiting parallelism across PNN operations and applying mixed-precision quantization. Evaluated on real-world PNN workloads, APU outperforms the state-of-the-art accelerator by 2.54× in speedup and 4.54× in energy saving. Chen Nie, Kang You, Yu Feng 0007, Limin Xiao 0002, Weifeng Zhang 0003, Zhezhi He |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | BiNeuroRAM: Energy-Efficient ReRAM-Based PIM for Accurate Bipolar Spiking Neural Network AccelerationabstractReRAM is a promising non-volatile memory for neuromor-phic accelerators, yet it faces challenges such as high sensing power and accuracy degradation. This work proposes BiNeuroRAM, a novel spiking neural network (SNN) accelerator leveraging ReRAM-based processing-in-memory (PIM), with three key contributions: (1) It is the first to support higher-accuracy spike-tracing bipolar-integrate-and-fire (ST-BIF) neurons, achieving 80.9% accuracy on ImageNet, 8.4% higher than the previous state-of-the-art; (2) It introduces a low-power voltage sense amplifier (LPVSA) that reduces ReRAM read power by 14.7~58.2×, enhancing energy efficiency; (3) It employs an asynchronous micro-architecture that fully exploits the event-driven nature of SNNs. Experimental results show that BiNeuroRAM improves throughput density and energy efficiency by 2.08× and 2.09× on ImageNet with ResNet-18, compared to traditional integrate-and-fire (IF) neuron-based SNN accelerators. Jun Yan Lee, Chen Nie, Kang You, Yueyang Jia, Zhezhi He |
DAC | 2 |
| 2025 | PICK: An SRAM-based Processing-in-Memory Accelerator for K-Nearest-Neighbor Search in Point CloudsabstractK-nearest neighbor (kNN) search is a fundamental operation in various point cloud applications, such as autonomous driving. However, the heavy computational intensity and memory demands of kNN search pose significant challenges for efficient implementation, especially in resource-constrained scenarios. To address these challenges, we propose PICK, a processing-in-memory (PIM) architecture designed to accelerate kNN search in point cloud applications. PICK leverages bit-serial-based PIM (BS-PIM) and customized circuits to efficiently handle key operations of kNN search: distance calculation and top-k selection. The run-time off-chip access is eliminated thanks to the large on-chip memory. For distance calculation, we introduce a bit-width clipping technique to reduce the latency of bit-serial execution with negligible accuracy degradation, providing flexible trade-offs between performance and precision. Besides, we propose a filtering-and-selection strategy that realizes approximately constant time complexity for arbitrary values of k. Furthermore, a two-stage pipeline is implemented to parallelize distance calculation and top-k search, effectively hiding latency and improving throughput. According to our experiments, PICK achieves $4.17 \times$ speedup and a $4.42 \times$ energy saving over the state-of-the-art design. Chen Nie, Liming Xiao, Weifeng Zhang 0003, Zhezhi He |
DAC | 1 |
| 2025 | MASIM: An Energy-Efficient Multi-Array Scheduler for SIMD Logic-in-Memory ArchitecturesabstractSingle instruction, multiple data (SIMD) is a popular design style of logic-in-memory (LiM) architectures, which enables memory arrays to perform logic operations to achieve low energy consumption and high throughput. To implement a target function on the data stored in memory, the function is first transformed into a netlist of the supported logic operations by logic synthesis. Then, a scheduler transforms the netlist into an instruction sequence given to the architecture, where an instruction either performs a logic operation in the netlist on memory rows within a single array or copies the data from one array to another. Most existing schedulers focus on optimizing the execution sequence of the operations to minimize the number of memory rows needed, neglecting the energy-consuming copy instructions that cannot be avoided when working with arrays with limited sizes. In this work, we focus on reducing the number of copy instructions to decrease the total energy consumption. We propose MASIM, a multi-array scheduler for SIMD logic-in-memory architectures. It consists of a priority-based scheduling algorithm and an iterative improvement process. Compared to the best existing scheduler, MASIM reduces the number of copy instructions by 63.2% on average, which leads to a 28.0% reduction in energy. The experiment also shows that MASIM can be applied to various SIMD LiM architectures, showing its wide applicability. Xingyue Qian, Chen Nie, Zhezhi He, Weikang Qian |
ICCAD | 2 |
| 2025 | PolymorPIC: Embedding Polymorphic Processing-in-Cache in RISC-V based Processor for Full-stack Efficient AI Inference
Ziling Wei, Jun Yan Lee, Chen Nie, Kang You, Zhezhi He |
MICRO | 4 |
| 2024 | PIMLC: Logic Compiler for Bit-Serial Based PIMabstractRecently, the bit-serial-based processing-in-memory (PIM) has evolved as a promising solution to enhance the computing performance of data-intensive applications, due to its high performance and programmability. However, it is absent that a compiler can automatically convert an arbitrary Boolean function (generic workload) into PIM instructions, with optimized scheduling w.r.t. the varying hardware resource and specification. To fill the gap, we develop a logic compiler for bit-serial-based PIM (PIMLC). In PIMLC, we propose a workload-resource-aware scheduling to minimize the execution latency of a given parallel workload. Thanks to PIMLC, PIM can achieve$15.55\times$and$19.03\times$speedup (geo-mean) for SRAM- and ReRAM-PIM respectively, compared to the naive scheduling of prior work. PIMLC is publicly available at: https://github.com/Intelligent-Computing-Research-GroupIPIMLC. Chenyu Tang, Chen Nie, Weikang Qian, Zhezhi He |
DATE | 2 |
| 2024 | SpikeZIP-TF: Conversion is All You Need for Transformer-based SNNabstractSpiking neural network (SNN) has attracted great attention due to its characteristic of high efficiency and accuracy. Currently, the ANN-to-SNN conversion methods can obtain ANN on-par accuracy SNN with ultra-low latency (8 time-steps) in CNN structure on computer vision (CV) tasks. However, as Transformer-based networks have achieved prevailing precision on both CV and natural language processing (NLP), the Transformer-based SNNs are still encounting the lower accuracy w.r.t the ANN counterparts. In this work, we introduce a novel ANN-to-SNN conversion method called SpikeZIP-TF, where ANN and SNN are exactly equivalent, thus incurring no accuracy degradation. SpikeZIP-TF achieves 83.82% accuracy on CV dataset (ImageNet) and 93.79% accuracy on NLP dataset (SST-2), which are higher than SOTA Transformer-based SNNs. The code is available in GitHub: https://github.com/Intelligent-Computing-Research-Group/SpikeZIP_transformer Kang You, Chen Nie, Zhijie Deng, Qinghai Guo, Zhezhi He |
ICML | 3 |
| 2024 | Social media use, social bot literacy, perceived threats from bots, and perceived bot control: a moderated-mediation modelabstractThe rapid development and widespread presence of social bots online has been transforming users' online news environment.This study adopts a human-centered perspective to investigate the impact of individuals' social media usage experiences on their social bot literacy, perception of threats posed from bots, and perceived social bot control within the context of China.We collected data from surveying 1159 Sina Weibo users and conducting interviews among 20 participants.The data were used to examine (1) the relationship between social media use and social bot literacy and perceived bot control, (2) the mediating role of social bot literacy between social media use and perceived bot control, and (3) the moderating role of perceived threat from bots in this relationship.The results of the analysis suggested a significant moderated mediation model in which social bot literacy mediated the correlation between social media use and perceived bot control and individuals' perceived threat further moderated this relationship.Specifically, for those with a higher level of perceived threat, their indirect effect of social media use was lower compared to those with a lower level of perceived threat from bots. Chen Nie |
Behav. Inf. Technol. | 2 |
| 2024 | VSPIM: SRAM Processing-in-Memory DNN Acceleration via Vector-Scalar OperationsabstractProcessing-in-Memory (PIM) has been widely explored for accelerating data-intensive machine learning computation that mainly consists of general-matrix-multiplication (GEMM), by mitigating the burden of data movements and exploiting the ultra-high memory parallelism. The two mainstreams of PIM, the analog- and digital-type, have both been exploited in accelerating machine learning workloads by numerous outstanding prior works. Currently, the digital-PIM is increasingly favored due to the broader computing support and the avoidance of errors caused by intrinsic non-idealities, e.g., process variation. Nevertheless, it still lacks further optimization considering the characteristics of the GEMM computation, including better efficient data layout and scheduling, and the ability to handle the sparsity of activations at the bit-level. To boost the performance and efficiency of digital SRAM PIM, we propose the architecture called VSPIM that performs the computation in a bit-serial fashion, with unique support of vector-scalar computing pattern. The novelties of the VSPIM can be concluded as follows: 1) support bit-serial based scalar-vector computing via ingenious parallel bit-broadcasting; 2) refine the GEMM mapping strategy and computing pattern to enhance performance and efficiency; 3) powered by the introduced scalar-vector operation, the bit-sparsity of activation is leveraged to halt unnecessary computation to maximize efficiency and throughput. Our comprehensive evaluation shows that, compared to the state-of-the-art SRAM-based digital-PIM design (Neural Cache), VSPIM can significantly boost the performance and energy efficiency by up to$8.87\times$and$4.81\times$respectively, with negligible area overhead, upon multiple representative neural networks. Chen Nie, Chenyu Tang, Jie Lin 0004, Chenyang Lv, Ting Cao 0007, Weifeng Zhang 0003, Li Jiang 0002, Xiaoyao Liang, Weikang Qian, Yanan Sun 0003, Zhezhi He |
IEEE Trans. Computers | 1 |
| 2023 | XMG-GPPIC: Efficient and Robust General-Purpose Processing-in-Cache with XOR-Majority-Graph
Chen Nie, Xianjue Cai, Chenyang Lv, Weikang Qian, Zhezhi He |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | GIM: Versatile GNN Acceleration with Reconfigurable Processing-in-MemoryabstractRecent boost of deep learning has revolutionized many machine learning tasks, including the graph neural networks (GNNs) that are specifically designed for non-Euclidean graph data. GNNs have been widely adopted in numerous real-world applications, such as the recommendation system. However, with increasingly enlarged graph size and complexity, GNN performance on conventional computers has been severely hindered by the memory bottleneck. The challenge attracts wide investigations, and the processing-in-memory (PIM) architecture arises as one of the most promising solutions. Prior works have leveraged the ReRAM crossbars as analog dot-product engines to accelerate the vector-matrix multiplications in GNN, and achieve prominent performance improvements over modern CPUs and GPUs. Nevertheless, analog computing is known to be variation-vulnerable, which hampers the inference accuracy of GNN. Besides, the mixed-signal peripherals (e.g., ADC) are hardware-expensive and specialize in dense computations, which makes the analog crossbar-based PIM not the ideal candidate for GNN inference whose computation is of great sparsity.In this work, we propose a novel digital-PIM architecture for GNN acceleration, namely GIM. Our compact yet efficient digital computing paradigm can greatly boost computing parallelism with a minimum budget. GIM integrates dedicated optimizations on both operand- and bit-sparsity, to eliminate sparse computations thus significantly boost the performance. Meanwhile, at the software level, we implement data-layout optimizations to minimize the inter-memory communications and maximize computing parallelism. Our design derives prominent performance improvements over the modern CPU, GPU, and state-of-the-art PIM-based accelerators. Compared to modern CPU and GPU, GIM averagely achieves 24485× and 778× of speedup, and 78480× and 8906× of energy reduction. Compared to the state-of-the-art PIM-based GNN accelerators ReFlip and PIMGCN, GIM averagely achieves 9.0× and 73.4× of throughput boost with 15.2× and 95.6× of efficiency improvements. Chen Nie, Guoyang Chen, Weifeng Zhang 0003, Zhezhi He |
ICCD | 1 |
| 2021 | Energy-Efficient Hybrid-RAM with Hybrid Bit-Serial based VMM SupportabstractThis work presents HRAM, a SRAM-based hybrid memory bit-cell for energy-efficient in-memory computing purpose. The HRAM bit-cell consists of conventional 6T-SRAM for static data storage, and extra one accessing transistor and capacitor for caching data temporarily then conduct the computation within the HRAM array. As the Vector-Matrix Multiplication (VMM) is the dominant operation of neural network inference, performing the VMM in bit-serial fashion is a popular method in recent works. Meanwhile, there are two variants of bit-serial VMM, digital and analog VMM respectively, which fits for varying network topology (e.g., ResNet and MobileNet correspondingly). Through designing re-configurable sensing module and peripherals, our HRAM can be configured to conduct both DVMM and AVMM efficiently. With 65nm technology, the cross-layer simulation indicates that the HRAM based in-memory computing accelerator outperforms the state-of-the-art CSRAM and MBC design by 1.94×/1.81× and 1.95×/11× respectively, in energy efficiency for ResNet-50/MobileNet-V2. Chen Nie, Jie Lin 0004, Li Jiang 0002, Xiaoyao Liang, Zhezhi He |
ACM Great Lakes Symposium on VLSI | 1 |