EDBT 2026 Demo / reviewers in the wild / expert
Erxiang Ren
dblp:235/0220
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0003-2459-4333ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Live Demonstration: An Efficient and Compact Visuo-Tactile Perception System
Cheng Qu, Erxiang Ren, Guangyuan Xu, Fei Qiao |
ISCAS | 2 |
| 2026 | GCC-CIM: A Charge-Domain Compute-in-Memory Macro using Grouped-Row Capacitors and C-2C Ladder with Improved Multi-Bit MAC Linearity
Erxiang Ren, Daniel Zheng Fang, Qi Wei 0001, Fei Qiao |
ISCAS | 2 |
| 2025 | Dcha: Distributed-Centralized Heterogeneous Architecture Enables Efficient Multi-Task Processing for Smart SensingabstractThe rapid development of artificial intelligence (AI) has accelerated the progression of IoT technology into the smart era. Integrating AI processing capabilities into IoT devices to create smart sensing systems holds significant promise. In this work, we propose a distributed-centralized heterogeneous architecture that enables efficient multitask processing for smart sensing. This architecture improves the operational efficiency of sensing systems and enhances the deployment scalability through collaborative computing across end, edge, and center nodes. Specifically, we partition the network in traditional centralized sensing systems into several parts and perform algorithm-hardware co-design for each part on its respective deployment platform. We developed a sample design to validate the proposed architecture. By implementing a lightweight image encoder, we achieved an 88x reduction in encoder parameters and up to 9873x energy gain, facilitating deployment on resource-constrained devices. Experimental results demonstrate that the proposed architecture effectively reduces overall energy consumption by 0.0573x to 0.0889x, while maintaining robust multitask inference capabilities. Moreover, energy consumption reductions of 2.88x to 3.22x on edge nodes and 6311.56x to 10037.23x on end nodes were observed. Erxiang Ren, Cheng Qu, Zheyu Liu, Xinghua Yang, Qi Wei 0001, Fei Qiao |
DATE | 1 |
| 2025 | Denoise on Sensor: A Near-Sensor Compute-in-Memory Macro for Visual Perception Denoising via Concatnation-EliminatingabstractNoise is one of the most common and significant factors leading to image degradation. In recent years, due to the rapid development of neural networks, the performance of denoising algorithms has seen a substantial improvement. However, state-of-the-art denoising models often entail large-scale models and heavy computational requirements, making deployment challenging. Additionally, we have observed that the location of denoisers in the entire image processing pipeline has a significant impact on resource consumption and denoising effectiveness. Deploying the denoiser closer to the image acquisition stage will be more effective in separating noise from the image. In this paper, we propose a near-sensor compute-in-memory macro for visual perception denoising (Denoise on Sensor, DoS) along with its corresponding Edge Denoise U-net (EDU) architecture. DoS employs a mixed-signal circuit implementation for neural network inference, offering a notable advantage in terms of high speed and low power consumption compared to FPGA or GPU based approaches, making it feasible to deploy denoising tasks at the near-sensor edge. The simulation results show that the energy efficiency of DoS can reach 21.98 TOPS/W, and EDU deployed on DoS can achieve around 30dB PSNR and 0.83 SSIM on KODAK, BSD300 and SET14 datasets. Aolin You, Erxiang Ren, Daniel Zheng Fang, Cheng Qu, Qi Wei 0001, Fei Qiao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | NS-Engine: Near-Sensor Neural Network Engine with SRAM-Based Compute-in-Memory MacroabstractSensing devices at edge nodes are usually resource-constrained, such as limited battery capacity and physical size. As a result, there is a substantial demand for enhancing energy efficiency in these devices, which can otherwise hinder the deployment of more complex neural networks. This work proposes an energy-efficient computing engine, which is equipped with appropriate computing power to deploy medium-size neural networks for smart sensing at near-sensor edge nodes. SRAM-based Compute-in-Memory (CIM) macro and end-to-end digital controller comprise the engine. We successfully prototyped this engine using an FPGA platform and conducted a demonstration of image classification on the CIFAR-10 dataset. Furthermore, we implement the above design on a TSMC 40nm process, with post-simulation results indicating an impressive throughput of 368.64GOPS and an energy efficiency of 25.7TOPS/W at 10MHz operating frequency, given the memory capacity is 576kb. Erxiang Ren, Xinghua Yang, Qi Wei 0001, Fei Qiao |
ISCAS | 1 |
| 2024 | Cambricon-M: A Fibonacci-Coded Charge-Domain SRAM-Based CIM Accelerator for DNN InferenceabstractCharge-domain SRAM-based Computing-in-memory (CIM) proves to be a promising method for DNN inference, and benefits from avoiding data movement between computing units and memory. However, the high resolution Analog-to-Digital Converters (ADCs) dominates the energy consumption (up to 64%), limiting the energy efficiency of SRAM-CIM architectures. The main reason is the wide range of input analog values, requiring high resolution ADCs to convert the high precision averaged analog voltages into high bitwidth digital data. In this paper, to reduce the ADC overhead, we propose Cambricon-M, a novel Fibonacci-coded SRAM-based charge-domain CIM accelerator for DNN inference. Cambricon-M features the Fibonacci coding, which guarantees low density of ‘1’ in operands (i.e., the adjacent two bits of each ‘1’ are both ‘0’), narrowing the output voltage range and enabling low resolution ADCs. Further, Cambricon-M exploits the high bit-level sparsity to address the extra energy and area overhead caused by the larger bitwidth in Fibonacci coding. Specifically, Cambricon-M proposes zero-skipping methods to reduce ineffectual input/output, and the bit-slice based compression method to reduce memory capacity/bandwidth pressure. Experimental results show that Cambricon-M reduces ADC energy by 68.7%, and improves the energy efficiency 3.48× and 1.62× compared to TPUv4 and an ISAAC-based charge-domain SRAM-CIM accelerator. Hongrui Guo, Mo Zou, Yifan Hao 0001, Zidong Du, Erxiang Ren, Yang Liu 0466, Yongwei Zhao 0001, Tianrui Ma, Rui Zhang 0040, Xing Hu 0001, Fei Qiao, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002 |
MICRO | 5 |