EDBT 2026 Demo / reviewers in the wild / expert
Rongxuan Shen
dblp:227/8738
· DBLP profile ↗
4ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A PulseWidth-IN-PulseWidth-Out Universal Nonlinear Processing Element for Time-Domain In-Memory Computing SystemsabstractTime-Domain In-Memory Computing (TD-IMC) has emerged as a promising analog computing architecture for edge AI applications. However, the lack of developed hardware operators, especially general nonlinear operators, necessitates frequent cross-domain data transmission in practical TD-IMC systems, significantly reducing energy efficiency. In this work, we propose a PulseWidth-IN-PulseWidth-OUT Universal Nonlinear Processing Element (PIPO-UNPE) to address the challenges of nonlinear processing in analog computing. By implementing an RRAM-based two-layer ReLU network, the PIPO-UNPE performs universal nonlinear operations entirely in the time domain. Algorithmically, we introduce Dynamic Loss-Responsive Subset Enhancement (DLRSE) to boost the performance of this low-cost network in function approximation tasks. From a hardware perspective, we design an RRAM-based pulse-driven programmable current source and a low-latency dispersion comparator-based voltage-to-time converter (VTC) to enhance both the energy efficiency and precision of the PIPO-UNPE. Hybrid simulations reveal that the PIPO-UNPE consumes 912 uW of power while delivering a throughput of $\mathbf{1 0 M}$ NOPS (Nonlinear Operations Per Second). Incorporating the PIPO-UNPE into the TD-IMC accelerator can increase energy efficiency by a factor of 9.5 to 25, keeping the accuracy loss below 0.1%. Pengcheng Feng, Rongxuan Shen, Huaxiang Lu, Xiaoxin Xu |
DAC | 5 |
| 2025 | CSIA-UIM: A Universal Ising Machine Based on CIM-friendly Spring-Ising AlgorithmabstractIsing machines are specialized processors designed to solve combinatorial optimization problems through the physical evolution of the Ising graphs. However, conventional Ising machines are restricted to solving problems with specific graph topologies and are further hindered by the inherent memory wall of the von Neumann architecture. In this work, we propose a Computing-In-Memory-friendly Spring-Ising Algorithm (CSIA) to solve Ising models with arbitrary graph topologies. Building on CSIA, we introduce a universal Ising machine (CSIA-UIM) capable of fully parallel spin updates. The CSIA-UIM adopts a Time-Domain Computing-In-Memory architecture, with key modules including the Generalized Momentum Update Module (GMUM), Generalized Coordinate Update Module (GCUM), and Hardware Inelastic Wall (HIW) working collaboratively in a parallel pipeline fashion. Hybrid simulation results show that CSIA-UIM achieves speed improvements of 468×, 64×, 2.5×, compared to GPU, RRAM-based, and CMOS-based universal Ising machines, respectively, when solving a fully connected Ising model with 1,000 spins. Zhelong Jiang, Pengcheng Feng, Jinke Yu, Rongxuan Shen, Huaxiang Lu |
ISCAS | 5 |
| 2025 | A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced DataflowabstractFPGA accelerators for lightweight convolutional neural networks (LWCNNs) have recently attracted significant attention. Most existing LWCNN accelerators focus on single-Computing-Engine (CE) architecture with local optimization. However, these designs typically suffer from high on-chip/off-chip memory overhead and low computational efficiency due to their layer-by-layer dataflow and unified resource mapping mechanisms. To tackle these issues, a novel multi-CE-based accelerator with balanced dataflow is proposed to efficiently accelerate LWCNN through memory-oriented and computing-oriented optimizations. Firstly, a streaming architecture with hybrid CEs is designed to minimize off-chip memory access while maintaining a low cost of on-chip buffer size. Secondly, a balanced dataflow strategy is introduced for streaming architectures to enhance computational efficiency by improving efficient resource mapping and mitigating data congestion. Furthermore, a resource-aware memory and parallelism allocation methodology is proposed, based on a performance model, to achieve better performance and scalability. The proposed accelerator is evaluated on Xilinx ZC706 platform using MobileNetV2 and ShuffleNetV2. Implementation results demonstrate that the proposed accelerator can save up to 68.3% of on-chip memory size with reduced off-chip memory access compared to the reference design. It achieves an impressive performance of up to 2092.4 FPS and a state-of-the-art MAC efficiency of up to 94.58%, while maintaining a high DSP utilization of 95%, thus significantly outperforming current LWCNN accelerators. Pengcheng Feng, Jixing Li, Rongxuan Shen, Huaxiang Lu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2018 | Tracking the multi-well surface dynamometer card state for a sucker-rod pump by using a particle filterabstractFor a non‐linear sucker‐rod pumping system, a surface dynamometer card estimation algorithm based on a particle filter is presented. The dynamometer card is a plot of the polished rod load at various positions of a pump stroke. Since the polished rod load measured by a load sensor is frequently affected by drift problems, a local characteristic correlation method is proposed while building the state‐space model for the pumping unit. The local characteristic correlation method makes the system insensitive to load drift problems. Moreover, the prior data recorded from different wells are used to construct the importance density. To make the k ‐time importance density closer to the real posterior distribution, current measurement information is used. The performance of the proposed algorithm is evaluated on the actual operating data of a Xinjiang oil field containing typical daily production activities that can cause sudden system state changes. The results show that the proposed algorithm can adapt to sudden changes of the underground environment caused by various human factors, and it can provide robust estimation for multi‐well long‐term state tracking. Guoliang Gong, Rongxuan Shen, Wenyu Mao, Huaxiang Lu |
IET Commun. | 3 |