EDBT 2026 Demo / reviewers in the wild / expert
Pengcheng Feng
dblp:141/2051
· DBLP profile ↗
7ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A PulseWidth-IN-PulseWidth-Out Universal Nonlinear Processing Element for Time-Domain In-Memory Computing SystemsabstractTime-Domain In-Memory Computing (TD-IMC) has emerged as a promising analog computing architecture for edge AI applications. However, the lack of developed hardware operators, especially general nonlinear operators, necessitates frequent cross-domain data transmission in practical TD-IMC systems, significantly reducing energy efficiency. In this work, we propose a PulseWidth-IN-PulseWidth-OUT Universal Nonlinear Processing Element (PIPO-UNPE) to address the challenges of nonlinear processing in analog computing. By implementing an RRAM-based two-layer ReLU network, the PIPO-UNPE performs universal nonlinear operations entirely in the time domain. Algorithmically, we introduce Dynamic Loss-Responsive Subset Enhancement (DLRSE) to boost the performance of this low-cost network in function approximation tasks. From a hardware perspective, we design an RRAM-based pulse-driven programmable current source and a low-latency dispersion comparator-based voltage-to-time converter (VTC) to enhance both the energy efficiency and precision of the PIPO-UNPE. Hybrid simulations reveal that the PIPO-UNPE consumes 912 uW of power while delivering a throughput of $\mathbf{1 0 M}$ NOPS (Nonlinear Operations Per Second). Incorporating the PIPO-UNPE into the TD-IMC accelerator can increase energy efficiency by a factor of 9.5 to 25, keeping the accuracy loss below 0.1%. Pengcheng Feng, Rongxuan Shen, Huaxiang Lu, Xiaoxin Xu |
DAC | 2 |
| 2025 | CSIA-UIM: A Universal Ising Machine Based on CIM-friendly Spring-Ising AlgorithmabstractIsing machines are specialized processors designed to solve combinatorial optimization problems through the physical evolution of the Ising graphs. However, conventional Ising machines are restricted to solving problems with specific graph topologies and are further hindered by the inherent memory wall of the von Neumann architecture. In this work, we propose a Computing-In-Memory-friendly Spring-Ising Algorithm (CSIA) to solve Ising models with arbitrary graph topologies. Building on CSIA, we introduce a universal Ising machine (CSIA-UIM) capable of fully parallel spin updates. The CSIA-UIM adopts a Time-Domain Computing-In-Memory architecture, with key modules including the Generalized Momentum Update Module (GMUM), Generalized Coordinate Update Module (GCUM), and Hardware Inelastic Wall (HIW) working collaboratively in a parallel pipeline fashion. Hybrid simulation results show that CSIA-UIM achieves speed improvements of 468×, 64×, 2.5×, compared to GPU, RRAM-based, and CMOS-based universal Ising machines, respectively, when solving a fully connected Ising model with 1,000 spins. Zhelong Jiang, Pengcheng Feng, Jinke Yu, Rongxuan Shen, Huaxiang Lu |
ISCAS | 3 |
| 2025 | A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced DataflowabstractFPGA accelerators for lightweight convolutional neural networks (LWCNNs) have recently attracted significant attention. Most existing LWCNN accelerators focus on single-Computing-Engine (CE) architecture with local optimization. However, these designs typically suffer from high on-chip/off-chip memory overhead and low computational efficiency due to their layer-by-layer dataflow and unified resource mapping mechanisms. To tackle these issues, a novel multi-CE-based accelerator with balanced dataflow is proposed to efficiently accelerate LWCNN through memory-oriented and computing-oriented optimizations. Firstly, a streaming architecture with hybrid CEs is designed to minimize off-chip memory access while maintaining a low cost of on-chip buffer size. Secondly, a balanced dataflow strategy is introduced for streaming architectures to enhance computational efficiency by improving efficient resource mapping and mitigating data congestion. Furthermore, a resource-aware memory and parallelism allocation methodology is proposed, based on a performance model, to achieve better performance and scalability. The proposed accelerator is evaluated on Xilinx ZC706 platform using MobileNetV2 and ShuffleNetV2. Implementation results demonstrate that the proposed accelerator can save up to 68.3% of on-chip memory size with reduced off-chip memory access compared to the reference design. It achieves an impressive performance of up to 2092.4 FPS and a state-of-the-art MAC efficiency of up to 94.58%, while maintaining a high DSP utilization of 95%, thus significantly outperforming current LWCNN accelerators. Pengcheng Feng, Jixing Li, Rongxuan Shen, Huaxiang Lu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | Correction to: Underwater target detection with an attention mechanism and improved scale
Long Yu 0001, Shengwei Tian, Pengcheng Feng, Xin Ning 0001 |
Multim. Tools Appl. | 4 |
| 2021 | Underwater target detection with an attention mechanism and improved scale
Long Yu 0001, Shengwei Tian, Pengcheng Feng, Xin Ning 0001 |
Multim. Tools Appl. | 4 |
| 2017 | A Taxi Order Dispatch Model based On Combinatorial OptimizationabstractTaxi-booking apps have been very popular all over the world as they provide convenience such as fast response time to the users. The key component of a taxi-booking app is the dispatch system which aims to provide optimal matches between drivers and riders. Traditional dispatch systems sequentially dispatch taxis to riders and aim to maximize the driver acceptance rate for each individual order. However, the traditional systems may lead to a low global success rate, which degrades the rider experience when using the app. In this paper, we propose a novel system that attempts to optimally dispatch taxis to serve multiple bookings. The proposed system aims to maximize the global success rate, thus it optimizes the overall travel efficiency, leading to enhanced user experience. To further enhance users' experience, we also propose a method to predict destinations of a user once the taxi-booking APP is started. The proposed method employs the Bayesian framework to model the distribution of a user's destination based on his/her travel histories. Lingyu Zhang 0001, Yue Min, Guobin Wu 0001, Pengcheng Feng, Pinghua Gong, Jieping Ye |
KDD | 6 |
| 2013 | ICAMF: Improved Context-Aware Matrix Factorization for Collaborative FilteringabstractContext-aware recommender system (CARS) can provide more accurate rating predictions and more relevant recommendations by taking into account the contextual in-formation. Yet the state-of-the-art context-aware matrix factorization approaches only consider the influence of con-textual information on item bias. Tensor factorization based Multiverse Recommendation deals with the contextual in-formation by incorporating user-item-context interaction into recommendation model. However, all of these approaches cannot fully capture the influence of contextual information on the rating. In this paper, we propose two improved context-aware matrix factorization approaches to fully capture the influence of contextual information on the rating. Both of the baseline predictors (user bias and item bias) and user-item-context interaction are fully concerned. Experimental results on three semi-synthetic datasets and one real world dataset show that the two proposed approaches outperform Multiverse Recommendation and the state-of-the-art context-aware matrix factorization methods in prediction performance. Jiyun Li, Pengcheng Feng, Juntao Lv |
ICTAI | 2 |