EDBT 2026 Demo / reviewers in the wild / expert
Chenxi Feng
dblp:239/7396
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Logic Gate Network Inference Acceleration with RISC-V Custom Instruction SetabstractLogic Gate Networks (LGNs) exploit the similarity between neural networks and logic circuit networks and replace the neurons with logic gates.Consequently, the computation inside the neurons can be replaced by Boolean operations (16 operations for two-input logic).LGNs can be implemented by logic-based instructions in processors and significantly reduce the computation overhead during inference.However, the encoding and decoding processes at the input and output stages of LGNs face efficiency challenges when using traditional RISC-V instruction sets.This limitation arises because these processes rely on one-bit operations, which cannot fully utilize the 32-bit bandwidth of standard instructions.In this work, we proposed four custom RISC-V-based instructions to accelerate the encoding and decoding processes of LGNs.An applicationspecific RISC-V processor, called RV-LGN, has been implemented on FPGA and synthesized using Synopsys® Design Compiler with the CMOS 55nm process.The custom instructions can be called via in-line assembly in C code, making RV-LGN highly promising for implementation in edge devices.Benchmark tests on MIT-BIH, MNIST, and CIFAR-10 classification tasks demonstrate that RV-LGN achieves a runtime reduction of over 87% compared to a generic RISC-V RV32IM ISA processor.Additionally, power consumption during LGN inference is significantly reduced.For the MIT-BIH dataset, the energy consumption is 0.098 µJ/Beat, while MNIST and CIFAR-10 tasks require 0.18 µJ/Image and 0.51 µJ/Image, respectively.These results highlight the superior efficiency of RV-LGN compared to other processors. Chenxi Feng, Xinyu Kang, Yuru Li, Yucong Huang, Terry Tao Ye |
CF | 2 |
| 2025 | Full-Body Pose Motion Tracking From Sparse Data via Morphology-Aware Constraints
Yinghao Yang 0002, Sanyi Zhang, Chixuan Wei, Chenxi Feng, Long Ye |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | RV-SCNN: A RISC-V Processor With Customized Instruction Set for SNN and CNN Inference Acceleration on Edge PlatformsabstractThe rapid advancement of artificial intelligence (AI) applications has driven an increasing demand for conducting inference tasks on edge devices. However, implementing computation-intensive neural networks on resource-constrained edge systems remains a significant challenge. In this article, we propose a novel processor architecture called RV-SCNN to address this challenge. The architecture is based on the RISC-V generic instruction set and incorporates various single instruction multiple data (SIMD) custom instruction extensions to accelerate the computation of spike neural networks (SNNs) and convolutional neural networks (CNNs), enabling efficient execution of complex neural network models. The core operators of the processor are shared by both SNN and CNN operations, thus supporting both computation modes. Other acceleration implementations include an internal hardware loop control unit that reduces the instruction overhead, an address calculation unit and an interlayer fusion unit that minimize the memory access overhead, as well as an image to column (IM2COL) unit that improves the computational efficiency of the$3 \times 3$convolutions in SNNs and CNNs. The custom instructions are called through inline assembly in the C program, providing higher flexibility compared to traditional ASICs and supporting custom complex SNN/CNN network structures. Compared to traditional instruction sets, the RV-SCNN processor reduces the execution time of CNNs and SNNs by over 90%. We validate the processor on FPGA platform and evaluate its performance under CMOS 55-nm process. The processor achieves an operational efficiency of 9.88 pJ/SOP in SNN network inference tasks, while the peak energy efficiency reaches 679 GOPS/W in CNN network inference. Chenxi Feng, Xinyu Kang, Qi Wang 0051, Yucong Huang, Terry Tao Ye |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | LCIQA: A Lightweight Contrastive-Learning Framework for Image Quality Assessment via Cross-Scale Consistency MinimizationabstractBlind image quality assessment (BIQA), which functions without the need for a reference image, is a challenging yet essential task in various image processing systems and downstream vision applications, ranging from semantic recognition to image enhancement. Traditionally, numerous BIQA models have been developed using supervised learning methodologies, which rely heavily on the availability and quality of ground truth data. To improve the generalization capability and robustness of these models, recent studies have explored the application of contrastive learning, aiming to enhance the quality representation capacity of model backbones through a self-supervised approach. However, the training process for contrastive learning is computationally intensive, posing significant challenges in resource-constrained environments. To mitigate this issue, we propose a Lightweight Contrastive-learning-based IQA (LCIQA) framework, designed to be efficiently trained on a single GPU without relying on ground truth data. This framework maintains a fixed vision backbone and focuses on optimizing the parameters of subsequent IQA heads through contrastive learning. To accommodate a lightweight framework, we incorporate a quality task adapter to eliminate semantic biases introduced by the features extracted from the fixed-parameter backbone. A coarse-to-fine contrastive learning strategy is then employed to train the quality regression module. Extensive experiments demonstrate the superior performance of our model in terms of both accuracy and complexity. In addition, ablation studies validate the effectiveness of each component within the proposed framework. Chenxi Feng, Xiongkuo Min, Long Ye, Yinghao Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | RV-GEMM: Neural Network Inference Acceleration with Near-Memory GEMM Instructions on RISC-VabstractGeneral Matrix Multiply (GEMM), as a fundamental operation in neural network, plays an important role in artificial intelligence and signal processing applications. In this paper, we proposed three SMID RISC-V custom instructions to accelerate GEMM computations, supporting multiple precisions including 32-bit, 16-bit and 8-bit fixed. Furthermore, we implemented address calculation and loop control units along with the GEMM acceleration module to reduce the memory access overhead. These three GEMM custom instructions, along with the near-memory optimization units, were incorporated in the RV-GEMM processor and implemented on the FPGA platform for speedup evaluation. It was also compiled in Synopsys Design Compiler with CMOS 55nm process for hardware overhead estimation. Compared to the baseline RISC-V processor, for GEMM computations under precisions of 32-bit, 16-bit and 8-bit fixed, the RV-GEMM processor achieved speedup ratios of 15.8×, 28.7× and 42.5×. The peak energy efficiency also reached 260 GOPS/W, 420 GOPS/W and 609 GOPS/W, respectively. Chenxi Feng, Bingzhen Chen, Qi Wang 0051, Yucong Huang, Terry Tao Ye |
CF | 2 |
| 2024 | Optimizing CNN Computation Using RISC-V Custom Instruction Sets for Edge PlatformsabstractBenefit from the custom instruction extension capabilities, RISC-V architecture can be optimized for many domain-specific applications. In this paper, we propose seven RISC-V SIMD (single instruction multiple data) custom instructions that can significantly optimize the convolution, activation and pool operations in CNN inference computation. More specifically, instruction CONV23 can greatly speed up the operation ofF(2 × 2, 3 × 3). With the adoption of Winograd algorithm, the number of multiplications can be reduced from 36 to 16, and the execution time is also reduced from 140 to 21 clock cycles. These custom instructions can be executed in batch mode within the acceleration module where the immediate data can be reused, so the latency and energy overhead associated with excess memory accesses can be eliminated. Using inline assembler in C language, the custom instructions can be called and compiled together with C source code. A revised RISC-V processor, RI5CY-Accel is constructed on FPGA to accommodate these custom instructions. Revised LeNet-5, VGG16 and ResNet18 model; called LeNet-Accel, VGG16-Accel and ResNet18-Accel are also optimized based on RI5CY-Accel architecture. Benchmark experiments demonstrated that the inference of LeNet-Accel, VGG16-Accel and ResNet18-Accel based on RI5CY-Accel can greatly reduce the execution latency by over 76.6%, 88.8% and 87.1%, with the total energy consumption saving of 74.8%, 87.8% and 85.1% respectively. Bingzhen Chen, Chenxi Feng, Qi Wang 0051, Terry Tao Ye |
IEEE Trans. Computers | 5 |
| 2023 | A Model-Agnostic Semantic-Quality Compatible Framework based on Self-Supervised Semantic DecouplingabstractBlind Image Quality Assessment (BIQA) is a challenging research topic that is critical for preprocessing and optimizing downstream vision tasks such as semantic recognition and image restoration. However, there has been a significant disconnect between BIQA research and other vision tasks. The primary cause of such disconnect is the incompatibility of existing BIQA models with other vision tasks, resulting in significant computational complexity. To address this issue, we propose a model-agnostic semantic-quality compatible framework that can simultaneously generate quality and semantic predictions. By incorporating a lightweight learning architecture, we demonstrate that a parameter-fixed semantic-oriented backbone can predict the perceptual quality of images as accurately as models trained end-to-end. We systematically study the major components of our framework, and our experimental results demonstrate the superiority of our model in terms of both complexity and accuracy. The source code of this work is available at https://github.com/MaxiaoyuHehe/SQCFNet. Chenxi Feng, Suiyu Zhang, Jinchi Zhu, Chang Liu 0114, Dingguo Yu |
ACM Multimedia | 2 |