VLDB 2026 Research / reviewers in the wild / expert
Lei Chen 0001
dblp:09/3666-1
· DBLP profile ↗
11ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0003-3568-4468ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dual-Thread Deflate/Inflate Accelerator With Multicheckpoint Control With High Throughput and Compression Ratio for Bandwidth-Efficient SystemsabstractWith the exponential growth of data volumes in AI training and prediction systems, the cost and resource demands of data transmission have emerged as critical challenges. Lossless data compression effectively reduces data size, transmission bandwidth, and latency while preserving data integrity. This article presents a fully pipelined lossless CODEC integrating Deflate compression and Inflate decompression accelerators. The proposed Deflate implementation employs match filtering and pair merging strategies to enhance compression ratios. We introduce three key innovations for the Inflate decompressor: 1) a dual-thread architecture with multicheckpoint control; 2) optimized end-of-block (EOB) handling in Huffman coding; and 3) a rewinding mechanism in LZ77 decoding. FPGA implementation results demonstrate that our Deflate compressor achieves 16 bytes/cycle throughput with an average compression ratio of 2.26, surpassing state-of-the-art implementations. The 28 nm CMOS implementation shows Inflate decompression throughputs of 1431.85 MB/s (dynamic Huffman) and 1324.26 MB/s (static Huffman) on the Calgary Corpus dataset. Notably, our 28 nm CMOS-based decompressor achieves$1.16\times $higher throughput than recent 14 nm implementations in spite of operating at half their maximum frequency. Yiwei Luo, Jiaqi Ouyang, Lei Chen 0001, Fengwei An |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | ReHIT: Reconfigurable High-Radix Iterative-Taylor Architecture for Ultraprecise Logarithm/Exponential Functions in FPGA-Based Softmax AcceleratorsabstractThe softmax function, as a pivotal component in neural network accelerators, imposes stringent demands on the precision-efficiency tradeoff for logarithmic and exponential computations. This article presents reconfigurable high-radix iterative-Taylor (ReHIT) architecture, a novel hardware framework that synergistically integrates high-radix iterative normalization with optimized Taylor expansion to achieve subunit-in-the-last-place (ULP) precision in floating-point transcendental functions. Our key innovation lies in the hierarchical pretreatment mechanism where high-radix iterations (radix-256/512) systematically decompose input operands into normalized subdomains, enabling subsequent quadratic Taylor approximations with guaranteed convergence. This codesign methodology reduces polynomial orders by 33% compared to conventional approaches while eliminating resource-intensive division operations through shift-and-add transformations. The implemented ReHIT-logarithm (ReHIT-L) and ReHIT-exponential (ReHIT-E) modules demonstrate configurable precision scaling from half to double precision (FP16/32/64), validated through exhaustive error analysis over 1 000 000 random test vectors with worst case errors bounded at 0.78 ULP. Field-programmable gate array (FPGA) implementations on Arria 10/Virtex-7 platforms can achieve up to 13.8% logic resources reduction and 32.3% latency improvement over state-of-the-art designs, with post-synthesis results in Taiwan Semiconductor Manufacturing Company (TSMC) 28-nm showing up to$1.34\times $giga operations per second (GOPS)/W energy efficiency and$3.97\times $GOPS/mm2area efficiency for softmax acceleration. The reconfigurable pipeline of the architecture permits dynamic precision/throughput adaptation, particularly beneficial for quantized neural networks requiring FP16–FP32 hybrid precision. Yangyi Zhang 0002, Lei Chen 0001, Fengwei An |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | Live Demonstration: A 1920×1080 129fps 4.3pJ/pixel Stereo-Matching Processor for Low-power ApplicationsabstractThis demonstration presents an advanced stereo vision system with high energy efficiency. An ov5640 binocular camera, operating at a maximum of 30 frames/second with FHD (1920×1080) resolution, is employed to capture image pairs. A Spartan-7 FPGA rectifies these images with a calibration map matrix and then channels the pixel stream to the stereo-matching processor in a 28nm CMOS process for depth estimation. The resulting depth map, crucial for tasks like obstacle detection and navigation, is displayed in real-time on the monitor for low-power stereo vision applications. Zhuoyu Chen, Shengming Zhou, Pingcheng Dong, Fengwei An, Lei Chen 0001 |
ISCAS | 7 |
| 2024 | Live Demonstration: A Video Denoising Co-processor with Non-local Means Algorithm for FHD 30fps Image SensorabstractIn this demonstration, a non-local means (NLM) video denoising co-processor with data reuse scheme and dual-clock domain for high resolution image sensor is presented. With an OV5640 camera, the real-time denoising processor can be performed at 30 frames/second for full high definition (FHD 1920×1080) RAW video, a debayer filter then decodes the RAW format to RGB format. Ruoheng Yao, Shengming Zhou, Zhiyue Gao, Yangyi Zhang 0002, Yiwei Luo, Lei Chen 0001, Fengwei An |
ISCAS | 6 |
| 2024 | Stereo Matching Accelerator With Re-Computation Scheme and Data-Reused Pipeline for Autonomous VehiclesabstractBinocular stereo vision is a depth estimation technique by imitating human eyes. It is widely used in various fields, such as self-driving cars, SLAM, and 3D reconstruction. However, designing a hardware architecture that can balance resource utilization, processing speed, and estimation accuracy remains a significant challenge. This paper proposes a compact and efficient hardware-based design that incorporates linear fitting-based cost fusion, disparity optimization with subpixel interpolation, and multi-directional occlusion filling techniques. Firstly, a gradient-enhanced pipelined matching costs architecture with a resource-saving scheme and re-computation paradigm is proposed to improve the accuracy of edge information. Then, we approximate the nonlinear exponential function by linear fitting to save the hardware resource. Moreover, we design the subpixel interpolation with an SRT radix-4 divider to refine the disparity, which significantly enhances the accuracy of the disparity map in real situations. Finally, we proposed a resource-reused architecture for synchronous hole filling and median filter in post-processing. The disparity map quality of the proposed architecture is evaluated on KITTI2015 datasets, which delivers leading accuracy compared to other state-of-the-art works. The architecture has been successfully implemented and demonstrated on the Stratix-V FPGA platform and achieved 54 frames per second operating at 112 MHz under a resolution of$1920\times 1080$. Xiwei Fang, Yunhao Ma, Pingcheng Dong, Zhuoyu Chen, Lei Chen 0001, Fengwei An |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | A Spatio-Temporal Video Denoising Co-Processor With Adaptive CodecabstractWith the increasing demand for high-resolution video and real-time processing, the limited efficiency of video-denoising algorithms has become a critical factor. This paper proposes a spatio-temporal video denoising co-processor to suppress an image sequence’s spatial and temporal noise. Temporal denoising is achieved by merging the current and previous frames at the pixel level in which the current frame is processed by a spatial filter. After exploiting noise estimation and motion detection, the Wiener filter calculates the merge ratio. Rather than buffering the entire previous frame, the JPEG-like codec can dynamically adjust the compression ratio through a predefined quantization table to satisfy the designed on-chip storage. The experimental results demonstrate that the spatio-temporal denoising co-processor can effectively eliminate the fluctuation of the grayscale value of the noise in videos. Simultaneously, the adaptive codec can reduce the storage space consumption for the frame buffer by at least 80% of the original size. To the best of our knowledge, this is the first fully integrated spatio-temporal denoising co-processor without any external memory. Additionally, the grayscale, RGB, and RAW versions of the co-processor are also implemented on the Stratix V FPGA platform and synthesized in 28nm CMOS technology. Yichen Ouyang, Ruoheng Yao, Zhuoyu Chen, Lei Chen 0001, Fengwei An |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | Anti-Aliasing and Anti-Color-Artifact Demosaicing for High-Resolution CMOS Image SensorabstractDemosaicing is a technique that reconstructs an RGB image from fragmentary color samples sensed by the image sensor. The color filter array (CFA), which is placed over the image sensor, determines the color of each pixel. The most used color filter array is the Bayer CFA. This paper proposes an Anti-Aliasing and Anti-Color-Artifact Demosaicing (AAACA) algorithm for the Bayer pattern and the resource-efficient very-large-scale integration (VLSI) architecture for the proposed algorithm. The AAACA comprises an anti-aliasing approach and a color artifacts filter named color difference-based median filter (CDMF). Compared to the traditional demosaicing methods that equally treat pixels on and not on edges, the AA reconstructs the pixels on edges with different strategies from those not on edges, significantly removing the aliasing around recovered edges. Then the CDMF is utilized to remove the color artifacts of the reconstructed RGB image based on the color difference after median filtering. We respectively simulate the quantitative evaluation and subjective visual quality on McMaster and Kodak datasets. Our experiments reveal that the proposed AAACA algorithm can significantly remove the visual aliasing around recovered edges and greatly reduce color artifacts in demosaiced images compared to the state-of-art demosaicing algorithms. The proposed VLSI architectures can achieve superior visual qualities compared with the previous VLSI implementations under the same process technology conditions with 180nm CMOS technology, an image resolution of$1280\times 720$(HD), and a working frequency of 200MHz. Yangyi Zhang 0002, Zizhao Peng, Lei Chen 0001, Fengwei An |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | A 4.29nJ/pixel Stereo Depth Coprocessor With Pixel Level Pipeline and Region Optimized Semi-Global Matching for IoT ApplicationabstractThe semi-global matching (SGM) algorithm in stereo vision is a well-known depth-estimation method since it can generate dense and robust disparity maps. However, the real-time processing and low power dissipation, the specifications of the Internet-of-Thing (IoT) applications, are challenging for their computational complexity. In this paper, we propose a hardware-oriented SGM algorithm with pixel-level pipeline and region-optimized cost aggregation for high-speed processing and low hardware-resource usage. Firstly, the matching costs in a region are integrated with an optimization strategy to significantly reduce memory usage and improve the processing speed of the cost aggregation. Then, a two-layer parallel two-stage pipeline (TPTP) architecture, which enables pixel-level processing, is designed to calculate two directions (0° and 135°) aggregation to further solve the crucial computational bottleneck of the SGM algorithm. Finally, the architecture is demonstrated on a low-cost XILINX Spartan-7 device and an advanced Stratix-V FPGA device for VGA ($640\times 480$) depth estimation. The experimental results show that the proposed architecture with compact hardware architecture also ensures accuracy. The pixel-level pipeline architecture enables a processing speed of 355 frames per second (fps) at 109MHz on the Spartan-7 FPGA device and 508 fps at 156MHz on the Stratix-V FPGA. Besides, the coprocessor respectively achieves an energy efficiency of 4.74 nJ/pixel with a power dissipation of 517mW and 4.29nJ/pixel with a power dissipation of 669mW on these two FPGAs. Pingcheng Dong, Zhuoyu Chen, Zhuoao Li, Yuzhe Fu, Lei Chen 0001, Fengwei An |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2018 | A Hardware Architecture for Cell-Based Feature-Extraction and Classification Using Dual-Feature SpaceabstractMany computer-vision and machine-learning applications in robotics, mobile, wearable devices, and automotive domains are constrained by their real-time performance requirements. This paper reports a dual-feature-based object recognition coprocessor that exploits both histogram of oriented gradient (HOG) and Haar-like descriptors with a cell-based parallel sliding-window recognition mechanism. The feature extraction circuitry for HOG and Haar-like descriptors is implemented by a pixel-based pipelined architecture, which synchronizes to the pixel frequency from the image sensor. After extracting each cell feature vector, a cell-based sliding window scheme enables parallelized recognition for all windows, which contain this cell. The nearest neighbor search classifier is, respectively, applied to the HOG and Haar-like feature space. The complementary aspects of the two feature domains enable a hardware-friendly implementation of the binary classification for pedestrian detection with improved accuracy. A proof-of-concept prototype chip fabricated in a 65-nm SOI CMOS, having thin gate oxide and buried oxide layers (SOTB CMOS), with 3.22-mm2core area achieves an energy efficiency of 1.52 nJ/pixel and a processing speed of 30 fps for 1024 × 1616-pixel image frames at 200-MHz recognition working frequency and 1-V supply voltage. Furthermore, multiple chips can implement image scaling, since the designed chip has image-size flexibility attributable to the pixel-based architecture. Fengwei An, Xiangyu Zhang 0002, Aiwen Luo, Lei Chen 0001, Hans Jürgen Mattausch |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Resource-Efficient Object-Recognition Coprocessor With Parallel Processing of Multiple Scan Windows in 65-nm CMOSabstractObject recognition offers a more general implementation for vision-based applications. This paper reports a resource-efficient recognition coprocessor with embedded cell-based simplified speeded up robust feature descriptor extraction unit and parallel scan-window (SW) recognition engine, applicable for various mobile scenarios and image sensor types. The feature extraction circuitry with pixel-based pipelined architecture describes the target objects among complex backgrounds, only relying on the pixel frequency from the image sensor. A cell-based SW algorithm enables parallelized recognition in multiple SWs and compatibility to different image sizes. The proposed hardware-friendly object-recognition coprocessor was implemented in 65-nm Silicon on thin BOX CMOS technology with 1.26 mm2core area and can operate down to low supply voltage of 0.5 V. For video graphics array image sizes, the energy efficiency is determined as 910 μJ per frame at 200 MHz and 1-V supply voltage. The coprocessor's classification performance is demonstrated for pedestrian and car detection. Aiwen Luo, Fengwei An, Xiangyu Zhang 0002, Lei Chen 0001, Hans Jürgen Mattausch |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Dynamically reconfigurable system for LVQ-based on-chip learning and recognitionabstractArtificial neural networks implement a simplified model of the human brain and thus specialize on pattern recognition. As an alternative to conventional single-instruction-multiple-data (SIMD) solutions with massive parallelism for self-organizing-map (SOM) neural network models, we report resource-efficient hardware architecture for 1-chip implementation of the learning vector quantization (LVQ) neural network algorithm, which is a variant of SOM. Dynamic configurability for two operation modes is realized through the same circuitry for recognition based on nearest-neighbor matching and on-chip learning based on error back-propagation. Switching between learning and recognition modes is carried out by a pipeline with multiplexers and parallel p-word input (P-MPPI). The multiplexers enable data-flow-path reconfiguration, resulting in a significant reduction of area and power consumption. Thus, the P-MPPI architecture achieves time-domain multiplexing between operation modes as well as area/energy-efficiency by reusing both memory arrays and arithmetic or logic units. Additionally, high flexibility for feature-vector dimension and reference-vector number allows the implementation of many different applications, including continuously adaptive neural systems, on the same hardware platform. A test chip in TSMC 65 nm CMOS has parallel 32-word inputs, 585 K-bit on-chip memory, and achieves high processing throughput of 76.8 Gbps and low power consumption of 27.92 mW (at 150 MHz, 1.0 V supply voltage). Fengwei An, Xiangyu Zhang 0002, Lei Chen 0001, Hans Jürgen Mattausch |
ISCAS | 3 |