Fupeng Chen

dblp:249/3156 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2023
0000-0002-1548-8243ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2023 AOS: An Automated Overclocking System for High-Performance CNN Accelerator Through Timing Delay Measurement on FPGA
abstract
With the inherent algorithmic error resilience of conventional neural networks (CNNs) and the worst-case design methodologies of current electronic design automation tools, overclocking-based timing speculation is a promising technique to improve the performance of CNN accelerators on FPGA by removing unnecessary timing margins. To avoid potential timing errors, timing delay measurement should be used during overclocking. However, current approaches are not yet good at measuring paths with more intense variability factors such as jitter and lack an automated process for testing circuit delays. In this article, we first propose 2-dimension multiframe fusion to deal with the sampling jitter, then present a timing delay measurement-based automatic overclocking system (AOS) running on heterogeneous FPGA for high-performance CNN accelerators. On the FPGA side, AOS is composed of timing delay monitors (TDMs) that can measure all types of timing paths, a TDM controller that converts the sampled values of TDMs into timing delay in terms of the ratio of path delay to the clock period. On the CPU side, AOS converts the path delay from clock period ratio to absolute delay value and decides the frequency of the accelerator in the next iteration. We demonstrate AOS with a SkyNet accelerator on the Xilinx ZCU104 board and achieve 657 FPS at 436 MHz without accuracy degradation, which is$1.41\times $performance compared to the baseline.
Weixiong Jiang, Heng Yu 0001, Fupeng Chen, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 HRFF: Hierarchical and Recursive Floorplanning Framework for NoC-Based Scalable Multidie FPGAs
abstract
Emerging applications are calling for significantly larger FPGAs with multi-dies. However, the interconnection architecture of existing FPGAs lacks scalability. The execution time and failure probability of their RTL-to-Silicon process increase dramatically with the growth of design and the number of dies. To address this issue, we propose both an NoC-based scalable multi-die FPGA architecture and a corresponding floorplanning framework, namely Hierarchical and Recursive Floorplanning Framework(HRFF). First, from the architecture side, we introduce an interconnection architecture with a class of scalable hierarchical topology. Second, for the algorithm side, we formulate the generic floorplanning problem for NoC-based architectures as a multi-objective Mixed Integer Linear Programming (MILP) problem, balancing the design timing and interconnection workload. Third, we develop a novel recursive approximate method to efficiently solve the multi-objective MILP formulation over the proposed architecture, with a configurable trade-off between solution quality and solver run time. Experimental results show that the scalability of our proposed technique is at least$1.5 \times $on all and$3 \times $on certain benchmarks as that of the state-of-the-art solutions with no loss of design throughput.
Jianwen Luo 0004, Fupeng Chen, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 Quality Optimization of Adaptive Applications via Deep Reinforcement Learning in Energy Harvesting Edge Devices
abstract
Applications with adaptability are widely available on the edge devices with energy harvesting capabilities. For their runtime quality optimization, however, current approaches cannot tackle the variations of quality modeling and harvested energy simultaneously. Therefore, in this article, we are the first to propose a deep reinforcement learning (DRL)-based dynamic voltage frequency scaling (DVFS) method that optimizes the application execution quality of energy harvesting edge devices to mitigate the variations. First, we propose a baseline DRL formulation that novelly migrates the objective of quality maximization into a reward function and constructs a DRL quality agent. Second, we devise a long short-term memory (LSTM)-based selector that performs DRL quality agent selection based on the energy harvesting history. Third, we further propose two optimization methods to alleviate the nonnegligible overhead of DRL computations: 1) an improved thinking-while-moving concurrent DRL scheme to compromise the “state drifting” issue during the DRL decision process and 2) a variable interstate duration decision scheme that compromises the DVFS overhead incurred in each action taken. The experiments take an adaptive stereo matching application as a case study. The results show that the proposed DRL-based DVFS method on average achieves 17.9% runtime reduction and 22.05% quality improvement compared to state-of-the-art solutions.
Fupeng Chen, Heng Yu 0001, Weixiong Jiang, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Ultra-Fast FPGA Implementation of Graph Cut Algorithm With Ripple Push and Early Termination
abstract
Graph cut has been a popular approach widely used to solve the minimum cut problem, which is prevalent in computer vision tasks, although not limited to this field. Push-relabel is considered as one of the promising algorithms of graph cut due to its good potential to be parallelized. However, existing implementations often not only fail to fully exploit the available parallelism but also fail to make full use of the application context to reduce redundant computations. Therefore, they are not competent for application scenarios with high resolution and real-time requirements. To address the issue, we propose three novel techniques to achieve an ultra-fast and efficient FPGA implementation of a push-relabel algorithm. First, we propose a ripple push technique that significantly parallelizes push operations so as to accelerate the push-relabel convergence process. Second, we propose an early-termination technique that effectively removes redundant computations of the push-relabel algorithm. We also theoretically prove the correctness of our early-termination technique. Third, we propose a highly parallelized search technique called flood irrigation search (FIS). It quickly judges early termination conditions based on a pixel parallel architecture. Our implementation focuses on performance-sensitive applications that divide large images into small graph tiles with a specific size. Compared to the state-of-the-art FPGA implementations of push-relabel algorithms, experimental results show that our method can at least achieve$8.93\times $improvement of execution time.
Guangyao Yan, Fupeng Chen, Hui Wang 0036, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 Bitwidth-Optimized Energy-Efficient FFT Design via Scaling Information Propagation
abstract
The Fast Fourier Transform (FFT) is an efficient algorithm widely used in digital signal processing to transform between the time domain and the frequency domain. For fixed-point VLSI implementations, dynamic range growth inevitably occurs at each stage of the FFT operation. However, current methods either waste bitwidth or consume excessive resources when dealing with the dynamic range growth issue. To address this issue, we propose an efficient scaling method called Scaling Information Propagation (SIP) to alleviate the problem of dynamic range growth, which makes full use of bitwidth with much less extra area consumed than the state-of-the-art solutions. In two consecutive transform operations, the SIP method extracts scaling information and makes scaling decisions in the former transform, then executes those in the latter one. We implement the FFT’s VLSI architecture in the orthogonal frequency division multiplexing (OFDM) and the holographic video compression (HVC) systems to verify the SIP method. Compared to the state-of-the-art, experimental results after VLSI synthesis show that our method achieves 9.38% energy reduction and 8.36% area savings when requiring 1.02 × 10-7bit error ratio (BER) of the OFDM system, and 33.47% energy reduction and 30.98% area savings when requiring 20dB signal-to-noise ratio (SNR) of the HVC system, respectively.
Fupeng Chen, Raees Kizhakkumkara Muhamad, David Blinder, Dessislava Nikolova, Peter Schelkens, Francky Catthoor, Yajun Ha
DAC2
2021 CLIF: Cross-Layer Information Fusion for Stereo Matching and its Hardware Implementation
abstract
The rapid advancement of intelligent systems, especially robotics and autonomous driving, is highly reliant on low-complexity and high-accuracy stereo matching algorithms. However, the performance of state-of-the-art stereo matching algorithms still has great space for improvement by gaining awareness of the implicit information hidden in the cost volume layers. In this paper, we propose a low-complexity local stereo matching algorithm named Cross-Layer Information Fusion (CLIF), to improve the matching accuracy by exploring the hidden information. First, we analyze and extract the hidden information into an auxiliary extractor using a novel fusion method. Second, we propose an information sharing strategy that transforms the extractor into a regularization term on each cost volume layer. Then we improve the design by re-constructing the information extractor between the adjacent cost volume layers and form a pipelined hardware architecture on the FPGA platform. Experimental results show that the proposed CLIF algorithm improves 6.53% average accuracy incurring negligible resources and performance impacts, compared to the state-of-the-art solutions.
Fupeng Chen, Heng Yu 0001, Yajun Ha
ISCAS1
2021 DVFS-Based Quality Maximization for Adaptive Applications With Diminishing Return
abstract
Application-level approximate computing exploits inherent resilience of adaptive applications, and trades off application output quality for runtime system resources. Existing methods treat computing quality as the number of clock cycles to execute a task, but they overlook the fact that the quality of many real-life applications exhibit the characteristic of diminishing return as the processor continues executing. The diminishing return of the quality is largely due to the features of iterative processing or successive refinement inherent in those applications. Ignoring it leads to large over-estimation in contemporary quality optimization approaches. In this article, we exploit the application adaptability to achieve quality maximization by taking both system resource constraints and diminishing return of the quality into account. We first reveal that the diminishing return of the quality is inherent in several well-known applications, and suggest an exponential model that accurately captures it. Second, we propose a dynamic frequency scaling (DFS) methodology to optimally decide the processor execution cycles for such applications, in order to maximize the output quality under system energy, timing, and temperature constraints. We transform the DFS problem to an iterative pseudo quadratic programming heuristic that can be efficiently solved. Third, we present a wrapping dynamic voltage scaling (wDVS) methodology to achieve further quality improvement, by judiciously adjusting the supply voltage to provide extra frequency scaling space. Compared to state-of-the-art algorithms, our approach produces at least 19.1 percent quality improvement on all evaluated cases, with negligible execution overhead.
Heng Yu 0001, Yajun Ha, Bharadwaj Veeravalli, Fupeng Chen, Hesham El-Sayed
IEEE Trans. Computers4
2020 Quality Estimation and Optimization of Adaptive Stereo Matching Algorithms for Smart Vehicles
abstract
Stereo matching is a promising approach for smart vehicles to find the depth of nearby objects. Transforming a traditional stereo matching algorithm to its adaptive version has potential advantages to achieve the maximum quality (depth accuracy) in a best-effort manner. However, it is very challenging to support this adaptive feature, since (1) the internal mechanism of adaptive stereo matching (ASM) has to be accurately modeled, and (2) scheduling ASM tasks on multiprocessors to generate the maximum quality is difficult under strict real-time constraints of smart vehicles. In this article, we propose a framework for constructing an ASM application and optimizing its output quality on smart vehicles. First, we empirically convert stereo matching into ASM by exploiting its inherent characteristics of disparity–cycle correspondence and introduce an exponential quality model that accurately represents the quality–cycle relationship. Second, with the explicit quality model, we propose an efficient quadratic programming-based dynamic voltage/frequency scaling (DVFS) algorithm to decide the optimal operating strategy, which maximizes the output quality under timing, energy, and temperature constraints. Third, we propose two novel methods to efficiently estimate the parameters of the quality model, namely location similarity-based feature point thresholding and street scenario-confined CNN prediction. Results show that our DVFS algorithm achieves at least 1.61 times quality improvement compared to the state-of-the-art techniques, and average parameter estimation for the quality model achieves 96.35% accuracy on the straight road.
Fupeng Chen, Heng Yu 0001, Yajun Ha
ACM Trans. Embed. Comput. Syst.1