Chao Chen 0042

dblp:66/3019-42 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0001-6654-8341ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DPAM: A Dual-Path Protected Approximate Multiplier for Reliable Neural Network Acceleration
Chao Chen 0042, Yanxi Lin, Haoan Yin, Yan Li 0084, Xiaoyang Zeng
ISCAS1
2026 Toward Exploring Fault-Tolerant Neural Architectures: A Hierarchical Codesign Optimization Framework
abstract
The increasing deployment of neural networks in safety-critical domains, such as autonomous driving and embodied artificial intelligence, has underscored the urgent need for fault-tolerant neural architectures. Hardware-induced faults stemming from soft errors, aging, or other disturbances can severely impair model performance. In this paper, we propose a hierarchical optimization framework that systematically designs fault-tolerant neural architectures from operator design to architecture search method, while minimizing both accuracy loss and computational cost. Specifically, we design a fully decoupled Winograd convolution operator (FD-WGC) that localizes the impact of bit-flip faults and reduces computational cost. We then expand the architecture search space by introducing fault-tolerant cells composed of the FD-WGC and complementary operators, enabling more flexible network construction. Within this expanded search space, we employ MOBO-NAS, a multi-objective Bayesian optimization based neural architecture search method, to efficiently explore neural architectures that balance accuracy, computational cost, and fault tolerance. Experimental results show that our framework enhances fault tolerance by up to 510× compared to state-of-the-art (SOTA) manually designed and automatically searched architectures, while maintaining comparable accuracy and reducing computational cost by up to 80%. Extensive evaluations across diverse hardware fault models further validate the generalizability and effectiveness of our proposed framework. All codes are available at https://github.com/cc-innocence/MOBO-NAS/tree/master.
Chao Chen 0042, Liang Wang 0024, Yan Li 0084, Xiaoyang Zeng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2026 Energy-Efficient Logarithmic Floating-Point Multipliers Using Truncation-Based Error Compensation for Fault-Tolerant Applications
abstract
The growing demand for energy-efficient computing in resource-constrained devices necessitates approximate arithmetic solutions that balance accuracy and hardware cost. This article presents a family of logarithmic truncation-based approximate multipliers (LTAMs) for floating-point (FP) operations, including two hardware cost-optimized architectures (LTAM1and LTAM2) and an accuracy-focused lookup table (LUT)-based design (LTAM-LUT). All designs effectively address the systematic negative bias inherent in classical logarithmic multipliers through advanced error compensation methods. Based on a comprehensive analysis, a 6-bit mantissa truncation is identified as the optimal configuration. Under the 6-bit configuration, LTAM1-6achieves 52.9% area and 77.6% power-delay product (PDP) reduction compared to conventional logarithmic approximate multipliers, while LTAM2-6 provides 45.1% area and 70.4% PDP reduction. LTAM-LUT achieves 56.4% accuracy improvement over conventional logarithmic approximate multipliers with a mean relative error distance (MRED) of 1.68%, the lowest among all compared designs. These hardware efficiency gains are also validated across diverse application domains. In HDR tone mapping, LTAM-LUT achieves up to 7.0dB higher peak signal-to-noise ratio (PSNR) than conventional designs, while LTAM1-6and LTAM2-6 provide 1.1- and 4.6-dB improvements, respectively. For single-image super-resolution (SISR) on DIV2K ($\times 4$scale), LTAM-LUT preserves reconstruction quality nearly identical to exact arithmetic on HAT (29.73dB) and achieves up to 2.3dB higher PSNR than conventional logarithmic multipliers on EDSR, while LTAM1-6and LTAM2-6 consistently outperform prior approximate designs across all tested architectures, demonstrating superior accuracy–efficiency tradeoffs for error-tolerant applications.
Baining Wu, Chao Chen 0042, Yan Li 0084, Xiaoyang Zeng
IEEE Trans. Very Large Scale Integr. Syst.4
2025 MAD-Flow: An Efficient Deployment Flow for Memory-Bank-Based Anomaly Detection on FPGA-SoCs
Weiju Wu, Chao Chen 0042
IEEE Internet Things J.2
2025 Acceleration of Fast Sample Entropy for FPGAs
abstract
Complexity measurement, essential in diverse fields like finance, biomedicine, climate science, and network traffic, demands real-time computation to mitigate risks and losses. Sample Entropy (SampEn) is an efficacious metric which quantifies the complexity by assessing the similarities among microscale patterns within the time-series data. Unfortunately, the conventional implementation of SampEn is computationally demanding, posing challenges for its application in real-time analysis, particularly for long time series. Field Programmable Gate Arrays (FPGAs) offer a promising solution due to their fast processing and energy efficiency, which can be customized to perform specific signal processing tasks directly in hardware. The presented work focuses on accelerating SampEn analysis on FPGAs for efficient time-series complexity analysis. A refined, fast, Lightweight SampEn architecture (LW SampEn) on FPGA, which is optimized to use sorted sequences to reduce computational complexity, is accelerated for FPGAs. Various sorting algorithms on FPGAs are assessed, and novel dynamic loop strategies and micro-architectures are proposed to tackle SampEn's undetermined search boundaries. Multi-source biomedical signals are used to profile the above design and select a proper architecture, underscoring the importance of customizing FPGA design for specific applications. Our optimized architecture achieves a 7x to 560x speedup over standard baseline architecture, enabling real-time processing of time-sensitive data.
Chao Chen 0042, Chengyu Liu 0001, Jianqing Li 0002, Bruno da Silva 0001
IEEE Trans. Computers1
2025 FASE: An FPGA-Based Accelerator for Lightweight Sample Entropy With Monte Carlo Sampling
abstract
Sample entropy (SampEn) is an algorithm within information entropy that enables effective analysis of biological signals. Due to the need for extensive similarity matching operations, the SampEn calculation process is time-consuming. Although a series of fast SampEn algorithms have been proposed, they remain time-intensive when processing large data volumes. Additionally, previous field-programmable gate array (FPGA)-based hardware accelerators designed for SampEn suffer from architectural design limitations, consuming substantial on-chip memory resources and operating at low frequencies. In this article, we propose FASE, an FPGA-based accelerator for lightweight sample entropy (LW-SampEn) with Monte Carlo (MC) sampling. The FASE design comprises two main parts: algorithm and hardware optimizations. On the algorithmic side, we introduce MC sampling into the merge-sort-based LW-SampEn algorithm, named MCLW-SampEn. MCLW-SampEn effectively reduces the computation load for large data volumes while maintaining algorithmic accuracy. For hardware, we first design efficient sorting and allocation modules to address boundary localization and load imbalance issues in previous accelerator designs. Then, we replicate the computation across the main phases to enable parallel processing. Finally, we deploy the design on the Pynq-Z2 board for validation. Experimental results show that the proposed MCLW-SampEn algorithm achieves an average speed up of$3\times $over the LW-SampEn algorithm, with accuracy losses kept within 0.5%. Compared to state-of-the-art (SOTA) designs, FASE achieves an average speed up of$12.8\times $while reducing power consumption by 89.3%. Ablation studies indicate that, for the same algorithm, FASE offers a$7.4\times $speedup over related FPGA designs.
Yuanhang Li, Zhengyang Huang, Chao Chen 0042, Ruiqi Chen 0001, Bruno da Silva 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2023 DMBF: Design Metrics Balancing Framework for Soft-Error-Tolerant Digital Circuits Through Bayesian Optimization
abstract
Radiation Hardened by Design (RHBD) is one of the main measures for solving the soft error issue in digital circuits. However, a multi-objective optimization (MOO) problem obviously appears when utilizing the hardened counterparts to replace the original unreliable cells. This paper proposes a MOO framework based on Bayesian Optimization (BO) for balancing design metrics like area, Longest Path Delay (LPD)/power, and Soft Error Rate (SER) while hardening digital circuits, including combinational and sequential circuits. This framework comprises two phases: 1) data preprocessing and 2) multi-objective Bayesian optimization. The first phase makes this framework much more applicable for large-scale circuits through data dimensionality reduction. The second phase is characterized by utilizing a black-box approach to greatly promote the efficiency and accuracy of MOO. Experimental results on benchmark circuits demonstrate that the framework achieves a 1.34x improvement in accuracy, an 11.47x enhancement in efficiency, and a 0.77x reduction in SER, while exhibiting a 4.27x and 0.72x increase in area for combinational and sequential benchmark circuits, respectively, along with a 0.54x increase in LPD and a 1.25x increase in power for Triple Modular Redundancy (TMR) techniques.
Yan Li 0084, Chao Chen 0042, Xu Cheng 0002, Jun Han 0003, Xiaoyang Zeng
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 Biological Activity Prediction of GPCR-targeting Ligands on Heterogeneous FPGA-based Accelerators
abstract
In the drug discovery process, the biological activity value (BAV) of G Protein-Coupled Receptors (GPCRs) targeting ligands is a large consideration. Past BAV prediction on CPU consumes tremendous time and power, yet there is rarely any related acceleration research. Therefore, this paper proposes a series of heterogeneous FPGA-based accelerators for well-performing algorithms to predict GPCRs ligands BAV. Communication delay is reduced by compressing the sparse matrix and directly coupling accelerators on the system BUS. Computation is accelerated by the remapping during the weight storage. Experimental results show that our FPGA accelerator implemented on Xilinx XCZU7EV performs 54.5× faster than CPU and 35.2× more energy-efficient than GPU.
Ruiqi Chen 0001, Yuhanxiao Ma, Shaodong Zheng, Shizhen Huang, Chao Chen 0042, Jun Yu 0010, Kun Wang 0005
FCCM5
2022 Acceleration of Fast Sample Entropy Towards Biomedical Applications on FPGAs
abstract
Sample Entropy (SampEn) is an information en-tropy algorithm widely used for complexity analysis and chaos estimation in many applications. In particular, SampEn measures complexity of time series by the conditional probability of the inner pattern. Unfortunately, the straightforward implementation of SampEn is quadratic time complexity, restricting its real-time analysis ability for health applications and long-term data analysis. Although researchers have proposed fast versions of SampEn to avoid unnecessary comparisons, they have not been accelerated yet due to their performance bottleneck in the complex similarity pair process. In this paper, we evaluate fast SampEn algorithms by employing multi-source biomedical signals on an Field-Programmable Gate Arrays (FPGA). Since fast SampEn algorithms based of a pre-sorting stage promise to outperform other SampEn algorithms, Lightweight SampEn based on Merge Sort is here implemented and optimized. Dif-ferent type of optimizations, that can be generalized for similar Lightweight-based SampEn algorithms, are used to reduce the overall latency while the data throughput is increased. A load balancing strategy for multi similarity pair modules is also proposed to solve the unbalancing loads, a bottleneck when increasing the execution parallelism of this type of algorithms. As a result, the proposed SampEn architecture runs 10 times faster than the fastest SampEn implementation on a modern CPU.
Chao Chen 0042, Bruno da Silva 0001, Jianqing Li 0002, Chengyu Liu 0001
FPT1