VLDB 2026 Research / reviewers in the wild / expert
Chao Chen 0022
dblp:66/3019-22
· DBLP profile ↗
13ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0001-6488-224XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 2 first-author · 10 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NMSAcc: An Arch-Decoupled Non-Maximum Suppression Accelerator on Edge Devices
Zheng Wang 0027, Qiawei Zheng, Jinghan Zhou, Chao Chen 0022 |
ISCAS | 5 |
| 2026 | Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoCabstractWhile recent advances in AI SoC design have focused heavily on accelerating tensor computation, the equally critical task of tensor manipulation (TM)—centered on high-volume data movement with minimal computation—remains underexplored. This work addresses that gap by introducing the TM unit (TMU): a reconfigurable, near-memory hardware block designed to execute data-movement-intensive (DMI) operators efficiently. The TMU manipulates long datastreams in a memory-to-memory fashion using a RISC-inspired execution model and a unified addressing abstraction, enabling support for both a wide range of coarse- and fine-grained tensor transformations. The proposed architecture integrates the TMU alongside a TPU within a high-throughput AI system-on-chip (SoC), leveraging double buffering and output forwarding to improve pipeline utilization. The TMU, synthesized under the SMIC 40-nm standard cell library, occupies only$0.019~\mathrm {\text {mm}^{2}}$while supporting over 10 representative DMI operators. Benchmarking shows that the TMU alone achieves up to$82.42\times $and$11.06\times $operator-level latency reduction over ARM A72 and NVIDIA Jetson TX2, respectively. When integrated with the in-house TPU, the complete system achieves a 22.89% reduction in end-to-end inference latency, demonstrating the effectiveness in reducing inference latency and the scalability of the TMU architecture across diverse tensor operators. Weiyu Zhou, Zheng Wang 0027, Chao Chen 0022, Yongkui Yang, Zhuoyu Wu, Anupam Chattopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | AttenPU: An Area Efficient Attention Processor with Reconfigurable FP8 Precision and DataflowabstractEfficient numerical representation is crucial for deep learning accelerators, especially for large language models (LLMs). The 8-bit-floating-point (FP8) data representation achieves higher precision and fewer quantization efforts than integer, which has been proven inevitable in attention-based accelerators for LLMs. Therefore, area-efficient design techniques for FP8 play a central role in lowering LLMs chip’s budget. This paper presents AttenPU, which is built upon reconfigurable FP8 units and supports E4M3 for inference and E5M2 for training. Bidirectional dataflow is exploited to enable AttenPU to interact with FP32 coprocessor to reduce latency. The design achieves a low FP8-to-INT8 area ratio of 1.63×, an area efficiency of 193.5 GFLOPS/mm2, with an 87.05% reduction in the latency of RTX3090 GPU. Qiawei Zheng, Zheng Wang 0027, Zhuoyu Wu, Zhihao Du, Chao Chen 0022, Yongkui Yang, Wenqi Fang, Anupam Chattopadhyay |
ACM Great Lakes Symposium on VLSI | 7 |
| 2025 | Swift: Fast Performance Tuning with GAN-Generated Configurations
Chao Chen 0022, Shixin Huang, Xuehai Qian, Zhibin Yu 0001 |
USENIX ATC | 1 |
| 2024 | PerFT-N: Low-overhead Permanent Fault-Tolerance Mechanism for Neural Processing UnitsabstractThe reliability of neural network (NN) accelerators is one key to ensuring inference accuracy. Relying solely on the robustness of NN can only tolerate transient faults, and once the circuit encounters permanent faults, it will lead to a serious accuracy drop. We, by employing processing elements (PEs) checking and re-scheduling of threads of functional units, propose a new fault-tolerant mechanism named PerFT-N for neural processing units (NPUs). Compared to previous NPU fault-tolerant methods, PerFT-N recovers the execution facing permanent faults with minimal hardware overhead. Specifically, we utilize the existing resources to achieve fault detection and localization, while achieving recovery by re-linking unfailed threads. Furthermore, an instruction-based programmable permanent fault emulation scheme is deployed on the FPGA platform for fast verification. Experimental results demonstrate that the PerFT-N architecture works in the extreme case with 98.5% of failed functional units while incurring the physical overheads of 2.7% in area and 3.6% in power consumption under 40nm SMIC standard cell library. Haojie Jian, Chao Chen 0022, Zheng Wang 0027 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2024 | Low-latency Buffering for Mixed-precision Neural Network Accelerator with MulTAP and FQPipeabstractPrevious work has proposed precision scalable accelerators to handle mixed-precision neural network (NN) inferences on the edge, which focus on designing reconfigurable MAC arrays while leaving the issue of time-costly data buffering procedure less discussed. Besides, integer-only inference is incapable of handling emerging NN models with various non-linear activation functions. In this work, we propose a mixed-precision NN accelerator supporting int8, int16 and fp32 arithmetic with two buffering techniques namely MulTAP and FQPipe, which jointly facilitate low-latency data movement. Experiment results show that MulTAP and FQPipe boost the baseline NN accelerator with 7.7 × and 1.5 × in speed respectively, which leads to the application performance of 473.9 (int8) and 252.5 (int16) inferences per second (IPS) on YOLOv3-Tiny. Post-layout netlist with SMIC 40nm standard-cell technology demonstrates a design with an area of 26.96mm2and a power estimate of 1.83W. Zheng Wang 0027, Wenhui Ou, Weiyu Zhou, Yongkui Yang, Chao Chen 0022 |
ISCAS | 7 |
| 2024 | TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data AnalyticsabstractRecently, experiment-driven machine-learning (ML) based configuration tuning for in-memory data analytics such as Apache Spark become popular because they can achieve high speedups. However, experiment-driven ML-based approaches naturally need alargenumber of iterations and each iteration generates a configuration with a probabilistic strategy and executes the program on a real cluster with the configuration. It therefore takes a long time to optimize the performance of an in-memory data analytics program, and thereby hinders these approaches from being widely used in practice.To address this issue, we propose a novel as well as simple approach dubbedTerminating-It-Early (TIE)to reduce the time needed to perform the experiment executions but to achieve speedups similar to those obtained by experiment-driven ML-based approaches. The key idea is that, during the process of searching for the optimal configuration which produces the shortest execution time for a program, weterminatean experiment program execution with a trial configuration as soon as possible when we find its execution time islonger than a predefined threshold(e.g., the shortest execution time thus far). In contrast, traditional experiment-driven ML-based approaches always run all experiment executions completely.We employ 19 Apache Spark programs running on a physical cluster as well as a virtual cluster to evaluate TIE. We compare thetuning timeused to find the optimal configuration of a program and theoptimized execution timeof a program obtained by TIE against those obtained byCherryPickand a reinforcement learning (RL) based approach. The experimental results show that on physical machines, TIE reduces the tuning time used byCherryPickand the RL-based approach by factors of 2.39× and 1.68× on average, respectively. On virtual machines, the corresponding factors are 2.79× and 1.71×. Moreover, the average optimized execution time of the 19 programs tuned by TIE is slightly shorter than those tuned byCherryPickand the RL-based approach. Chao Chen 0022, Jinhan Xin, Zhibin Yu 0001 |
IEEE Trans. Computers | 1 |
| 2023 | COMPACT: Co-processor for Multi-mode Precision-adjustable Non-linear Activation FunctionsabstractNon-linear activation functions imitating neuron behaviors are ubiquitous in machine learning algorithms for time series signals while also demonstrating significant gain in precision for conventional vision-based deep learning networks. State-of-the-art implementation of such functions on GPU-like devices incurs a large physical cost, whereas edge devices adopt either linear interpolation or simplified linear functions leading to degraded precision. In this work, we design COMPACT, a co-processor with adjustable precision for multiple non-linear activation functions including but not limited to exponent, sigmoid, tangent, logarithm, and mish. Benchmarking with state-of-the-arts, COMPACT achieves a 26% reduction in the absolute error on a 1.6x widen approximation range taking advantage of the triple decomposition technique inspired by Hajduk's formula of Padé approximation. A SIMD-ISA-based vector co-processor has been implemented on FPGA which leads to a 30% reduction in execution latency but the area overhead nearly remains the same with related designs. Furthermore, COMPACT is adjustable to 46% latency improvement when the maximum absolute error is tolerant to the order of 1E-3. Wenhui Ou, Zhuoyu Wu, Zheng Wang 0027, Chao Chen 0022, Yongkui Yang |
DATE | 4 |
| 2021 | OR-ML: Enhancing Reliability for Machine Learning Accelerator with Opportunistic RedundancyabstractReliability plays a central role in deep sub-micron and nanometre IC fabrication technology and has recently been reported to be one of the key issues affecting the inference phase of neural networks. State-of-the-art machine learning (ML) accelerators exploit massively computing parallelism observed in neural networks to achieve high energy efficiency. The topology of ML engines' computing fabric, which constitutes large arrays of processing elements (PEs), has been increasing dramatically to incorporate the huge size and heterogeneity of the rapid evolving ML algorithm. However, it is commonly observed that activations of zero value lead to reduced PE utilization. In this work, we present a novel and low-cost approach to enhance the reliability of generic ML accelerators by Qpportunistically exploring the chances of runtime Redundancy provided by neighbouring PEs, named as OR-ML. In contrast to conventional redundancy techniques, the proposed technique introduces no additional computing resources, therefore significantly reduces the implementation overhead and achieves obvious level of protection. The design prototype is evaluated using emulated fault injection on FPGA, executing mainstream neural networks for objectionclassification and detection. Zheng Wang 0027, Wenxuan Chen, Chao Chen 0022, Yongkui Yang, Zhibin Yu 0001 |
DATE | 4 |
| 2021 | CNN-DMA: A Predictable and Scalable Direct Memory Access Engine for Convolutional Neural Network with Sliding-window FilteringabstractMemory bandwidth utilization has become the key performance bottleneck for state-of-the-art variants of neural network kernels. Current structures such as depth-wise, point-wise and atrous convolutions have already introduced diverse and discontinuous memory access patterns, which impact efficient activation supply due to more frequent cache misses and consequently high-penalty DRAM pre-charging. To handle this, GPU achieves efficient parallelization with sophisticated optimization of CUDA program to reduce memory footprints, which demands high engineering efforts. In this work, we in contrast propose a programmable direct memory access engine for convolutional neural networks (CNN-DMA) supporting a fast supply of activation for independent and scalable computing units. The CNN-DMA favours a predictable activation streaming approach which completely avoids penalties by bus contention, cache misses and less carefully designed low-level programs. Furthermore, we enhance the baseline DMA with the capability of out-of-order data supply to filter out unique sliding-windows to boost the performance of the computing infrastructure. Experiments on state-of-the-art neural networks show that CNN-DMA achieves optimal DRAM access efficiency for point-wise convolution layers, while reduces 30% to 70% rounds of computation with sliding-window filtering. Zheng Wang 0027, Chao Chen 0022, Yongkui Yang, Weiguang Chen, Wenxuan Chen, Weiyu Guo, Zhibin Yu 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2020 | BBS: Micro-Architecture Benchmarking Blockchain Systems through Machine Learning and Fuzzy SetabstractDue to the decentralization, irreversibility, and traceability, blockchain has attracted significant attention and has been deployed in many critical industries such as banking and logistics. However, the micro-architecture characteristics of blockchain programs still remain unclear. What's worse, the large number of micro-architecture events make understanding the characteristics extremely difficult. We even lack a systematic approach to identify the important events to focus on. In this paper, we propose a novel benchmarking methodology dubbed BBS to characterize blockchain programs at micro-architecture level. The key is to leverage fuzzy set theory to identify important micro-architecture events after the significance of them is quantified by a machine learning based approach. The important events for single programs are employed to characterize the programs while the common important events for multiple programs form an importance vector which is used to measure the similarity between benchmarks. We leverage BBS to characterize seven and six benchmarks from Blockbench and Caliper, respectively. The results show that BBS can reveal interesting findings. Moreover, by leveraging the importance characterization results, we improve that the transaction throughput of Smallbank from Fabric by 70% while reduce the transaction latency by 55%. In addition, we find that three of seven and two of six benchmarks from Blockbench and Caliper are redundant, respectively. Chao Chen 0022, Zihao Su, Weiguang Chen, Tao Li 0006, Zhibin Yu 0001 |
HPCA | 2 |
| 2014 | System-level reliability exploration framework for heterogeneous MPSoCabstractPower density of digital circuits increased at alarming rate for deep sub-micron CMOS technology, turning reliability into a serious design concern. On the other hand, ever-growing task complexity with strict performance budget forced designers to adopt complex, heterogeneous MPSoCs as the implementation choice. Several commercial system-level design platforms exist currently for design, exploration and implementation of MPSoC. In this paper, we propose a system-level reliability exploration framework by extending a commercial system-level design flow. Using this framework, a heterogeneous MPSoC is designed which can accept a custom mapping algorithm based on the MPSoC topology before the actual task deployment. The dynamic reliability-aware task management is able to consider the desired reliability constraints of tasks as well as reliability levels of the system components. We report our experimental findings using state-of-the-art benchmark applications. Zheng Wang 0020, Chao Chen 0022, Piyush Sharma, Anupam Chattopadhyay |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | Accurate and efficient reliability estimation techniques during ADL-driven embedded processor designabstractThe downscaling of technology features has brought the system developers an important design criteria, reliability, into prime consideration. Due to external radiation effects and temperature gradients, the CMOS device is not guaranteed anymore to function flawlessly. On the other hand, admission for errors to occur allows extending the power budget. The power-performance-reliability trade-off compounds the system design challenge, for which efficient design exploration framework is needed. In this work, we present a high-level processor design framework extended with two reliability estimation techniques. First, a simulation-based technique, which allows a generic instruction-set simulator to estimate reliability via high-level fault injection capability. Second, a novel analytical technique, which is based on the reliability model for coarse arithmetic logical operator blocks within a processor instruction. The techniques are tested with a RISC processor and several embedded application kernels. Our results show the efficiency and accuracy of these techniques against a HDL-level reliability estimation framework. Zheng Wang 0020, Kapil Singh, Chao Chen 0022, Anupam Chattopadhyay |
DATE | 3 |