EDBT 2026 Demo / reviewers in the wild / expert
Chenyang Dai
dblp:336/8362
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0001-9609-0343ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FPGA-Par: An Efficient Algorithm for Elegant Partitioning in Multi-FPGA SystemsabstractThis work introduces FPGA-Par, an efficient graph partitioning algorithm for multi-FPGA systems. FPGA-Par utilizes an iterative balanced partitioning and supernodes transferring algorithm that transforms imbalanced partitioning problems into balanced ones, addressing the inefficiencies associated with traditional single-node movement methods, as well as eliminating the need for brute-force exploration of imbalance factors. This approach achieves partitions with flexible size ratios that satisfy resource constraints, resulting in improved partition quality. Experimental results demonstrate that FPGA-Par reduces time overhead by 55% compared to state-of-the-art imbalanced partitioning algorithms. Furthermore, it improves partition quality metrics, with Cut and Cutmaxdecreasing by 38% and 25%, respectively, and achieves a 1.25x increase in system frequency after mapping and routing on a multi-FPGA system. Hengyue Gao, Chenyang Dai, Jinlun Ji, Qiyue Zhao, Jiangtao Yuan, Li Li 0003 |
ISCAS | 2 |
| 2025 | An Energy Efficient Residual Spiking Neural Network Accelerator With Ternary SpikesabstractSpiking neural networks (SNNs) use discrete binary spikes to transfer information between neurons, which is different from artificial neural networks (ANNs). Although event-based characteristics bring potential computation power and efficiency to SNNs, the long processing time window of discrete spikes leads to high latency. In this brief, a spike version of the residual network using ternary spikes is proposed. A shorter time window is required to achieve competitive performance because the ability to transfer information of the ternary spikes is strengthened. An SNN accelerator based on the proposed residual network with ternary spikes is designed and implemented with 28 nm CMOS technology, and the core area is 0.63 mm2. The proposed SNN accelerator achieves the classification accuracy of 92.07% on CIFAR-10 dataset with SResNet20 and only 6 time steps. The accelerator achieves 0.39 mJ energy consumption per frame with a throughput of 165.7 FPS when running at 500 MHz. Congyi Sun, Wenqing Song, Qinyu Chen, Chenyang Dai, Li Li 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | FAS-NoC: A Real-Time Fused Approximation Scheme Coordinating Communication and Computation for NoC-Based NN AcceleratorsabstractNetwork-on-Chip (NoC) is a scalable on-chip communication architecture widely used in neural network accelerators. However, data-intensive applications like machine learning place significant demands on the NoC’s communication and computation, and often have a degree of resilience to data noise, which allows to use approximation techniques to reduce execution time and energy consumption for both computation and communication, under the constraints of acceptable quality loss. Traditional approximate NoCs do not consider the data distribution characteristics of the neural networks, resulting in a lower approximate rate. Moreover, these schemes do not take the synergistic optimization of computation and communication, which limits reductions in execution time. In this paper, we propose a Fused Approximation Scheme of NoC (FAS-NoC) that incorporates the characteristics of data distribution in neural networks. FAS-NoC includes an approximate compression and recovery scheme based on data hierarchy, and uses a congestion-aware scheme to adjust the approximate rate of the node. Additionally, by leveraging the characteristics of recovered data after approximate communication, the scheme optimizes the design of computing units within the computing array. FAS-NoC collaboratively optimizes approximate communication and computation, organically integrating the two aspects. Compared with the state-of-the-art approximate framework ACDC (ACDC_ABDTR and ACDC_APPROX), the execution time of FAS-NoC is reduced by$48.94\%$and$47.72\%$, respectively. The experimental results show that the additional area overhead of the FAS-NoC only accounts for$0.81\%$of the original node, the additional power consumption overhead only accounts for$0.75\%$of the original node. Chuanzhu Liu, Wenjie Fan 0004, Heng Zhang 0025, Chenyang Dai, Congyi Sun, Xinyu Wang 0027, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | HAS-RL: A Hierarchical Approximate Scheme Optimized With Reinforcement Learning for NoC-Based NN AcceleratorsabstractNetwork-on-Chip (NoC) is a scalable on-chip communication architecture for the NN accelerator, but with the increase in the number of nodes, the communication delay becomes higher. Applications such as machine learning have a certain resilience to noisy/erroneous transmitted data. Therefore, approximate communication becomes a promising solution to improving performance by reducing traffic loads under the constraint of the acceptable maximum accuracy loss of neural networks. It is a key issue to balance the result quality and the communication delay for approximate NoC systems. The traditional approximate NoC only considers the node-to-node approximation-based dynamic traffic regulation. However, the dynamically changing traffic patterns across different nodes, different times, and different applications lead to a huge search space, which makes it hard to explore an optimal global approximation solution. In this paper, we propose a quality model for different neural networks, which presents the relationship between the quality loss and the data approximate rate. Then, a hierarchical approximate scheme optimized with reinforcement learning (HAS-RL) is proposed and we reduce the complexity of the HAS-RL by reducing the state space and action space, which will reduce the resource overhead as well. After that, we embed a global approximate controller in the NoC system, in which we deploy a policy network trained with the offline reinforcement learning algorithm to adjust the data approximate rates of each node at run time. Compared with the state-of-the-art method, the proposed scheme reduces the average network delay by 13.5% while their accuracies are similar. The proposed HAS-RL only causes an additional area overhead of 1.24% and power consumption of 0.77% compared with the traditional router design. Shize Zhou, Yongqi Xue, Wenjie Fan 0004, Tong Cheng, Jinlun Ji, Chenyang Dai, Wenqing Song, Qinyu Chen, Chang Gao 0002, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |