EDBT 2026 Demo / reviewers in the wild / expert
Kai Chen 0034
dblp:181/2839-34
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0007-9970-804XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast Modular Reduction Algorithm and Reconfigurable Domain-Specific Architecture Design Based on Generalized Mersenne PrimesabstractModular arithmetic enjoys a broad spectrum of applications. In recent years, the advancement of post-quantum cryptography (PQC) has imposed growing demands on the flexibility and scalability of domain-specific accelerators. This article presents a fast modular reduction algorithm based on generalized Mersenne (GM) primes, which employs approximate scaling and iterative compression approaches to rapidly converge the quotient value while reducing the precomputation complexity to rely solely on the modulus itself. Building upon this algorithm, we have designed a reconfigurable modular reduction array using multiplier units with smaller word length. Operating at 1GHz, the proposed array achieves an area reduction of 45.14% compared to Barrett-based structures, and 13.93% compared to Montgomery-based structures. The array has been integrated into a complete number-theoretic transform (NTT) acceleration architecture. The resulting reconfigurable GM/general modular reduction domain-specific architecture elevates the chip frequency to$0.42\sim 1$GHz under the same process technology. Under identical test conditions, it improves area efficiency by 10.6% and reduces energy consumption by 17.0% compared to the state-of-the-art ASIC design. When compared to the latest field-programmable gate array (FPGA) implementations, it achieves a reduction in area-time product (ATP) by 7.2%~18.4% for Kyber and by 13.4% for Dilithium. These results strongly demonstrate the notable advantages of the hardware-friendly GM algorithm. Xinyu Wang 0027, Guoqiang He, Congyi Sun, Kai Chen 0034, Li Li 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | Automatic Generation and Optimization Framework of NoC-Based Neural Network Accelerator Through Reinforcement LearningabstractChoices of dataflows, which are known as intra-core neural network (NN) computation loop nest scheduling and inter-core hardware mapping strategies, play a critical role in the performance and energy efficiency of NoC-based neural network accelerators. Confronted with an enormous dataflow exploration space, this paper proposes an automatic framework for generating and optimizing the full-layer-mappings based on two reinforcement learning algorithms including A2C and PPO. Combining soft and hard constraints, this work transforms the mapping configuration into a sequential decision problem and aims to explore the performance and energy efficient hardware mapping for NoC systems. We evaluate the performance of the proposed framework on 10 experimental neural networks. The results show that compared with the direct-X mapping, the direct-Y mapping, GA-base mapping, and NN-aware mapping, our optimization framework reduces the average execution time of 10 experimental NNs by 9.09$\%$, improves the throughput by 11.27$\%$, reduces the energy by 12.62$\%$, and reduces the time-energy-product (TEP) by 14.49$\%$. The results also show that the performance enhancement is related to the coefficient of variation of the neural network to be computed. Yongqi Xue, Jinlun Ji, Xinming Yu, Shize Zhou, Tong Cheng, Shiping Li, Kai Chen 0034, Zhonghai Lu, Li Li 0003 |
IEEE Trans. Computers | 9 |
| 2023 | A DSP-Purposed REconfigurable Acceleration Machine (DREAM) for High Energy Efficiency MIMO Signal ProcessingabstractThe wireless baseband processing algorithms are still developing and show a great diversity. The development of ASIC implementations cannot quickly adapt to the evolution of algorithms and standards. Meanwhile, the general-purpose processors cannot meet the real-time requirements in some scenarios. This paper proposes a DSP-purposed REconfigurable Acceleration Machine (DREAM) core for wireless baseband digital signal processing, which has a good trade-off between flexibility and performance. First, we abstract a set of shared operators with a moderate granularity from a variety of wireless MIMO signal processing algorithms. Then, we propose a two-step configuration process to reduce the size of the required reconfiguration bits. Besides, we design a conflict-free address generator to transfer data between the on-chip scratchpad memory and reconfiguration processing elements with high efficiency and high throughput. Finally, the prototype DREAM core has been implemented in TSMC CMOS 28 nm, and its area and power consumption have been analyzed. The chip has great flexibility in supporting a variety of wireless MIMO processing algorithms and a wide range of MIMO scales. The proposed DREAM core can achieve the normalized area efficiency and the normalized energy efficiency of$0.67~Gbps/MGE$and$15.05~Gbps/W$, which are$1.56\times $and$4.18\times $those of state-of-the-art reconfigurable implementations when running the WeJi-based MIMO detection algorithm. Kai Chen 0034, Wenqing Song, Guoqiang He, Sirui Shen, Huizheng Wang, Chuan Zhang 0001, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Unsupervised Learning Based on Temporal Coding Using STDP in Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) have been recognized as one of the next generation of Neural Networks (NNs), showing a great potential in a variety of applications. Spiking-Timing Dependent Plasticity (STDP) underlies the brain’s learning mechanisms, and trains SNNs with great energy efficiency. In this paper, we propose a low-cost spike-time based unsupervised learning method. It constructs a SNN with one fully-connected excitatory layer structure without inhibitory layer, and trains the SNN with STDP using a first-spike-based temporal coding scheme where input information is directly encoded into spike times. It only updates the synaptic weights connected to the neuron that first generates a spike in a forward propagation step, which reduces the frequency of the synaptic weight updates significantly. The forward propagation process can be stopped once a neuron fires whether in the training mode or the inference mode, by which many unnecessary computations are just avoided and the latency in the inference mode is reduced. The method was used to train on the classification task on MNIST dataset and achieved an accuracy of 90.4% with 800 excitatory neurons. Congyi Sun, Qinyu Chen, Kai Chen 0034, Guoqiang He, Li Li 0003 |
ISCAS | 3 |
| 2019 | Congestion-Aware Dynamic Elevator Assignment for Partially Connected 3D-NoCsabstractThe combination of Network-on-Chips (NoCs) and 3D IC technology, 3D NoCs, has been proven to be able to achieve a great improvement in both network performance and power consumption compared to 2D NoCs. In the traditional 3D NoC, all routers are vertically connected. Due to the large overhead of Through-Silicon-Via (TSV, e.g., low fabrication yield and the occupied silicon area), the partially connected 3D NoC has emerged. The assignment method determines the traffic loads of the vertical links (elevators), thus has a great impact on 3D-NoCs' performance. In this paper, we propose a congestion-aware dynamic elevator assignment (CDA) scheme, which takes both the distance factors and network congestion information into account. Experiments show that the performance of the proposed CDA scheme is improved by 67% to 87% compared to the random selection scheme, 8% to 25% compared to SelByDis-1, and 13% to 18% compared to SelByDis-2. Qinyu Chen, Guoqiang He, Kai Chen 0034, Zhonghai Lu, Chuan Zhang 0001, Li Li 0003 |
ISCAS | 4 |