EDBT 2026 Demo / reviewers in the wild / expert
Dongsheng Liu 0001
dblp:07/6337-1
· DBLP profile ↗
11ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0001-5571-1932ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Lightweight and Efficient Post Quantum Crypto-Processor for ML-DSA
Jiahao Lu 0002, Lei Chen 0102, Mingbo Wang, Xianqi Mei, Hao Li 0098, Dongsheng Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2026 | A 240 × 180 Event-Based Vision Sensor ROIC With Global Threshold Voltage Calibration TechnologyabstractThis paper presents a$240\times 180$event-based vision sensor (EVS) readout integrated circuit (ROIC) that incorporates global threshold voltage calibration technology. Conventional EVS suffer from error events and intra-array mismatch. This paper delves into four types of error events and presents effective solutions for them. This paper proposes a global threshold voltage calibration (GTVC) technique to minimize the impact of mismatch in the pixel array. The pixel size measures$10\times 10~\mu $m${}^{\mathbf {2}}$, using 40 nm CMOS technology. The analog pixel circuits work with supply voltages of 2.5 V and 1.1 V, while the power consumption of this chip is 12.16 mW. By using the global threshold voltage calibration technology, which incorporates eight reference pixels for determining the input voltage of event comparators, the error in the detection result is reduced to 2.84 mV. The maximum event rate reaches 360 Meps with a 40 MHz system clock, boasting a dynamic range of 97.3 dB. Further, the event power efficiency stands at 29.6 Ge/W. Yanwen Su, Hao Li 0098, Kaiyue Li, Ang Hu, Zhichen Yang, Luxin Yan, Dongsheng Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | An Efficient and Reconfigurable Post-Quantum Crypto-Processor for SPHINCS+abstractSPHINCS+ is the sole hash-based digital signature scheme among the selected post-quantum cryptography (PQC) in 2022. This algorithm possesses the ability to resist attacks from both classical and quantum computers. Due to the extensive computations and different data widths for various parameters, its hardware implementation faces the weakness of long operation time, large area requirement, and low flexibility. This paper presents an efficient and reconfigurable SPHINCS+ processor. The proposed on-the-fly WOTS+ public key generation scheme with unified chain address generator accelerated the most time-consuming operations. This optimization achieves efficient resource utilization. A security switch mechanism resolves the bit misalignment among different data widths with resource reduction. Finally, we introduce a grouped subtree and segmented signature streaming scheme. They reduce the memory to 16k bytes. The processor consumes 29410 LUTs, 14090 FFs, 4 BRAMs on Artix-7 FPGA and achieves$1.04\times/2.41\times $ATPs (area-time-product) optimizations in Sign/Verify with the advantage of supporting all security levels of SPHINCS+. Jiahao Lu 0002, Dongsheng Liu 0001, Zhixiang Luo, Lei Chen 0102, Xiang Li 0220 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | A Timing Attack Resistant Lightweight Post-Quantum Crypto-Processor for SPHINCS+abstractSPHINCS+ is a selected hash-based post-quantum cryptography (PQC) for digital signature due to its promised security and simple cryptographic operations. However, the excessive storage requirements and potential side channel attack may be the obstacle in its practical application. In this paper, a compact signature streaming scheme is proposed, saving storage consumption by 88%. Besides, two timing attacks: breakpoint and saving signature attack are presented. These crack critical information for grafting tree attack to carry out a universal forgery. A fake signature compensation protection is designed as their countermeasure with 2×128 bits register and free storage. The lightweight, timing attack resistant processor consumes 5832 LUT, 2794 FF, 1 BRAM, representing lowest/second-lowest consumption in LUT(FF)/BRAM among existing designs, and achieves 1.11×/7.25×/20× area time product (ATP) optimizations in LUT/FF/BRAM. Jiahao Lu 0002, Dongsheng Liu 0001, Lei Chen 0102, Xiang Li 0220 |
ISCAS | 3 |
| 2023 | A Flexible and High-Performance Lattice-Based Post-Quantum Crypto Secure CoprocessorabstractProgress of quantum computing technology seriously threaten the industrial information security based on traditional public-key cryptosystem. Thus, the cryptosystem with anti-quantum attack characteristics is gradually becoming a significant research in the security field. In this article, a flexible and high-performance secure coprocessor is designed for security in industrial processes, which can execute the post-quantum cryptographic algorithm Saber efficiently. Custom instruction set and arithmetic accelerators are proposed to effectively optimize the flexibility of system architecture, and improve the performance of calculation. The hardware implementation results show that the maximum operating frequency of the coprocessor can reach 345 MHz. Compared with related state-of-the-art works, it achieves the highest operating frequency on the same Xilinx UltraScale+ FPGA platform, performing the encryption and decryption operations within 13.5 and 15.4μs, respectively. Meanwhile, this article achieves 1.7/3.1/5.9× area-time product improvements in look-up table flip-flop block memory storage with good flexibility. Dongsheng Liu 0001, Xiang Li 0220, Xingjie Liu, Jiahao Lu 0002, Xuecheng Zou, Ang Hu, Tianming Ni |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | An Efficient Unstructured Sparse Convolutional Neural Network Accelerator for Wearable ECG Classification DeviceabstractConvolution neural network (CNN) with pruning techniques has shown remarkable prospects in electrocardiogram (ECG) classification. However, efficiently deploying the existing pruned neural network to wearable devices for ECG classification is a great challenge due to the limited hardware resource and randomly distributed sparse weights. To address this issue, an efficient unstructured sparse CNN accelerator is proposed in this paper. A tile-first dataflow with compressed data storage format is presented to skip zero weight multiplications and increase the computing efficiency during inference of small-scale model with large sparsity. The two-level weight index matching structure in the dataflow exploits shifting operation to select valid data pairs and maintain the fully-pipelined calculation process. A configurable processing element (PE) array with 32-bit instruction control is proposed to increase the flexibility of the accelerator. Verified in FPGA and post-synthesis simulations in SMIC 40nm process, the proposed sparse CNN accelerator consumes$3.93~\mu $J/classification at 2MHz clock frequency and it achieves an averaged ECG classification accuracy of 98.99%. A computing efficiency of 118.75% is realized which is improved by 48% compared to the dense baseline. In brief, the proposed efficient CNN accelerator is especially suitable for wearable ECG classification device. Jiahao Lu 0002, Dongsheng Liu 0001, Ang Hu, Xuecheng Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Towards Efficient Hardware Implementation of NTT for Kyber on FPGAsabstractKyber is a promising lattice-based post-quantum cryptography (PQC) for key encapsulation mechanisms. Since number theoretic transform (NTT) is the most computationally expensive operation in Kyber, this paper focuses on novel optimization techniques for efficient hardware implementation of NTT for Kyber. Benefiting from the proposed fast modular multiplication method and a doubled bandwidth ping-pong memory access scheme, our NTT architecture can complete an NTT operation in Kyber in 490 cycles using only 609 LUTs, 640 FFs, 2 DSPs on a Xlinx Artix-7 FPGA. The proposed NTT architecture is 3.95 times faster than the state-of-the-art design for Kyber while achieves an improvement of more than 1.5 times in the area time product (ATP). Compared with the state-of-the- art NTT designs for other algorithms, our NTT architecture reduces 24% FFs and 50% DSPs and ranks second smallest in ATPs, which can also confirm the high efficiency of our design. Dongsheng Liu 0001, Xingjie Liu, Xuecheng Zou, Guangda Niu, Bo Liu 0063, Quming Jiang |
ISCAS | 2 |
| 2021 | Efficient Hardware Architecture of Convolutional Neural Network for ECG Classification in Wearable Healthcare DeviceabstractNowadays, with the increasing shortage of traditional medical resources, the existing portable monitoring healthcare device is no longer satisfactory. Thus, wearable healthcare device with diagnostic capability is becoming much more desirable. However, the design of wearable healthcare device faces the challenge of limited hardware resource and high diagnostic accuracy. In this paper, an efficient hardware architecture is proposed to implement a 1-D CNN with global average pooling (GAP) specially for embedded electrocardiogram (ECG) classification. The GAP is implemented by substituting division into shifting operation without extra computing resource consumption and it can largely reduce the parameters of the network. The fully pipelined processing unit (PU) array is designed to increase computing efficiency. A sign bit based dynamic activation strategy is developed for removing redundant multiplications and resource consumption of ReLU. The proposed efficient hardware architecture is implemented on Xilinx Zynq ZC706 board and achieves an average performance of 25.7 GOP/s under 200-MHz with resource consumption of 1538 LUT, which makes resource efficiency improved by more than 3× compared with non-optimized case. The averaged classification accuracy of five ECG beats classes is 99.10%. In brief, the proposed efficient hardware design is prospective for wearable healthcare device especially in ECG classification area. Jiahao Lu 0002, Dongsheng Liu 0001, Xuecheng Zou, Bo Liu 0063 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2020 | A Flexible and Generic Gaussian Sampler With Power Side-Channel Countermeasures for Quantum-Secure Internet of ThingsabstractPost-quantum cryptography (PQC) great potential in providing reliable communication security for Internet-of-Things (IoT) devices against the quantum computer in the future. The Gaussian sampler is a crucial part in lattice-based post-quantum cryptosystems, thus being the most vulnerable module to side-channel attack as well. However, research on the countermeasures for the Gaussian sampler against power side-channel attacks is almost blank. In this article, a flexible and generic cumulative distribution table (CDT)-based Gaussian sampler using the hardware-software approach is proposed. The proposed CDT sampler has an AHB interface and can be reconfigured to support various parameter sets, while utilizing just 77 Slices on a Xilinx Spartan-6 FPGA with constant response time. Additionally, the first simple power analysis (SPA) attack on the CDT sampler is presented. The presented attack mainly takes advantage of the chosen input and the SPA vulnerability associated with the binary search method, hence the attacker is able to recover every sampled value by comparing a few pairs of power consumption traces. To further protect against chosen input SPA attack, this article identifies the vulnerability associated with three main operations in every binary search state and construct an effective countermeasure based on randomization at the cost of only extra 58.4% Slices. Compared to other related works, the merits of the proposed CDT sampler are the high hardware flexibility, side-channel security, and suitability for resource-constrained IoT nodes. Jiahao Lu 0002, Dongsheng Liu 0001 |
IEEE Internet Things J. | 5 |
| 2020 | Architecture of Cobweb-Based Redundant TSV for Clustered FaultsabstractIn this brief, a cobweb-based redundant through-silicon-via (TSV) design is proposed with efficient hardware as well as high repair rate to repair clustered faulty TSVs (FTSVs). The experimental simulation results demonstrate that for highly clustered faults, the repair rate of the proposed RTSV method is 48.59% and 1.75% higher than that of the ring-based and router-based RTSV methods, respectively. Furthermore, the proposed design can achieve 63.93% and 16.34% hardware reductions compared with the router-based and the ring-based design, respectively. Tianming Ni, Dongsheng Liu 0001, Qi Xu 0004, Zhengfeng Huang, Huaguo Liang, Aibin Yan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | A Methodology for Design of Unbuffered Router Microarchitecture for S-Mesh NoC
Feifei Cao, Dongsheng Liu 0001, Xuecheng Zou |
NPC | 3 |