EDBT 2026 Demo / reviewers in the wild / expert
Zibin Dai
dblp:86/2928
· DBLP profile ↗
12ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-0359-8434ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 since 2021Security and privacy · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PEDC: A High-Efficacy, Parallel, and Configurable Distributed Control Architecture Design Approach of Cryptographic CGRA Utilizing the Control Subgraph Partitioning and Expansion MethodabstractCoarse-grained reconfigurable architectures (CGRAs) are increasingly employed as cryptographic accelerators due to their efficiency and flexibility. Existing studies on security-oriented CGRAs primarily focus on scaling or optimizing the data path, while comparatively little attention has been given to the control path. Recently, a general-purpose processing core or configuration system has been widely adopted as the controller of CGRA. While this approach significantly reduces the design and application complexity of CGRAs, it does so at the expense of control efficacy and flexibility. To address this issue, a configurable distributed control (D-C) architecture design approach (referred to as PEDC) is proposed, which enhances the parallel processing capability of CGRA by improving control flexibility. There are four key technologies in PEDC. First, control subgraphs are automatically partitioned to define the control scopes of controllers and extract the control nodes of the control framework. Second, control dependency relationships are extracted from the control flow graph to link the control nodes. Third, a control architecture graph is constructed by establishing master-slave relationships between process control nodes capable of independent task process control and other nodes. Lastly, design models for the input scheduling controller, output scheduling controller, cluster controller, and task process controller are presented. The PEDC approach essentially transforms the CGRA into a multi-instruction stream, multi-data stream processor. D-C architectures with various scales are implemented based on 40-nm CMOS technology. With the PEDC approach, multiple pipelines and independent tasks can be processed simultaneously, regardless of the array structure or algorithm type. Compared with the traditional control method, the PEDC achieves a 4.5 × execution efficiency. Compared with related reconfigurable architectures, PEDC enables CGRAs to be more functionally flexible and achieve a better full-load throughput. Zibin Dai, Yanjiang Liu, Danping Jiang, Zhaoxu Zhou |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2026 | DPTM: An Adaptive Scheduler Design Utilizing Timeslot Matching and Release Methods for Concurrent and Multi-task Interleaved Pipelining-oriented CGRAabstractCoarse-grained reconfigurable architectures (CGRAs) are increasingly employed as domain-specific accelerators due to their efficiency and flexibility. However, the existing CGRA architectures suffer from low hardware resource utilization and performance due to the limitations of the scheduling scheme. In this article, an adaptive scheduler (denoted as DPTM) for concurrent and multi-task interleaved pipelining-oriented CGRA is introduced, which exploits timeslot matching and release methods to avoid the pipeline conflicts and improve the scheduling performance. The characteristics of task scheduling based on directed acyclic graph (DAG) are analyzed, and several performance-influencing factors are extracted to build a scheduling performance model for reducing the time cost of scheduling and guiding the design of scheduling schemes. Moreover, the scoreboard method of dynamic instruction schedulers is optimized to control the entry time of multiple tasks into the pipeline, and then a timeslot matching method is proposed to provide non-conflict pipelining for the multiple tasks. Further, a timeslot release method is presented to release the timeslots for unscheduled sub-tasks dynamically, which can adapt the parallel processing of multiple tasks and decrease the scheduling time. Then, an adaptive scheduling scheme combines the dynamic priority-based task assignment method, timeslot matching method, and timeslot release method to schedule massive tasks for CGRA. Finally, the overall architecture of DPTM is introduced and designed to validate the efficacy of the proposed scheduling scheme. Experimental results show that the proposed timeslot matching/release approach reduces 84% total scheduling time and decreases 40% average scheduling time at most compared to the non-timeslot-matching scheduling schemes, the proposed task assignment approach decreases 8% total scheduling time and lowers 3% average scheduling time compared to the existing approaches, and the proposed scheduler decreases 51% critical path delay, lowers 35% area overhead, and reduces 12% power consumption at most compared with the existing schedulers. Danping Jiang, Zibin Dai, Yanjiang Liu, Zhaoxu Zhou |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | CRM_BF: A Low-Overhead, High-Efficient and Reconfigurable Operation Unit Design Approach Using the Customized Reed-Muller Unit For Boolean Functions of Sequence Cipher AlgorithmsabstractSequence ciphers algorithms encrypt or decrypt information at a low cost and high speed compared to other cryptographic algorithms, which are widely applied to critical applications and sensitive fields. As the core component of sequence ciphers, Boolean functions generate the random number or implement the update process of random numbers. The existing implementations of Boolean functions cause a great waste of area resources and generate several long critical paths that limit the hardware performance of sequence ciphers. To address this issue, a 64-bit Boolean Function Reconfigurable Operation Unit (BFROU) is proposed to reduce the area overhead, lower the delay latency, and enhance the operation efficacy of Boolean functions. Through statistical characterization analysis and cutting experiments of Boolean functions, a 64 bits BFROU based on CRM-3 units has been designed, which has the advantage of low-cost and high-efficient。The CRM unit is customized based on RM logic. A theoretical framework for Boolean functions is proposed by combining CRM units with mathematical expressions, which encompasses Boolean functions for any variable. On the platform of synthesis software, based the theoretical architecture, a CRM-OPT optimization algorithm is proposed, which can achieve the conversion of And Inverter Graph (AIG) to Customized Reed Muller Graph (CRMG).This Customized Reed-Muller (CRM) unit achieved at least 22.4% and 25.1% optimization in delay and area compared to Universal Reed-Muller (URM) units. The experimental results show that the Area Delay Product (ADP) is minimized when the CRM-3 unit is the optimal maximum cutting size. Ultimately, the BFROU design was realized utilizing CRM units, achieving an area of 195.4um² and a critical path delay of 0.35ns. This BFROU can achieve special Boolean functions involving 64 variables at maximum,with 91% of these functions being mapped within two iterations. Moreover, this BFROU has significant advantages over other known schemes regarding area, critical path delay, ADP, and number of iterations consumed. Zhaoxu Zhou, Junwei Li 0007, Yanjiang Liu, Zibin Dai |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2024 | RGMU: A High-flexibility and Low-cost Reconfigurable Galois Field Multiplication Unit Design Approach for CGRCAabstractFinite field multiplication is a non-linear transformation operator that appears in the majority of symmetric cryptographic algorithms. Numerous specified finite field multiplication units have been proposed as a fundamental module in the coarse-grained reconfigurable cipher logic array to support more cryptographic algorithms; however, it will introduce low flexibility and high overhead, resulting in reduced performance of the coarse-grained reconfigurable cipher logic array. In this article, a high-flexibility and low-cost reconfigurable Galois field multiplication unit (RGMU) is proposed to balance the tradeoffs between the function, delay, and area. All the finite field multiplication operations, including maximum distance separable matrix multiplication, parallel update of Fibonacci linear feedback shift register, parallel update of Galois linear feedback shift register, and composite field multiplication, are analyzed and two basic operation components are abstracted. Further, a reconfigurable finite field multiplication computational model is established to demonstrate the efficacy of reconfigurable units and guide the design of RGMU with high performance. Finally, the overall architecture of RGMU and two multiplication circuits are introduced. Experimental results show that the RGMU can not only reduce the hardware overhead and power consumption but also has the unique advantage of satisfying all the finite field multiplication operations in symmetric cryptography algorithms. Danping Jiang, Zibin Dai, Yanjiang Liu, Zongren Zhang |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | Extending the classical side-channel analysis framework to access-driven cache attacks
Yingjian Yan, Fan Zhang 0010, Chunsheng Zhu, Zibin Dai |
Comput. Secur. | 6 |
| 2023 | CBDC-PUF: A Novel Physical Unclonable Function Design Framework Utilizing Configurable Butterfly Delay Chain Against Modeling AttackabstractPhysical unclonable function (PUF) is a promising security-based primitive, which provides an extremely large number of responses for key generation and authentication applications. Various PUFs have been developed as central building blocks in cryptographic protocols and security architectures, however, the existing PUFs and their improvements are still vulnerable to modeling attacks (MA) with refined machine learning algorithms. In this article, a configurable butterfly delay chain-based PUF design framework is proposed to meet the requirements of randomness, reliability, uniqueness, and MA-resistance metrics. A configurable butterfly delay chain is introduced to create multiple pairs of symmetric paths and a strong PUF relying on the intrinsic delay fluctuations of two identical paths is built. Furthermore, a secure hash function is used to insert non-linearities into the PUF, and a BCH-based error correction algorithm is utilized to recover the actual responses under noisy environments. The proposed PUF is implemented on Xilinx FPGAs and three machine learning algorithms are used to evaluate the resistance against MA. Experimental results show that the randomness, reliability, and uniqueness of the proposed PUF are close to the ideal value (49.6%, 99.9%, and 49.9%, respectively), and the prediction accuracy reaches 50% that indicating a desirable resilient to MA. Yanjiang Liu, Junwei Li 0007, Tongzhou Qu, Zibin Dai |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2022 | A Comprehensive Evaluation of Integrated Circuits Side-Channel Resilience Utilizing Three-Independent-Gate Silicon Nanowire Field Effect Transistors-Based Current Mode LogicabstractSide-channel attack (SCA) is one of the physical attacks, which will reveal the confidential information from cryptographic circuits by statistically analyzing physical manifestations. Various circuit-level countermeasures have been proposed as fundamental solutions to eliminate the correlations between side-channel information and circuit’s internal operations. The existing solutions, however, will introduce nonnegligible power and area overheads, making them difficult to be deployed in resource-constrained applications. In this article, a novel three-independent-gate silicon nanowire field effect transistor (TIGFET) with the intrinsic SCA-resilience characteristics is introduced to balance the tradeoffs among cost, performance, and security of cryptographic implementations. We construct six TIGFET-based current mode logic (CML) gates that can retain lower power variation under all possible transitions compared to the CMOS counterparts. As a proof of concept, advanced encryption standard (AES), SM4 block cipher algorithm (SM4), and lightweight cryptographic algorithm PRESENT are implemented utilizing the TIGFET-based CML gates. Correlation power attack is performed to evaluate the improvement of SCA resilience. Simulation results verify that the TIGFET-based cryptographic implementations decrease 42.37% area usage, lower 61.16% energy efficiency, reduce$5.35\times $power variation, and achieve a similar level of SCA resistance compared to the CMOS counterpart, which is applicable for the resource-constrained applications. Yanjiang Liu, Jiaji He 0001, Haocheng Ma, Tongzhou Qu, Zibin Dai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | A Low-Overhead and High-Security Cryptographic Circuit Design Utilizing the TIGFET-Based Three-Phase Single-Rail Pulse Register against Side-Channel AttacksabstractSide-channel attack (SCA) reveals confidential information by statistically analyzing physical manifestations, which is the serious threat to cryptographic circuits. Various SCA circuit-level countermeasures have been proposed as fundamental solutions to reduce the side-channel vulnerabilities of cryptographic implementations; however, such approaches introduce non-negligible power and area overheads. Among all of the circuit components, flip-flops are the main source of information leakage. This article proposes a three-phase single-rail pulse register (TSPR) based on the three-independent-gate field effect transistor (TIGFET) to achieve all desired properties with improved metrics of area and security. TIGFET-based TSPR consumes a constant power (MCV is 0.25%), has a low delay (12 ps), and employs only 10 TIGFET devices, which is applicable for the low-overhead and high-security cryptographic circuit design compared to the existing flip-flops. In addition, a set of TIGFET-based combinational basic gates are designed to reduce the area occupation and power consumption as much as possible. As a proof of concept, a simplified advanced encryption algorithm (AES), SM4 block cipher algorithm (SM4), and light-weight cryptographic algorithm (PRESENT) are built with the TIGFET-based library. SCA is implemented on the cryptographic implementations to prove its SCA resilience, and the SCA results show that the correct key of cryptographic circuits with TIGFET-based TSPRs is not guessed within 2,000 power traces. Yanjiang Liu, Tongzhou Qu, Zibin Dai |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2020 | PVHArray: An Energy-Efficient Reconfigurable Cryptographic Logic Array With Intelligent MappingabstractThis article presents a coarse-grained reconfigurable cryptographic logic array named PVHArray and an intelligent mapping algorithm for cryptographic algorithms. We propose three techniques to improve energy efficiency without affecting performance. First, the coarse-grained pipeline variable reconfigurable operation units balance the system critical path delay and number of algorithm operations to ensure the best performance. Second, the hierarchical interconnect network overcomes the shortcomings of a single network, providing PVHArray with good interconnectivity and scalability while managing the network hardware resource overhead. Third, the distributed control network supports accurate period-oriented control with a lightweight hardware structure, preserving hardware resources for other performance enhancements. We combine these advances with deep learning to propose a type of smart ant colony optimization mapping algorithm to improve algorithm mapping performance. We implemented our PVHArray on a 12.25 mm2silicon square with 55-nm CMOS technology, with each algorithm working at its optimum frequency. Experiments show that PVHArray improved performance by about 12.9% per unit area and 13.9% per unit power compared with the reconfigurable cryptographic logic array REMUS_LPP and other state-of-the-art cryptographic structures. For cryptographic algorithm mapping, our smart ant colony optimization (SACO) algorithm reduced compilation time by nearly 38%. Finally, PVHArray supports a variety of types of cryptographic algorithms. Yiran Du, Wei Li 0131, Zibin Dai, Longmei Nan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | A highly efficient reconfigurable rotation unit based on an inverse butterfly networkabstractWe propose a reconfigurable control-bit generation algorithm for rotation and sub-word rotation operations. The algorithm uses a self-routing characteristic to configure an inverse butterfly network. In addition to being highly parallelized and inexpensive, the algorithm integrates the rotation-shift, bi-directional rotation-shift, and sub-word rotation-shift operations. To our best knowledge, this is the first scheme to accommodate a variety of rotation operations into the same architecture. We have developed the highly efficient reconfigurable rotation unit (HERRU) and synthesized it into the Semiconductor Manufacturing International Corporation (SMIC)’s 65-nm process. The results show that the overall efficiency (relative area×relative latency) of our HERRU is higher by at least 23% than that of other designs with similar functions. When executing the bi-directional rotation operations alone, HERRU occupies a significantly smaller area with a lower latency than previously proposed designs. Zibin Dai, Wei Li 0131, Hai-juan Zang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2009 | The Research of NULL Convention Logic Circuit Computing Model Targeted at Block Cipher ProcessingabstractOn the basis of the operation characteristics of NULL Convention Logic (NCL) circuit and popular block cipher algorithms, this paper proposes a model of NCL circuits computing, which targeted at characteristic of block cipher, and then maps Rijndael algorithm on the model based on NCL circuit by two ways. The circuit model has been validated in Balsa system, and synthesized with Balsa using 1 mum example NCL Cell Library. The experiment results indicate that the circuit model has high flexibility for block cipher algorithms. Yalei Cui, Zibin Dai |
IAS | 2 |
| 2007 | Design and Implementation of a High-Speed Reconfigurable Modular Arithmetic Unit
Wei Li 0131, Zibin Dai, Tao Chen 0047, Xuan S. Yang |
APPT | 2 |