EDBT 2026 Demo / reviewers in the wild / expert
Zhaoxu Zhou
dblp:249/8957
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0009-5801-8952ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PEDC: A High-Efficacy, Parallel, and Configurable Distributed Control Architecture Design Approach of Cryptographic CGRA Utilizing the Control Subgraph Partitioning and Expansion MethodabstractCoarse-grained reconfigurable architectures (CGRAs) are increasingly employed as cryptographic accelerators due to their efficiency and flexibility. Existing studies on security-oriented CGRAs primarily focus on scaling or optimizing the data path, while comparatively little attention has been given to the control path. Recently, a general-purpose processing core or configuration system has been widely adopted as the controller of CGRA. While this approach significantly reduces the design and application complexity of CGRAs, it does so at the expense of control efficacy and flexibility. To address this issue, a configurable distributed control (D-C) architecture design approach (referred to as PEDC) is proposed, which enhances the parallel processing capability of CGRA by improving control flexibility. There are four key technologies in PEDC. First, control subgraphs are automatically partitioned to define the control scopes of controllers and extract the control nodes of the control framework. Second, control dependency relationships are extracted from the control flow graph to link the control nodes. Third, a control architecture graph is constructed by establishing master-slave relationships between process control nodes capable of independent task process control and other nodes. Lastly, design models for the input scheduling controller, output scheduling controller, cluster controller, and task process controller are presented. The PEDC approach essentially transforms the CGRA into a multi-instruction stream, multi-data stream processor. D-C architectures with various scales are implemented based on 40-nm CMOS technology. With the PEDC approach, multiple pipelines and independent tasks can be processed simultaneously, regardless of the array structure or algorithm type. Compared with the traditional control method, the PEDC achieves a 4.5 × execution efficiency. Compared with related reconfigurable architectures, PEDC enables CGRAs to be more functionally flexible and achieve a better full-load throughput. Zibin Dai, Yanjiang Liu, Danping Jiang, Zhaoxu Zhou |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2026 | DPTM: An Adaptive Scheduler Design Utilizing Timeslot Matching and Release Methods for Concurrent and Multi-task Interleaved Pipelining-oriented CGRAabstractCoarse-grained reconfigurable architectures (CGRAs) are increasingly employed as domain-specific accelerators due to their efficiency and flexibility. However, the existing CGRA architectures suffer from low hardware resource utilization and performance due to the limitations of the scheduling scheme. In this article, an adaptive scheduler (denoted as DPTM) for concurrent and multi-task interleaved pipelining-oriented CGRA is introduced, which exploits timeslot matching and release methods to avoid the pipeline conflicts and improve the scheduling performance. The characteristics of task scheduling based on directed acyclic graph (DAG) are analyzed, and several performance-influencing factors are extracted to build a scheduling performance model for reducing the time cost of scheduling and guiding the design of scheduling schemes. Moreover, the scoreboard method of dynamic instruction schedulers is optimized to control the entry time of multiple tasks into the pipeline, and then a timeslot matching method is proposed to provide non-conflict pipelining for the multiple tasks. Further, a timeslot release method is presented to release the timeslots for unscheduled sub-tasks dynamically, which can adapt the parallel processing of multiple tasks and decrease the scheduling time. Then, an adaptive scheduling scheme combines the dynamic priority-based task assignment method, timeslot matching method, and timeslot release method to schedule massive tasks for CGRA. Finally, the overall architecture of DPTM is introduced and designed to validate the efficacy of the proposed scheduling scheme. Experimental results show that the proposed timeslot matching/release approach reduces 84% total scheduling time and decreases 40% average scheduling time at most compared to the non-timeslot-matching scheduling schemes, the proposed task assignment approach decreases 8% total scheduling time and lowers 3% average scheduling time compared to the existing approaches, and the proposed scheduler decreases 51% critical path delay, lowers 35% area overhead, and reduces 12% power consumption at most compared with the existing schedulers. Danping Jiang, Zibin Dai, Yanjiang Liu, Zhaoxu Zhou |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2025 | CRM_BF: A Low-Overhead, High-Efficient and Reconfigurable Operation Unit Design Approach Using the Customized Reed-Muller Unit For Boolean Functions of Sequence Cipher AlgorithmsabstractSequence ciphers algorithms encrypt or decrypt information at a low cost and high speed compared to other cryptographic algorithms, which are widely applied to critical applications and sensitive fields. As the core component of sequence ciphers, Boolean functions generate the random number or implement the update process of random numbers. The existing implementations of Boolean functions cause a great waste of area resources and generate several long critical paths that limit the hardware performance of sequence ciphers. To address this issue, a 64-bit Boolean Function Reconfigurable Operation Unit (BFROU) is proposed to reduce the area overhead, lower the delay latency, and enhance the operation efficacy of Boolean functions. Through statistical characterization analysis and cutting experiments of Boolean functions, a 64 bits BFROU based on CRM-3 units has been designed, which has the advantage of low-cost and high-efficient。The CRM unit is customized based on RM logic. A theoretical framework for Boolean functions is proposed by combining CRM units with mathematical expressions, which encompasses Boolean functions for any variable. On the platform of synthesis software, based the theoretical architecture, a CRM-OPT optimization algorithm is proposed, which can achieve the conversion of And Inverter Graph (AIG) to Customized Reed Muller Graph (CRMG).This Customized Reed-Muller (CRM) unit achieved at least 22.4% and 25.1% optimization in delay and area compared to Universal Reed-Muller (URM) units. The experimental results show that the Area Delay Product (ADP) is minimized when the CRM-3 unit is the optimal maximum cutting size. Ultimately, the BFROU design was realized utilizing CRM units, achieving an area of 195.4um² and a critical path delay of 0.35ns. This BFROU can achieve special Boolean functions involving 64 variables at maximum,with 91% of these functions being mapped within two iterations. Moreover, this BFROU has significant advantages over other known schemes regarding area, critical path delay, ADP, and number of iterations consumed. Zhaoxu Zhou, Junwei Li 0007, Yanjiang Liu, Zibin Dai |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2021 | Performance analysis of polling-based MAC protocol with retrial for Internet of ThingsabstractSummary We consider and analyze a single‐server multiqueue polling model with inner arrivals. Customers arriving at the queue before polling instant could receive service in the current polling round; furthermore, each one could be retried (turns into an inner arrival) a given number of times with a specified probability. Such polling model can be used to study the performance of certain scheduling data transmission in the Internet of Things (IoT) and the relationship between data retransmission and delay. We obtain the closed‐form expression for the generating function of the amount of customers, which are presented at polling instants. Then, it is used to derive the precise closed‐form formula of mean queue length and mean waiting time in symmetric system. Wenhua Qian, Zhaoxu Zhou |
Concurr. Comput. Pract. Exp. | 5 |