EDBT 2026 Demo / reviewers in the wild / expert
Donghui Guo
dblp:56/2415
· DBLP profile ↗
24ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S2M: Spatiotemporal-Cooperative Symbiotic Memory for Heterogeneous PQC Acceleration
Sizhao Li, Chenyu Zhai, Bingrui Guo, Donghui Guo |
APPT | 6 |
| 2026 | Uni-Winograd: A Massively Resource-Efficient and Unified Radix-4 NTT Architecture for PQC Algorithms
Danni Wang, Sizhao Li, Guisheng Yin, Donghui Guo |
APPT | 6 |
| 2026 | Spiking neural networks with uncertainty model of stochastic sampling for circuit yield enhancement
Zenan Huang, Wenrun Xiao, Haojie Ruan, Shan He 0003, Donghui Guo |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | RePM: Reconfigurable Elastic Computing for Polynomial Multiplier with Hybrid NTT AlgorithmabstractPolynomial multiplication, a core component of lattice-based cryptography, has demonstrated impressive performance in lattice-based cryptographic chips. However, costly Number Theoretic Transform (NTT) makes efficient and flexible hardware design extremely challenging, particularly for ASIC/FPGA-based hardware acceleration solutions that often struggle with high resource consumption and low computational efficiency. To address these issues, we propose a lightweight reconfigurable polynomial multiplication accelerator for NTT termed RePM. The hardware architecture adheres to the principles that are applicable to various post-quantum cryptography (PQC) algorithms and operates under very strict power constraints. A conflict-free near-memory mapping scheme is adopted to reconstruct the computation topology, significantly reducing the frequency of data relocation. Moreover, this work pioneers an elastic computing array architecture for butterfly units that eliminates pre-computing requirements. By integrating the mixed radix-2/4 NTT algorithms proposed in this paper, we eliminate both pre-processing and post-processing stages in the computational workflow. This innovation achieves a 2.1 × acceleration in ML-KEM (i.e. CRYSTALS-Kyber) execution speed with ML-KEM-512/1024, positioning it as a leading implementation of NIST’s fourth-round post-quantum cryptography standard. Experimental results demonstrate that this architecture achieves up to 75% power reduction and 65.8% improvement in area efficiency compared to state-of-the-art solutions, while preserving a 17× higher computational throughput under strict power constraints. Sizhao Li, Chenyu Zhai, Zhujun Guo, Shan He 0003, Donghui Guo |
ICCAD | 6 |
| 2025 | GSNN: A Neuromorphic Computing Model for the Flexible Path Planning in Various Constraint EnvironmentsabstractSpiking Neural Networks (SNNs) represent a new generation of artificial neural networks that draw inspiration from biological systems. However, due to the intricate dynamics they exhibit and the discontinuity inherent in spike signals, SNNs often encounter performance limitations when addressing optimization problems. In this paper, we introduce the Graph-connected Spiking Neural Network model (GSNN), an extension of the SNN framework. The GSNN model holds the potential for integration with various existing path planning methods, rendering it applicable to a wide array of common path planning tasks. We specifically present two fundamental models within the GSNN framework. The first model employs GSNN to extract heuristic information from constrained pixel maps. This extracted data is then amalgamated with a novel sampling method, resulting in enhanced planning efficiency when compared to conventional techniques. The second model leverages GSNN to map a weighted graph, effectively utilizing plasticity methods to ascertain the shortest path within the graph. Moreover, this model facilitates path planning under diverse constraint environments, encompassing dynamic considerations, cost-awareness, and the collision dimensions of moving objects. Recognizing that the size of pixel maps or the number of nodes within weighted graphs might constrain GSNN’s capabilities, we propose a partitioning strategy to address this limitation. Empirical results unequivocally demonstrate the superiority of both GSNN models in resolving static path planning problems. Furthermore, the second GSNN model demonstrates rational performance across various constrained scenarios.Note to Practitioners—The primary motivation behind this study is to explore the utilization of neural morphic computation methods in addressing path planning challenges. To accomplish this, we introduce an encompassing model named GSNN. Given the limited coverage of neural morphic computation within this field, we present a comprehensive overview of its potential applications, with the intention of providing a valuable reference for both researchers and practitioners. Moreover, GSNN has the capability to seamlessly integrate with existing advanced optimization methods, thereby leading to enhanced performance and the capacity to tackle even more intricate problems. Haojie Ruan, Yinghui Chang, Weikang Wu, Zenan Huang, Yabin Deng, Hongyan Luo, Shan He 0003, Donghui Guo |
IEEE Trans Autom. Sci. Eng. | 9 |
| 2023 | Refined Self-calibration of an Inductorless Low-noise Amplifier with Non-intrusive Circuit
Wenrun Xiao, Jidong Diao, Yanping Qiao, Xianming Liu 0002, Shan He 0003, Donghui Guo |
J. Electron. Test. | 6 |
| 2021 | Particle Swarm Optimization Algorithm With Self-Organizing Mapping for Nash Equilibrium Strategy in Application of Multiobjective OptimizationabstractIn this article, the Nash equilibrium strategy is used to solve the multiobjective optimization problems (MOPs) with the aid of an integrated algorithm combining the particle swarm optimization (PSO) algorithm and the self-organizing mapping (SOM) neural network. The Nash equilibrium strategy addresses the MOPs by comparing decision variables one by one under different objectives. The randomness of the PSO algorithm gives full play to the advantages of parallel computing and improves the rate of comparison calculation. In order to avoid falling into local optimal solutions and increase the diversity of particles, a nonlinear recursive function is introduced to adjust the inertia weight, which is called the adaptive particle swarm optimization (APSO). In addition, the neighborhood relations of current particles are constructed by SOM, and the leading particles are selected from the neighborhood to guide the local and global search, so as to achieve convergence. Compared with several advanced algorithms based on the eight multiobjective standard test functions with different Pareto solution sets and Pareto front characteristics in examples, the proposed algorithm has a better performance. Chenhui Zhao, Donghui Guo |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Global Exponential Stability of Hybrid Non-autonomous Neural Networks with Markovian Switching
Chenhui Zhao, Donghui Guo |
Neural Process. Lett. | 2 |
| 2019 | Energy efficiency clustering based on Gaussian network for wireless sensor networkabstractIn this study, the authors consider the sensors' connection model design and energy optimisation of routing protocol in wireless sensor networks based on the Gaussian network connection model. Accordingly, by the node‐symmetric and four different directions to four adjacent nodes of each node in the Gaussian network model, the study recommends a new wireless sensor network connection model, on which the network area will be divided into some virtual square grids. In the new wireless sensor network connection model, they describe each virtual square grid as a node in the Gaussian network, therefrom, they propose a routing method, which is a combination of the shortest path routing protocol in the Gaussian network and clustering protocol to improve the routing efficiency of the wireless sensor network. Some simulations were implemented in NS2, the results show that the proposed routing is very good. Dung Nguyen Quoc, Lvqing Bi, Yangbing Wu, Shan He 0003, Donghui Guo |
IET Commun. | 6 |
| 2019 | A Hardware-Efficient Block Matching Algorithm and Its Hardware Design for Variable Block Size Motion Estimation in Ultra-High-Definition Video EncodingabstractVariable block size motion estimation has contributed greatly to achieving an optimal interframe encoding, but involves high computational complexity and huge memory access, which is the most critical bottleneck in ultra-high-definition video encoding. This article presents a hardware-efficient block matching algorithm with an efficient hardware design that is able to reduce the computational complexity of motion estimation while providing a sustained and steady coding performance for high-quality video encoding. A three-level memory organization is proposed to reduce memory bandwidth requirement while supporting a predictive common search window. By applying multiple search strategies and early termination, the proposed design provides 1.8 to 3.7 times higher hardware efficiency than other works. Furthermore, on-chip memory has been reduced by 96.5% and off-chip bandwidth requirement has been reduced by 39.4% thanks to the proposed three-level memory organization. The corresponding power consumption is only 198mW at the highest working frequency of 500MHz. The proposed design is attractive for high-quality video encoding in real-time applications with low power consumption. Jianwei Zheng 0002, Chao Lu 0005, Jiefeng Guo, Deming Chen, Donghui Guo |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2018 | Race-Condition-Aware and Hardware-Oriented Task Partitioning and Scheduling Using Entropy MaximizationabstractIn a multithreaded execution environment, race condition leads to computational errors and system hazards. Up to date, a series of task scheduling strategies have been presented in the literature to reduce the risk of race condition. Because of the increasing complexity of this problem in multi-core systems, existing task scheduling approaches are not very efficient. To deal with this challenge, in this work, we develop a race-condition-aware and hardware-oriented task partitioning and scheduling algorithm using entropy maximization model. We model uncertainty as a probabilistic occurrence within a time interval, and hence the characteristics of event ordering are analyzed through an uncertainty matrix. Next, a metric is developed to measure the uncertainty of task execution in various execution environments. Finally, a maximum entropy model is generated to ensure the lowest probability of race condition during task execution. The smallest one among maximum entropy values is chosen and used in our proposed task scheduling algorithm. Experimental results show that the proposed task scheduling strategy based on our maximum entropy model outperforms existing state-of-the-art approaches. For example, in a 128-core computing system, the task execution time, CPU utilization ratio, and throughput of our proposed task scheduling is improved by$15.3\sim 36.4$percent,$8.2\sim 17.6$percent, and$20.7\sim 41.4$percent, respectively. Moreover, our proposed scheduling algorithm exhibits low computational complexity and good adaptivity to diverse execution environments. Sizhao Li, Yuanzhi Zhang 0004, Hongyin Luo, Chao Lu 0005, Donghui Guo |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2017 | A high quality compiler tool for application-specific instruction-set processors with library and parallel supports
Benbin Chen, Chung-Ta King, Donghui Guo |
Multim. Tools Appl. | 4 |
| 2016 | Gene regulatory network inference using PLS-based methodsabstractBACKGROUND: Inferring the topology of gene regulatory networks (GRNs) from microarray gene expression data has many potential applications, such as identifying candidate drug targets and providing valuable insights into the biological processes. It remains a challenge due to the fact that the data is noisy and high dimensional, and there exists a large number of potential interactions. RESULTS: We introduce an ensemble gene regulatory network inference method PLSNET, which decomposes the GRN inference problem with p genes into p subproblems and solves each of the subproblems by using Partial least squares (PLS) based feature selection algorithm. Then, a statistical technique is used to refine the predictions in our method. The proposed method was evaluated on the DREAM4 and DREAM5 benchmark datasets and achieved higher accuracy than the winners of those competitions and other state-of-the-art GRN inference methods. CONCLUSIONS: Superior accuracy achieved on different benchmark datasets, including both in silico and in vivo networks, shows that PLSNET reaches state-of-the-art performance. Shun Guo, Qingshan Jiang, Lifei Chen, Donghui Guo |
BMC Bioinform. | 4 |
| 2016 | Enhanced pipelined architecture of H.264/AVC intra prediction
Jiefeng Guo, Jianwei Zheng 0002, Donghui Guo |
Signal Process. Image Commun. | 4 |
| 2016 | Modeling of Gaussian Network-Based Reconfigurable Network-on-Chip DesignsabstractIn network on chips (NoCs) design, reconfiguration of NoC is a very effective option for minimizing power consumption, and Gaussian networks can provide significant advantage over the mesh networks in terms of network diameter, average hop distance and so on. In this paper, based on the special topology structure and the static connection rules within Gaussian networks, we present the reconfiguration representation for Gaussian networks and the nature of reconfigurable Gaussian networks. Furthermore, reconfigurable rules of Gaussian networks are proposed to design the constraints for automatic reconfiguration of NoC. Yangbing Wu, Deming Chen, Donghui Guo |
IEEE Trans. Computers | 4 |
| 2016 | FCUDA-NoC: A Scalable and Efficient Network-on-Chip Implementation for the CUDA-to-FPGA FlowabstractHigh-level synthesis (HLS) of data-parallel input languages, such as the Compute Unified Device Architecture (CUDA), enables efficient description and implementation of independent computation cores. HLS tools can effectively translate the many threads of computation present in the parallel descriptions into independent, optimized cores. The generated hardware cores often heavily share input data and produce outputs independently. As the number of instantiated cores grows, the off-chip memory bandwidth may be insufficient to meet the demand. Hence, a scalable system architecture and a data-sharing mechanism become necessary for improving system performance. The network-on-chip (NoC) paradigm for intrachip communication has proved to be an efficient alternative to a hierarchical bus or crossbar interconnect, since it can reduce wire routing congestion, and has higher operating frequencies and better scalability for adding new nodes. In this paper, we present a customizable NoC architecture along with a directory-based data-sharing mechanism for an existing CUDA-to-FPGA (FCUDA) flow to enable scalability of our system and improve overall system performance. We build a fully automated FCUDA-NoC generator that takes in CUDA code and custom network parameters as inputs and produces synthesizable register transfer level (RTL) code for the entire NoC system. We implement the NoC system on a VC709 Xilinx evaluation board and evaluate our architecture with a set of benchmarks. The results demonstrate that our FCUDA-NoC design is scalable and efficient and we improve the system execution time by up to 63× and reduce external memory reads by up to 81% compared with a single hardware core implementation. Yao Chen 0008, Swathi T. Gurumani, Yun Liang 0001, Guofeng Li, Donghui Guo, Kyle Rupnow, Deming Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2015 | Uncertainty Analysis of Race Conditions in Real-Time SystemsabstractRace conditions in real-time systems may cause unexpected computing result. Due to the uncertainty of realtime systems, a race condition detected by many static and dynamic approaches may occur in one execution environment but may not occur in another execution environment. In this paper, an easy and practical approach based on probabilistic models is presented to analyze the uncertainties of race conditions of real-time systems in various execution environments. The approach adopts a probabilistic occurrence within a time interval to represent the uncertainty of event occurrences. The confidence level is defined to measure the accuracy of the time interval observed, and then the uncertainties of event orders are analyzed according to the relations of time intervals. We propose a T-matrix to describe the uncertainties of event orders, and a metric is presented to measure the uncertainties of executions of real-time systems in various environments. Moreover, another metric is introduced to measure the total risk of the real-time system caused by race conditions in various environments. Shan He 0003, Sizhao Li, Donghui Guo |
QRS | 4 |
| 2014 | Hybrid circuit-switched network for on-chip communication in large-scale chip-multiprocessors
Hongyin Luo, Shaojun Wei, Deming Chen, Donghui Guo |
J. Parallel Distributed Comput. | 4 |
| 2013 | Robustness analysis of full implication inference method
Songsong Dai, Daowu Pei, Donghui Guo |
Int. J. Approx. Reason. | 3 |
| 2011 | Ties within Fault Localization rankings: Exposing and Addressing the ProblemabstractSoftware fault localization techniques typically rank program components, such as statements or predicates, in descending order of their suspiciousness (likelihood of being faulty). During debugging, programmers may examine these components, starting from the top of the ranking, in order to locate faults. However, the assigned suspiciousness to each component may not always be unique, and thus some of them may be tied for the same position in the ranking. In such a scenario, the total number of components that a programmer needs to examine in order to find the faults may vary considerably. The greater the variability, the harder it is for a programmer to decide which component to examine first, and the harder it is to accurately compute the expected effectiveness of a fault localization technique. In this paper, we first conduct a case study, based on three fault localization techniques across four sets of programs, which reveals that the phenomenon of assigning the same suspiciousness to multiple components is not limited to any technique or program in particular. Thus, to reduce variability and alleviate this problem, four tie-breaking strategies are discussed and evaluated empirically in our second case study. Results indicate that the strategies can not only reduce the number of ties in the rankings, but also maintain the effectiveness of the fault localization techniques. We also propose a new metric for evaluating fault localization techniques called CScore, which takes the notion of ties into account. Finally, an additional slicing-based approach to breaking ties is discussed briefly, which aims to provide further insights into tie-breaking and stimulate further research in the area. Vidroha Debroy, W. Eric Wong, Donghui Guo |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2010 | An Evaluation of Tie-Breaking Strategies for Fault Localization Techniques
Vidroha Debroy, W. Eric Wong, Donghui Guo |
SEKE | 4 |
| 2006 | TCP Performance Improvement through Inter-layer Enhancement with Mobile IPv6abstractTo enable higher layer transparent, Mobile IPv6 (MIPv6) hides mobility from the transport layer such as TCP, but it has serious implications on TCP due to mobility issues including the packet loss and the deviation of end-to-end transport delay, which may seriously degrade the transport performance of TCP. This paper explores the impact of MIPv6 on TCP and proposes mobility control mechanisms through the inter-layer enhancement to enable TCP aware of mobility-related behaviours, and react correctly. The changes are made to the network layer and transport layer at the endpoints, and preserve the end-to-end semantics of TCP. One part of the modifications, the inter-layer control module, monitors the mobility- related procedures executed by MIPv6 and transfers mobility-related information via the inter-layer control module to TCP. The second part is to response to mobility-related behaviours correctly by TCP to alleviate the problems caused by MIPv6 mobility issues. As we show through performance measurements, our approach significantly improves the transport performance of TCP in MIPv6-based mobile computing environments. Deguang Le, Donghui Guo |
GLOBECOM | 2 |
| 2003 | Mobile IPv6 in WLAN mobile networks and its implementationabstractThis paper is to introduce the WLAN mobile networks and analyze the handover problem in WLAN mobile networks. We will propose a new solution of the integration of WLAN and mobile IPv6, which can realize the seamless mobile communications. In order to evaluate our solution, we had set up a physical environment of IPv6 WLAN for testing mobile communications. The testing results show that our new solution is stable and qualified for mobile communications running in WLAN. Deguang Le, Donghui Guo, Gerard P. Parr |
PIMRC | 2 |
| 1999 | A New Symmetric Probabilistic Encryption Scheme Based on Chaotic Attractors of Neural Networks
Donghui Guo, Lee-Ming Cheng, L. L. Cheng |
Appl. Intell. | 1 |