Guowei Zhu

dblp:127/4823 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Computer networks · 7 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 An Agile Deployment System for Password Recovery on FPGA
abstract
Hardware-based acceleration of password recovery remains a pressing challenge, as CPUs and GPUs struggle to efficiently process modern cryptographic primitives. Although FPGAs offer superior performance-per-watt, their widespread adoption is limited by long development cycles, manual optimization, and the absence of an end-to-end deployment framework that jointly accelerates password generation and verification. To address this gap, we propose the first agile, end-to-end FPGA-based password recovery system that unifies deep-learning-driven password generation and cryptographic verification within a single deployment workflow. The framework consists of: (1) a customized Neural Processing Unit (NPU) that accelerates GAN-based password generation models such as PassGAN; (2) an automated, template-based accelerator generator for verification kernels, built on reusable Chisel hardware primitives; and (3) a multi-objective Design Space Exploration (DSE) engine that co-optimizes kernel-level parameters (e.g., loop unrolling) and system-level parallelism to determine globally optimal FPGA configurations. We deploy the system on a heterogeneous platform combining a Zynq MPSoC with dual Virtex UltraScale+ FPGAs. Experimental results show that the NPU outperforms an NVIDIA Tesla V100 by 82.16% in PassGAN inference throughput. The full system achieves 1.90× higher end-to-end throughput and 2.32× better energy efficiency than GPU-based implementations, and delivers an average 32.58% speedup over state-of-the-art FPGA-only verification designs. These results demonstrate the practicality and scalability of our architecture for real-world password recovery workflows.
Liming Deng, Guowei Zhu, Xitian Fan, Guangwei Xie, Mingqian Sun, Xuegong Zhou, Wei Cao 0002, Fan Zhang 0044, Xinsheng Yu 0001
IEEE Trans. Computers2
2025 Agile Design Flow for Cryptographic Hardware Accelerators
abstract
This paper presents an agile design flow for cryptographic hardware accelerators, which supports the automated generation of RTL code for efficient hardware accelerators intended for FPGA deployment. An automated design method for hardware accelerators based on domain-specific hardware design templates (DST) is proposed. Based on the characteristics of cryptographic algorithms, we designed a parameterized DST and constructed a corresponding specific hardware operator library (SHWOL). We adopted a novel operator matching strategy based on subgraph isomorphism to realize the mapping of user input algorithms in the operator library, thereby deriving the DST parameters for code generation thus finalizing the RTL code generation with design space exploration (DSE). When implemented on an FPGA, compared with the existing high-level synthesis (HLS) tools, the code generated by our proposed design flow has an LUT efficiency (throughput/number of LUTs) of up to 483× and energy efficiency (throughput/power) of up to 676×.
Liming Deng, Guowei Zhu, Wei Cao 0002, Xitian Fan, Xuegong Zhou
ICCD2
2025 AHCA: Agile Design Framework for Hashcat Acceleration Based on FPGA
abstract
This article presents AHCA, an agile design framework for Field Programmable Gate Array (FPGA)-based Hashcat acceleration that automates the generation of optimized register transfer level (RTL) code. Our approach is centered on a proposed automated design method using a parameterized domain-specific template (DST) and a specific hardware operator library. The framework analyzes an algorithm’s graph to extract key hardware operators and their interconnection network. To support diverse user inputs, we introduce an innovative operator matching strategy using subgraph isomorphism, which maps algorithms to our operator library. This matched information, combined with design space exploration (DSE), is used to configure the DST and generate the final RTL code, avoiding redundancy for previously implemented algorithms. Compared to state-of-the-art high-level synthesis (HLS) tools, AHCA demonstrates a maximum performance enhancement of 797×, a Look-Up table (LUT) efficiency improvement of up to 105×, and an energy efficiency gain of up to 676×. When deployed on an FPGA for password cracking, the AHCA-generated hardware achieves a 63.95× enhancement in energy efficiency over CPUs and a 4.71× improvement over GPUs.
Liming Deng, Guowei Zhu, Xitian Fan, Wei Cao 0002, Xuegong Zhou, Fan Zhang 0044, Shaobo Yang
ACM Trans. Reconfigurable Technol. Syst.2
2025 DVHetero: A Framework for Designing and Validating Heterogeneous SoC with RISC-V Processor and CGRA
abstract
CGRA, as a coprocessor in SoCs, has been widely studied. However, there is limited research on how to efficiently debug and verify SoCs composed of CGRAs and processors during the design process. To address this gap, we introduce DVHetero. DVHetero incorporates a simulation and validation framework, SoCDiff, which enables comprehensive SoC simulation, debugging, and rapid error localization. Using this verification framework, we successfully implemented and validated the entire SoC. The SoC includes a Chisel-based CGRA generator and provides a pipelined CGRA architecture template. The CGRA is tightly integrated with the RISC-V processor, allowing for efficient DMA-based data transfer and MMIO support within the SoC. The pipelined CGRA architecture generated by DVHetero shows a 1.27× improvement in area efficiency and a 10.54× increase in mapping speed compared to the state-of-the-art CGRA framework, HierCGRA. Additionally, compared to state-of-the-art CGRA-SoC systems FDRA, DVHetero demonstrates a 1.67× increase in execution speed and a 4.34× improvement in area efficiency.
Guowei Zhu, Liming Deng, Kaisen Zhang, Wang Fan, Boyin Jin, Wei Cao 0002, Fengzhe Zhang, Xuegong Zhou, Fan Zhang 0044, Xinsheng Yu 0001
ACM Trans. Reconfigurable Technol. Syst.1
2024 An End-to-End Agile Design Framework to Improve Energy Efficiency on CGRAs
abstract
In this paper, we propose a domain-specific frame-work that integrates Chisel-based Coarse-grained reconfigurable architecture (CGRA) Modeling, RTL generation, Architecture Graph Intermediate Representation (IR), dataflow graph (DFG) Mapping, interconnect exploration, and physical implementation. Within this framework, we propose an interconnect exploration flow based on a novel interconnect architecture called Matrix, which realizes a heterogeneous interconnect architecture for a set of specific applications by application mapping, design space exploration (DSE) and pruning. We design an agile mapper built on a graph-based two-level architectural IR, which better adapts to the flexible interconnect model and enables greater interconnect exploration to improve the mapping success rate. Experiments show a significant reduction in architecture area, improved energy efficiency, and high PE utilization compared to the state-of-the-art tool, with high-quality mapping results due to architecture tuning. Additionally, our pruning strategies reduce interconnect paths in the Matrix, ensuring interconnect efficiency and further improving PE utilization.
Yazhou Yan, Guowei Zhu, Wenbo Yin, Lingli Wang
ASAP3
2024 Urban Taxi Dispatching Optimization: A Reinforcement Learning-Based Approach
abstract
With the expansion of urban scale and the increase motor vehicles, traffic congestion is becoming more serious. The development of public transportation is the main way to alleviate congestion. As part of public transportation, taxis are popular because of their all-weather, comfort and high accessibility, but currently they mainly rely on driver experience to find passengers, resulting in imbalance between supply and demand, affecting passenger experience. In this paper, we propose a taxi dispatch model using reinforcement learning method, which utilizes the Dueling DQN algorithm to optimize the dispatch strategy to alleviate the imbalance between supply and demand. Using the GPS data of taxis in Shenzhen, after data cleaning and analysis, the ARIMA model is used to predict the future boarding and alighting volume, and a supply and demand matrix is constructed. Experiments show that the dispatch efficiency of Dueling DQN can reach 100%, and the average dispatch efficiency is 85.3%.
Guowei Zhu
MSN3
2024 A Spatial-Temporal Neural Network for Short-Term Traffic Flow Prediction
abstract
To assist intelligent traffic management, traffic flow prediction, which plays a crucial role in intelligent transportation system, involves forecasting future traffic flow based on road characteristics and historical traffic data. Due to the inherent complexity of traffic systems, achieving high accuracy in long-term traffic flow prediction poses significant challenges. Therefore, we propose a novel neural network model, which is able to capture both temporal and spatial dependencies in the traffic flow data using the combination of GCN layer and LSTM layer. Experiments have demonstrated that our model is effective and accurate when used to predict the short-term traffic flow.
Shuaiyu Wang, Chang Yan, Buliao Jia, Guowei Zhu
MSN5
2024 Anomaly Detection for Coal-fired Boiler via Adversarial VAE Based Method
abstract
As the most important part of thermal power plant, a sudden shut down of a coal-fired boiler may cause a huge loss of the operating revenue. Thus, the coal-fired boiler needs to be monitored timely and comprehensively to enhance its stability and reduce shut down times. Today, it has become the norm rather than the exception to use learning-based anomaly detection methods to identify faults in a real time. The various components of a coal-fired boiler are typically monitored and generating massive multivariate time series. However, due to the complex patterns and little useful labeled data, it is a great challenge to detect anomalies from these time series data. Conventional methods generally require a large number of labeled data, but cannot guarantee high accuracy because training sets contain sparse labeled anomaly samples in most cases. In this paper, we propose an anomaly detection framework using adversarial variational auto-encoder (VAE). We first use two discriminators to adversarially train an auto encoder to learn the normal pattern of multivariate time series, and then use the reconstruction error to detect anomalies. Furthermore, in order to assist operators better diagnose anomalies, we provide anomaly interpretation based on reconstruction error of univariate time series. We con-duct experiments on the dataset collected from a coal-fired boiler. The results show that our method significantly outperforms other state-of-the-art methods.
Guowei Zhu, Hexiao Zhou, Qingchun Wang
MSN1
2024 HierCGRA: A Novel Framework for Large-scale CGRA with Hierarchical Modeling and Automated Design Space Exploration
abstract
Coarse-grained reconfigurable arrays (CGRAs) are promising design choices in computation-intensive domains, since they can strike a balance between energy efficiency and flexibility. A typical CGRA comprises processing elements (PEs) that can execute operations in applications and interconnections between them. Nevertheless, most CGRAs suffer from the ineffectiveness of supporting flexible architecture design and solving large-scale mapping problems. To address these challenges, we introduce HierCGRA, a novel framework that integrates hierarchical CGRA modeling, Chisel-based Verilog generation, LLVM-based data flow graph (DFG) generation, DFG mapping, and design space exploration (DSE). With the graph homomorphism (GH) mapping algorithm, HierCGRA achieves a faster mapping speed and higher PE utilization rate compared with the existing state-of-the-art CGRA frameworks. The proposed hierarchical mapping strategy achieves 41× speedup on average compared with the ILP mapping algorithm in CGRA-ME. Furthermore, the automated DSE based on Bayesian optimization achieves a significant performance improvement by the heterogeneity of PEs and interconnections. With these features, HierCGRA enables the agile development for large-scale CGRA and accelerates the process of finding a better CGRA architecture.
Sichao Chen, Su Zheng, Guowei Zhu, Jingyuan Li 0003, Yazhou Yan, Yuan Dai, Wenbo Yin, Lingli Wang
ACM Trans. Reconfigurable Technol. Syst.5
2023 THRAM: A Template-based Heterogeneous CGRA Modeling Framework Supporting Fast DSE
abstract
Coarse-grained reconfigurable architecture (CGRA), composed of word-level processing elements (PEs) and interconnects, has emerged as a promising architecture due to its high performance, energy efficiency, and flexibility. Although multiple CGRA frameworks have been proposed, a complete heterogeneous CGRA exploration framework with tunable interconnect flexibility and fast design space exploration (DSE) is still lacking. In this paper, we propose an open-source template-based CGRA exploration framework that integrates the modeling of heterogeneous PEs and interconnects, RTL generation, DFG mapping, automatic simulation and verification, and fast DSE based on a CGRA framework TRAM. Moreover, we present a novel resource-efficient shared reconfigurable delay unit (RDU) for data synchronization, which can save the CGRA area by 7%, compared with the separated RDU. Further, the explored optimal heterogeneous architecture can reduce the area and power by 44.7% and 42.9% respectively, and improve the PE utilization by 20.4%, compared with the 8 × 8 baseline architecture in TRAM.
Jingyuan Li 0003, Yunhui Qiu, Guowei Zhu, Qilong Zhu, Wenbo Yin, Lingli Wang
ISCAS3
2022 User Mapping Strategy in Multi-CDN Streaming: A Data-Driven Approach
abstract
Using content delivery networks (CDNs) for video distribution has become the normalemde factoapproach for video streaming today, because they are easy to use (e.g., video chunks can be delivered as files via HTTP) and have good scalability. Today, it has become the norm rather than the exception for video providers to hire multiple CDNs for their video services in a pay-per-use manner—not only to serve users at different locations, but also to reduce operational costs. Given that multiple CDNs and their peering servers exist at many different locations, selecting different CDNs for different users in the same online video system has become a critical decision that can significantly affect the users’ Quality of Experience (QoE). Conventional strategies are generally rule based, e.g., assigning users to CDNs according to their locations or ISPs, but cannot guarantee any particular QoE level because QoE is affected by a combination of complicated factors. In this article, we propose using a data-driven approach to study the factors determining users’ QoE, including both Quality of Service (QoS) and user factors. Our findings indicate that QoE is affected by both the QoS provided by the CDNs and user preferences for the video content. We design a machine learning-based predictive model to capture the “utility” of a video for a user given a particular QoS guarantee. Based on that, we strategically assign CDNs to users to maximize the overall QoE. One month trace-driven experiments are used to demonstrate the effectiveness and efficiency of our design.
Guowei Zhu, Weixi Gu
IEEE Internet Things J.1
2020 APPLE: a new compression scheme for bitmap indexes: poster abstract
abstract
Compressed bitmap indexes are increasingly used in databases and search engines. By exploiting bit-level parallelism and bitwise operations, e.g. AND/OR operations, they can significantly accelerate the development of many areas. The Word Aligned Hybrid (WAH) bitmap compression scheme using run-length encoding (RLE), is commonly recognized as the most efficient scheme in terms of CPU-performance. This paper presents a new form of compressed bitmap indexes named Adaptive Partitioned Position List Encoding (APPLE), which uses packed position lists for compression. For experiments, we compare it with Huffman encoding, and two enhanced variants of WAH : Concise and COMPAX. Our empirical results show this scheme achieves significant improvement.
Ge Ma, Guowei Zhu, Kan Lv, Qiyang Huang, Weixi Gu
SenSys3
2019 Analysis of active faults based on natural earthquakes in Central north China
Guowei Zhu, Qingchao Zhang, Zhenqiang Yang
J. Vis. Commun. Image Represent.2
2018 Understanding Gaming Experience in Mobile Multiplayer Online Battle Arena Games
abstract
Online mobile game (OMG) is booming recently, which has driven user expectations for high-quality game service. Therefore, it is crucial for game operators to understand if and how system factors (i.e. network quality metrics such as delay and mobile phone's rendering performance such as frame rate) affect gaming experience and how to optimize resource allocation to improve it. This paper is a first step towards addressing these problems. Despite the rich literature on multimedia services and Quality of Experience (QoE) measurement, the understanding of gaming experience is limited because the game platform shifts from traditional PC end to the mobile end, which has yet to be explored in depth. Based on a large-scale dataset collected from Honour of Kings, the world's top grossing mobile game, we carry out elaborate studies to explore gaming experience from the aspects of user behavior and game quality. Our key findings are as follows. First, user behavior is mainly limited to the game logic itself, such as I of a game is restricted by game rules. Therefore, it cannot well represent gaming experience. Second, among all system factors, the biggest impact on game quality comes from network performance rather than the frame rate of device, which is different from the influence mode in PC games. Third, some context factors, such as AP/BaseStations used by mobile phones, can have an indirect but huge impact on gaming experience. Based on the above observations, we provide insights that can enhance OMG's system from the perspective of resource allocation and real-time gaming experience monitoring. To the best of our knowledge, we are the first to conduct large-scale measurement to study gaming experience of OMGs in the wild. We believe our study is not only crucial for the understanding of gaming experience, but also helpful for game operators to optimize their systems.
Chou Mo, Guowei Zhu, Zhi Wang 0001, Wenwu Zhu 0001
NOSSDAV2
2016 User Mapping Strategies in Multi-Cloud Streaming: A Data-Driven Approach
abstract
Using content delivery networks (CDNs) for video distribution has become a de facto approach for today's video streaming, due to the easy usage and good scalability. Today, it has become a norm rather than an exception for video providers to hire multiple cloud CDNs for their video services in a pay-per-use manner, to not only serve users at different locations, but also reduce the operation costs. Given the multiple CDNs and their peering servers at many different locations, mapping a user to an edge CDN server has become a critical decision that can affect the quality of experience (QoE) of users. Conventional user mapping strategies are generally rule-based, e.g., assigning users to CDN servers according to only their locations or ISPs, which cannot guarantee any QoE. In this paper, we first propose to use a data-driven approach to study factors determining the streaming QoE in the multi-cloud CDN paradigm. Our findings suggest that the streaming QoE is affected by a combination of not only network factors but also user factors including their preference of video content. Then, we design a machine learning based predictive model to capture the QoE given the network conditions and user preference. Finally, we formulate the user mapping problem as an optimization problem and design algorithms to solve it: our algorithms identify users whose QoE are mostly affected by QoS and assign users to CDN servers so that the overall QoE can be maximized. Trace-driven experiments further verify the effectiveness of our design.
Guowei Zhu, Chou Mo, Zhi Wang 0001, Wenwu Zhu 0001
GLOBECOM1