VLDB 2026 Research / reviewers in the wild / expert
Husheng Zhou
dblp:160/7471
· DBLP profile ↗
11ranked-venue papers
5as first author
3since 2021 · last 2021
0000-0003-0614-4980ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
GPUs and heterogeneous computing · 48% Embedded and real-time systems · 43% Energy-efficient computing · 9% | |
| Network and information security
2 papers |
Systems and software security · 50% Security and privacy of machine learning · 50% | |
| Software engineering, system software, and programming languages
2 papers |
Software testing · 100% | |
| Artificial intelligence
2 papers |
Autonomous driving · 51% Trustworthy machine learning · 31% Deep learning architectures and training · 18% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Embedded and real-time systems
real-time scheduling |
0.5 | 2 | 2018 | PredJoule: A Timing-Predictable Energy Optimization Framework for Deep Neural Networks · RTSS 2018 PASS: priority assignment of real-time tasks with dynamic suspending behavior under fixed-priority scheduling · DAC 2015 |
Systems and software security › memory safety
memory isolation |
0.5 | 1 | 2021 | gGuard: Enabling Leakage-Resilient Memory Isolation in GPU-accelerated Autonomous Embedded Systems · DAC 2021 |
Security and privacy of machine learning
model stealing |
0.5 | 1 | 2021 | Hermes Attack: Steal DNN Models with Lossless Inference Accuracy · USENIX Security Symposium 2021 |
GPUs and heterogeneous computing
GPU memory management |
0.5 | 1 | 2021 | gGuard: Enabling Leakage-Resilient Memory Isolation in GPU-accelerated Autonomous Embedded Systems · DAC 2021 |
Robotics › Autonomous driving
perception |
0.4 | 1 | 2020 | DeepBillboard: systematic physical-world testing of autonomous driving systems · ICSE 2020 |
Software testing
deep learning testing |
0.4 | 1 | 2020 | DeepBillboard: systematic physical-world testing of autonomous driving systems · ICSE 2020 |
Software testing
fuzzing |
0.4 | 1 | 2020 | Simulee: detecting CUDA synchronization bugs via memory-access modeling · ICSE 2020 |
GPUs and heterogeneous computing › GPU programming
CUDA |
0.4 | 1 | 2020 | Simulee: detecting CUDA synchronization bugs via memory-access modeling · ICSE 2020 |
Embedded and real-time systems › real-time scheduling
fixed-priority scheduling |
0.2 | 1 | 2015 | PASS: priority assignment of real-time tasks with dynamic suspending behavior under fixed-priority scheduling · DAC 2015 |
Machine learning › Deep learning architectures and training › neural network inference
DNN inference |
0.1 | 1 | 2021 | Hermes Attack: Steal DNN Models with Lossless Inference Accuracy · USENIX Security Symposium 2021 |
Embedded and real-time systems
automotive embedded systems |
0.1 | 1 | 2021 | gGuard: Enabling Leakage-Resilient Memory Isolation in GPU-accelerated Autonomous Embedded Systems · DAC 2021 |
GPUs and heterogeneous computing › GPU computing
GPU-accelerated systems |
0.1 | 1 | 2021 | gGuard: Enabling Leakage-Resilient Memory Isolation in GPU-accelerated Autonomous Embedded Systems · DAC 2021 |
Machine learning › Trustworthy machine learning › robustness
adversarial attack |
0.1 | 1 | 2020 | DeepBillboard: systematic physical-world testing of autonomous driving systems · ICSE 2020 |
Machine learning › Trustworthy machine learning
robustness |
0.1 | 1 | 2020 | DeepBillboard: systematic physical-world testing of autonomous driving systems · ICSE 2020 |
Energy-efficient computing › energy-efficient architecture
GPU energy optimization |
0.1 | 1 | 2018 | PredJoule: A Timing-Predictable Energy Optimization Framework for Deep Neural Networks · RTSS 2018 |
Energy-efficient computing
power management |
0.1 | 1 | 2018 | PredJoule: A Timing-Predictable Energy Optimization Framework for Deep Neural Networks · RTSS 2018 |
Methods — techniques the papers use, named apart from their topics
compiler-level data shredding · 1.0application-aware data shredding · 1.0OS-level data shredding · 1.0physical-world testing · 0.9memory-access modeling · 0.9adversarial perturbation · 0.9LLVM bytecode interpretation · 0.9layer-aware energy optimization · 0.3simulation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | gGuard: Enabling Leakage-Resilient Memory Isolation in GPU-accelerated Autonomous Embedded SystemsabstractGraphics processing units (GPUs) are being widely used as co-processors for performance acceleration in many autonomous embedded systems such as robotics and autonomous vehicles. However, current GPU hardware and systems software, including GPU device drivers, compilers, and operating systems, do not implement proper memory protection mechanisms due to performance and proprietary reasons, causing severe vulnerabilities such as information leakage. In this paper, we present gGuard, a leakage-resilient GPU memory management system with strong isolation. Based on the intrinsic characteristics of information leakage vulnerabilities on GPUs, gGuard develops a set of efficient and accurate data shredding techniques implemented at the compiler, library, and operating system levels, with the core idea of exploring the data access patterns and dependencies for efficient application-aware data shredding. Our implementation and evaluation show that gGuard can provide effective mitigation on GPU data leakage issues through efficient GPU data shredding while introducing less than 6% overhead in all tested scenarios. Yaswanth Yadlapalli, Husheng Zhou, Yuqun Zhang, Cong Liu 0005 |
DAC | 2 |
| 2021 | Hermes Attack: Steal DNN Models with Lossless Inference Accuracy
Yuankun Zhu, Yueqiang Cheng, Husheng Zhou, Yantao Lu |
USENIX Security Symposium | 3 |
| 2021 | Efficient algorithms for task mapping on heterogeneous CPU/GPU platforms for fast completion time
Zexin Li 0001, Yuqun Zhang, Husheng Zhou, Cong Liu 0005 |
J. Syst. Archit. | 4 |
| 2020 | Simulee: detecting CUDA synchronization bugs via memory-access modelingabstractWhile CUDA has become a mainstream parallel computing platform and programming model for general-purpose GPU computing, how to effectively and efficiently detect CUDA synchronization bugs remains a challenging open problem. In this paper, we propose the first lightweight CUDA synchronization bug detection framework, namely Simulee, to model CUDA program execution by interpreting the corresponding LLVM bytecode and collecting the memory-access information for automatically detecting general CUDA synchronization bugs. To evaluate the effectiveness and efficiency of Simulee, we construct a benchmark with 7 popular CUDA-related projects from GitHub, upon which we conduct an extensive set of experiments. The experimental results suggest that Simulee can detect 21 out of the 24 manually identified bugs in our preliminary study and also 24 previously unknown bugs among all projects, 10 of which have already been confirmed by the developers. Furthermore, Simulee significantly outperforms state-of-the-art approaches for CUDA synchronization bug detection. Mingyuan Wu, Yicheng Ouyang, Husheng Zhou, Lingming Zhang 0001, Cong Liu 0005, Yuqun Zhang |
ICSE | 3 |
| 2020 | DeepBillboard: systematic physical-world testing of autonomous driving systemsabstractDeep Neural Networks (DNNs) have been widely applied in autonomous systems such as self-driving vehicles. Recently, DNN testing has been intensively studied to automatically generate adversarial examples, which inject small-magnitude perturbations into inputs to test DNNs under extreme situations. While existing testing techniques prove to be effective, particularly for autonomous driving, they mostly focus on generating digital adversarial perturbations, e.g., changing image pixels, which may never happen in the physical world. Thus, there is a critical missing piece in the literature on autonomous driving testing: understanding and exploiting both digital and physical adversarial perturbation generation for impacting steering decisions. In this paper, we propose a systematic physical-world testing approach, namely DeepBillboard, targeting at a quite common and practical driving scenario: drive-by billboards. DeepBillboard is capable of generating a robust and resilient printable adversarial billboard test, which works under dynamic changing driving conditions including viewing angle, distance, and lighting. The objective is to maximize the possibility, degree, and duration of the steering-angle errors of an autonomous vehicle driving by our generated adversarial billboard. We have extensively evaluated the efficacy and robustness of DeepBillboard by conducting both experiments with digital perturbations and physical-world case studies. The digital experimental results show that DeepBillboard is effective for various steering models and scenes. Furthermore, the physical case studies demonstrate that DeepBillboard is sufficiently robust and resilient for generating physical-world adversarial billboard tests for real-world driving under various weather conditions, being able to mislead the average steering angle error up to 26.44 degrees. To the best of our knowledge, this is the first study demonstrating the possibility of generating realistic and continuous physical-world tests for practical autonomous driving systems; moreover, DeepBillboard can be directly generalized to a variety of other physical entities/surfaces along the curbside, e.g., a graffiti painted on a wall. Husheng Zhou, Wei Li 0159, Zelun Kong, Yuqun Zhang, Bei Yu 0001, Lingming Zhang 0001, Cong Liu 0005 |
ICSE | 1 |
| 2018 | GRU: Exploring Computation and Data Redundancy via Partial GPU Computing Result ReuseabstractGraphics processing units (GPUs) have been widely adopted by major cloud vendors for better performance and energy efficiency. Recent research has observed a considerable degree of redundancy in managing computation and data in many datacenters, particularly for several important categories of GPU-accelerated applications such as log mining and machine learning. In this paper, we present GRU, an ecosystem that smartly manages and shares GPU resources through exploiting redundancy. GRU transparently interprets GPU-accelerated computing requests and memoizes results for potential future reuse. To enhance reusability, GRU implements a partial result reuse idea, where GPU computation requests even with different input data and functionality may become reusable w.r.t. each other. To guarantee correctness of partial reuse, GRU employs a compiler-assisted approach that analyzes general data parallel patterns that are reliable for the reuse purpose, and is capable of smartly recognizing such reusable data parallel patterns of incoming requests. We have fully implemented GRU and conducted extensive sets of experiments running micro-benchmarks on local machines and real-world applications including Spark-based uses cases in an AWS cluster. Evaluation results show that GRU is effective in identifying and eliminating redundant GPU computations, achieving up to 5x (2.5x) speedup for compute-intensive (data-intensive) benchmarks. In addition, GPU-managed Spark observes a reduction of 25.3% (39.8%) on average w.r.t. turnaround time (GPU occupation time) over state-of-the-art solutions. Husheng Zhou, Soroush Bateni, Cong Liu 0005 |
ICS | 1 |
| 2018 | S^3DNN: Supervised Streaming and Scheduling for GPU-Accelerated Real-Time DNN WorkloadsabstractDeep Neural Networks (DNNs) are being widely applied in many advanced embedded systems that require autonomous decision making, e.g., autonomous driving and robotics. To handle resource-demanding DNN workloads, graphic processing units (GPUs) have been used as the main acceleration engine. Although much research has been conducted to algorithmically optimize the efficiency of applying DNN to applications such as object recognition, limited attention has been given to optimizing the execution of GPU-accelerated DNN workloads at the system level. In this paper, we propose S^3DNN, a system solution that optimizes the execution of DNN workloads on GPU in a real-time multi-tasking environment, which simultaneously optimizes the two (sometimes) conflicting goals of real-time correctness and throughput. S^3DNN contains a governor that selectively gathers system-wide DNN requests to perform smart data fusion, and a novel supervised streaming and scheduling framework that combines a deadline-aware scheduler with the concurrency-enabled CUDA stream technique. To simultaneously maximize concurrency-induced benefits and real-time performance, S^3DNN explores a rather interesting and unique characteristic of DNN workloads, where multiple layers of a DNN instance often exhibit a gradually decreased GPU resource utilization pattern. We have fully implemented S^3DNN in a GPU-accelerated system and have conducted extensive sets of experiments evaluating the efficacy of S^3DNN under a wide range of system and workload scenarios. The results show that S^3DNN significantly improves upon state-of-the-art GPU-accelerated DNN processing frameworks, e.g., up to 37% and over 40% improvements in real-time performance and throughput, respectively. Husheng Zhou, Soroush Bateni, Cong Liu 0005 |
RTAS | 1 |
| 2018 | PredJoule: A Timing-Predictable Energy Optimization Framework for Deep Neural NetworksabstractThe revolution of deep neural networks (DNNs) is enabling dramatically better autonomy in autonomous driving. However, it is not straightforward to simultaneously achieve both timing predictability (i.e., meeting job latency requirements) and energy efficiency that are essential for any DNN-based autonomous driving system, as they represent two (often) conflicting goals. In this paper, we propose PredJoule, a timing-predictable energy optimization framework for running DNN workloads in a GPU-enabled automotive system. PredJoule achieves both latency guarantees and energy efficiency through a layer-aware design that explores specific performance and energy characteristics of different layers within the same neural network. We implement and evaluate PredJoule on the automotive-specific NVIDIA Jetson TX2 platform for five state-of-the-art DNN models with both high and low variance latency requirements. Experiments show that PredJoule rarely violates job deadlines, and can improve energy by 65% on average compared to five existing approaches and 68% compared to an energy-oriented approach. Soroush Bateni, Husheng Zhou, Yuankun Zhu, Cong Liu 0005 |
RTSS | 2 |
| 2015 | PASS: priority assignment of real-time tasks with dynamic suspending behavior under fixed-priority schedulingabstractSelf-suspension is becoming an increasingly prominent characteristic in real-time systems such as: (i) I/O-intensive systems, where applications interact intensively with I/O devices, (ii) multi-core processors, where tasks running on different cores have to synchronize and communicate with each other, and (iii) computation offloading systems with coprocessors, like Graphics Processing Units (GPUs). In this paper, we show that rate-monotonic (RM), deadline-monotonic (DM) and laxity-monotonic (LM) scheduling will perform rather poor in dynamic self-suspending systems in terms of speed-up factors. On the other hand, the proposed PASS approach is guaranteed to find a feasible priority assignment on a speed-2 uniprocessor, if one exists on a unit-speed processor. We evaluate the feasibility of the proposed approach via a case study implementation. Furthermore, the effectiveness of the proposed approach is also shown via extensive simulation results. Wen-Hung Kevin Huang, Jian-Jia Chen, Husheng Zhou, Cong Liu 0005 |
DAC | 3 |
| 2015 | GPES: a preemptive execution system for GPGPU computingabstractGraphics processing units (GPUs) are being widely used as co-processors in many application domains to accelerate general-purpose workloads that are computationally intensive, known as GPGPU computing. Real-time multi-tasking support is a critical requirement for many emerging GPGPU computing domains. However, due to the asynchronous and non-preemptive nature of GPU processing, in multi-tasking environments, tasks with higher priority may be blocked by lower priority tasks for a lengthy duration. This severely harms the system's timing predictability and is a serious impediment limiting the applicability of GPGPU in many real-time and embedded systems. In this paper, we present an efficient GPGPU preemptive execution system (GPES), which combines user-level and driverlevel runtime engines to reduce the pending time of high-priority GPGPU tasks that may be blocked by long-freezing low-priority competing workloads. GPES automatically slices a long-running kernel execution into multiple subkernel launches and splits data transaction into multiple chunks at user-level, then inserts preemption points between subkernel launches and memorycopy operations at driver-level. We implement a prototype of GPES, and use real-world benchmarks and case studies for evaluation. Experimental results demonstrate that GPES is able to reduce the pending time of high-priority tasks in a multitasking environment by up to 90% over the existing GPU driver solutions, while introducing small overheads. Husheng Zhou, Guangmo Tong, Cong Liu 0005 |
RTAS | 1 |
| 2014 | Task mapping in heterogeneous embedded systems for fast completion timeabstractGraphics processing units are being widely used in embedded systems as they can achieve high performance and energy efficiency. In such systems, the problem of computation and data mapping for multiple applications while minimizing the completion time is quite challenging due to a large size of the policy space, including heterogeneous application characteristics, complex application structure, data communication costs, and data partitioning. To achieve fast competition time, a fine-grain mapping framework that explores a set of critical factors is needed for heterogeneous embedded systems. In this paper, we consider this mapping problem by presenting a theoretical framework that yields an optimal integer programming solution. Moreover, based upon several interesting measurements-based case studies, we design three practical mapping algorithms with low time complexity, each of which explores a specific set of factors that may affect the completion time performance. We evaluated the proposed algorithms by implementing them on a real heterogeneous system and using a large set of popular benchmarks for evaluation. Experimental results demonstrate that our proposed algorithms can achieve up to 30% faster completion time compared to the state-of-the-art mapping techniques, and can perform consistently well across different workloads. Husheng Zhou, Cong Liu 0005 |
EMSOFT | 1 |