Xu Zhou 0004

dblp:66/5686-4 · DBLP profile ↗
← Back
38ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0002-0075-5003ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 12 · 1 first-author · 9 since 2021Systems, architecture and hardware · 9 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Not All Paths Are Equal: Multi-path Optimization for Directed Hybrid Fuzzing
abstract
Directed Grey-Box Fuzzing (DGF) can improve bug exposure efficiency by stressing bug-prone areas. Recent studies have modeled DGF as the problem of finding and optimizing paths to reach target sites. However, they still face the “ multi-path ” challenge. When a target site is reachable by multiple paths, it is crucial to comprehensively evaluate and effectively select these paths, as this affects the fuzzer’s choice between reaching target sites via optimal paths and enhancing path diversity toward targets to expose hidden bugs in non-optimal paths. In this article, we propose MultiGo, a directed hybrid fuzzer designed for multi-path optimization. First, we propose a new fitness metric called path difficulty to comprehensively evaluate the promising paths. This metric uses the Poisson distribution to estimate the probability of exploring basic blocks along execution paths based on statistical block frequency, distinguishing between optimal and challenging paths. With path difficulty as a key factor, a customized Contextual Multi-Armed Bandit (CMAB) model is employed to efficiently optimize path scheduling by comprehensively considering the impact of testing conditions on path scheduling. We introduce the concept of the fuzzing context to represent and evaluate testing conditions, which encompass factors such as path characteristics (e.g., path difficulty), the testing agent (e.g., fuzzing or symbolic execution), and the testing goal (e.g., path exploitation or exploration). Then, the CMAB model predicts the expected rewards for scheduling paths under different testing agents and goals, thereby optimizing path scheduling. By leveraging the CMAB model, MultiGo enhances DGF’s capability to explore easier paths and symbolic execution’s capacity to handle more complex ones, enabling efficient target reaching through optimal paths while ensuring sufficient coverage of non-optimal paths. MultiGo is evaluated on 136 target sites of 41 real-world programs from 3 benchmarks. The experimental results show that MultiGo outperforms the state-of-the-art directed fuzzers (AFLGo, SelectFuzz, Beacon, WindRanger, and DAFL) and hybrid fuzzers (SymCC and SymGo) in reaching target sites and exposing known vulnerabilities. Moreover, MultiGo also discovered 14 undisclosed vulnerabilities.
Peihong Lin, Pengfei Wang 0010, Xu Zhou 0004, Wei Xie 0007, Gen Zhang, Kai Lu 0001
ACM Trans. Softw. Eng. Methodol.3
2025 Learning from the Packet Sequences: Diffusion Model-Based Protocol Greybox Fuzzing
abstract
Network protocol fuzzing is crucial for ensuring the security and stability of protocols. Traditional specification-based methods face severe challenges to generate test cases in scenarios where protocols are complex and specifications are unavailable. Furthermore, the statefulness of protocol implementations requires input packets to satisfy sequential dependencies, further complicates fuzzing. To address these challenges, we developed SPIREFuzz, a novel fuzzer that automatically learns protocol formats and temporal relationships directly from real traffic. SPIREFuzz uses reverse engineering to build a session pattern dataset and employs a Discrete Denoising Diffusion Probabilistic Model (D3PM) to generate packet sequences. Leveraging the model's strengths in stable training and mitigating mode collapse, the method achieves an optimal balance between the accuracy and diversity of generated sequences, which is crucial for generating high-fidelity and state-compliant patterns and simultaneously ensures enhanced state space exploration. Experimental results indicate that on 10 protocols SPIREFuzz's key field identification capability is superior to NetPlier. In fuzzing targeting 8 protocol implementations, compared to AFLNet, NSFuzz, and GANFuzz, its average state coverage, average state transitions, average bitmap coverage, and average unique crashes achieved improvements of up to$\text{78.57 \%}, \text{73.91 \%}, \text{3.94 \%}$, and 216.67 % respectively.
Peihong Lin, Xu Zhou 0004, Wei Xie 0007
IPCCC4
2025 SimFuzz: Conflict-Aware Parallel Fuzzing via Incremental Path Similarity Clustering
abstract
Parallel fuzzing boosts throughput by distributing testcase generation across multiple fuzzing instances. However, this architecture often suffers from task conflict—redundant exploration of similar execution paths—due to the lack of path-level awareness in seed scheduling. These conflicts waste computation and limit overall effectiveness.We present SIMFUZZ, a conflict-aware scheduling framework that mitigates redundancy by integrating path similarity into the fuzzing workflow. SIMFUZZ encodes seeds as branch-level coverage bitmaps and incrementally clusters them based on execution path overlap. It then applies a two-stage scheduling policy that assigns similar seeds to the same instance, while preserving global prioritization for high-potential inputs.We evaluate SIMFUZZ on 19 real-world programs and benchmark targets. Compared to a state-of-the-art baseline, it achieves a 6.7% average increase in branch coverage and reduces task conflict by 3.9%. In several cases, it also discovers substantially more unique crashes. Additionally, SIMFUZZ has uncovered 15 previously unknown vulnerabilities in widely used software projects, all of which have been assigned CVE identifiers.
Xuan Meng, Danjun Liu, Xu Zhou 0004, Peihong Lin, Chenyifan Liu, Lei Zhou 0023, Wei Xie 0007
ISSRE3
2025 When Control Flows Deviate: Directed Grey-box Fuzzing with Probabilistic Reachability Analysis
abstract
Directed grey-box fuzzing (DGF) steers testing toward high-value targets, but developing effective DGF for commercial off-the-shelf (COTS) binaries is challenging due to the lack of accurate structural information (e.g., control-flow graphs and call graphs), which can cause control flows to deviate and misguide DGF’s reachability analysis. In this paper, we introduce BinGo, a tailored binary-level directed grey-box fuzzer, which can accommodate the flawed control-flow graphs (CFGs) of COTS binaries and enable accurate and efficient reachability analysis. First, to quantify the inevitable inaccuracies of uncovered indirect edges and analyze their impact on the reachability of basic blocks, we propose a Bayesian-based method. This method combines prior knowledge from static analysis with dynamic observations from fuzzing to estimate the confidence in correctly recovering indirect edges. Then, we present a new concept called a region, which redefines granularity for efficient reachability analysis by transforming the CFG into a region graph. Using the Bayesian results and region graph, we propose a custom fitness metric for binary-level DGF, termed probabilistic reachability. This metric, based on a dynamically updated region graph and reachability scores, is adaptive, lightweight, and accommodates inaccurate binary-level CFGs. We implemented a prototype tool, BinGo, and evaluated it on the CGC dataset, CVE-Benchmark, and UniBench benchmark. Experimental results show that BinGo surpasses baseline fuzzers (AFL++, AFLGo, PDGF, UAFuzz, and 1dVul) in reaching target locations and exposing known vulnerabilities. Additionally, BinGo discovered three new vulnerabilities in the real-world application cscope-15.9.
Peihong Lin, Xu Zhou 0004, Wei Xie 0007, Kai Lu 0001
ASE3
2025 Constructing arbitrary write via puppet objects and delivering gadgets in Linux kernel
Danjun Liu, Xuan Meng, Pengfei Wang 0010, Xu Zhou 0004, Wei Xie 0007
Comput. Secur.4
2025 Efficient Forward-Edge Control-Flow Integrity for COTS Binaries via Arm BTI
abstract
Control-Flow Integrity (CFI) has been widely recognized as an effective technique for mitigating control-flow hijacking attacks. However, many binary-level CFI approaches suffer from weaknesses in safeguarding forward edges, particularly for the obfuscated binaries, due to the imprecision in binary analysis or heuristic algorithms. Moreover, these approaches often involve non-negligible overhead and are challenging to deploy, as they instrument plenty of code or employ hardware tracing to enforce the CFI policies. This paper introduces Mobius, the first complete implementation of security-instruction-based binary-only CFI solution on commercial processors. Mobius leverages the Branch Target Identification (BTI) technology in Arm v8.5 to safeguard the forward edges of binaries and shared libraries efficiently. It determines the forward-edge targets without false negatives and carefully instruments the bti instructions to conduct the CFI checking efficiently. Then, it mounts a runtime monitor to detect potential attacks. We deploy Mobius on an Alibaba Cloud server with Yitian 710 processors in practice without modifying the kernel or loader. Remarkably, Mobius successfully provides efficient protection for real-world applications, including obfuscated code, with marginal overhead (5.78% on SPEC2006).
Tai Yue, Kai Lu 0001, Zhenyu Ning, Pengfei Wang 0010, Lei Zhou 0023, Xu Zhou 0004, Fengwei Zhang, Gen Zhang
IEEE Trans. Inf. Forensics Secur.6
2024 Speed is Not All You Need When Fuzzing Stateful Network Servers
abstract
Current network protocol fuzzing is an efficient mechanism for uncovering protocol vulnerabilities, yet it faces several challenges. For instance, bugs in stateful protocols can significantly hinder the generation of effective fuzzing testcases. Additionally, factors such as network transmission and session synchronization can lead to substantial delays in the fuzzing process. However, we have observed a common phenomenon that existing fuzzing techniques often focus on optimizing either execution speed or state inference, but not both, which limits their overall effectiveness in analyzing protocol defects. We verify such a common issue by experimental analysis and seek to understand the specific reasons behind this. Then, we design an enhanced network protocol fuzzer that harnesses shared memory-based message transmission and session state synchronization to alleviate the heavy post-execution analysis in StateAFL state inference, named S2fuzzer, to optimize both speed and state inference simultaneously. Our experiments reveal that S2fuzzer enhances execution speed by 4.7x compared to StateAFL. Simultaneously, we find that S2Fuzzer can infer a more comprehensive state model. This leads to a 9.38% increase in code coverage and 371 additional crashes (1.42x) compared to StateAFL across tested programs. This state inference amplification, which has not been discovered in previous research, allows S2fuzzer to reach nearly equivalent coverage and superior bug finding ability, even if its execution speed is merely 1/10 of HNPFuzzer, the latest speedup scheme.
Lei Zhou 0023, Xu Zhou 0004, Danjun Liu
HPCC3
2024 DeepGo: Predictive Directed Greybox Fuzzing
Peihong Lin, Pengfei Wang 0010, Xu Zhou 0004, Wei Xie 0007, Gen Zhang, Kai Lu 0001
NDSS3
2024 Efficiently Rebuilding Coverage in Hardware-Assisted Greybox Fuzzing
abstract
Coverage-based greybox fuzzing (CGF) is an efficient technique for detecting vulnerabilities, but its coverage-feedback mechanism introduces significant overhead in binary-only fuzzing. Although hardware-assisted greybox fuzzing (HGF) has been proposed to address this issue, existing approaches struggle to achieve a balance between the efficiency and sensitivity of coverage, as well as to cope with trace buffer overflow.
Tai Yue, Yibo Jin 0006, Fengwei Zhang, Zhenyu Ning, Pengfei Wang 0010, Xu Zhou 0004, Kai Lu 0001
RAID6
2024 HyperGo: Probability-based directed hybrid fuzzing
Peihong Lin, Pengfei Wang 0010, Xu Zhou 0004, Wei Xie 0007, Kai Lu 0001, Gen Zhang
Comput. Secur.3
2024 Towards adaptive graph neural networks via solving prior-data conflicts
abstract
Graph neural networks (GNNs) have achieved remarkable performance in a variety of graph-related tasks. Recent evidence in the GNN community shows that such good performance can be attributed to the homophily prior; i.e., connected nodes tend to have similar features and labels. However, in heterophilic settings where the features of connected nodes may vary significantly, GNN models exhibit notable performance deterioration. In this work, we formulate this problem as prior-data conflict and propose a model called the mixture-prior graph neural network (MPGNN). First, to address the mismatch of homophily prior on heterophilic graphs, we introduce the non-informative prior, which makes no assumptions about the relationship between connected nodes and learns such relationship from the data. Second, to avoid performance degradation on homophilic graphs, we implement a soft switch to balance the effects of homophily prior and non-informative prior by learnable weights. We evaluate the performance of MPGNN on both synthetic and real-world graphs. Results show that MPGNN can effectively capture the relationship between connected nodes, while the soft switch helps select a suitable prior according to the graph characteristics. With these two designs, MPGNN outperforms state-of-the-art methods on heterophilic graphs without sacrificing performance on homophilic graphs.
Xugang Wu, Huijun Wu 0001, Ruibo Wang, Xu Zhou 0004, Kai Lu 0001
Frontiers Inf. Technol. Electron. Eng.4
2024 The progress, challenges, and perspectives of directed greybox fuzzing
abstract
Summary Greybox fuzzing is a scalable and practical approach for software testing. Most greybox fuzzing tools are coverage‐guided as reaching high code coverage is more likely to find bugs. However, since most covered codes may not contain bugs, blindly extending code coverage is less efficient, especially for corner cases. Unlike coverage‐guided greybox fuzzing which increases code coverage in an undirected manner, directed greybox fuzzing (DGF) spends most of its time allocation on reaching specific targets (e.g. the bug‐prone zone) without wasting resources stressing unrelated parts. Thus, DGF is particularly suitable for scenarios such as patch testing, bug reproduction, and special bug detection. For now, DGF has become an active research area. However, DGF has general limitations and challenges that are worth further studying. Based on the investigation of 42 state‐of‐the‐art fuzzers that are closely related to DGF, we conducted the first in‐depth study to summarize the empirical evidence on the research progress of DGF. This paper studies DGF from a broader view, which takes into account not only the location‐directed type that targets specific code parts but also the behavior‐directed type that aims to expose abnormal program behaviors. By analyzing the benefits and limitations of DGF research, we try to identify gaps in current research, meanwhile, reveal new research opportunities and suggest areas for further investigation.
Pengfei Wang 0010, Xu Zhou 0004, Tai Yue, Peihong Lin, Kai Lu 0001
Softw. Test. Verification Reliab.2
2024 Armor: Protecting Software Against Hardware Tracing Techniques
abstract
Many modern processors have embedded hardware tracing techniques (e.g., Intel Processor Trace or ARM CoreSight). While these techniques are widely used due to their transparency and low overhead, they also bring serious security threats. Attackers can utilize hardware tracing to trace the trusted applications from a non-secure application. Existing protection techniques fail to effectively protect the runtime information when hardware tracing is employed. To counter these threats, in this paper, we propose a novel direction called anti-hardware tracing. Our key idea is to exploit the limitations of hardware tracing: trace buffer overflow can cause trace data loss. We build a model to analyse the overflow and outline three principles for efficient triggering overflows and achieving anti-hardware tracing: numerous branches in the program, high-speed execution of the program, and the high-water mark of the trace buffer. We develop a framework called Armor on ARM Juno R2 to realize our approach. Armor protects software against the trace unit Embedded Trace Macrocell (ETM) in CoreSight by instrumenting protection and loop functions. The protection function detects runtime environments, efficiently fills the trace buffer, and employs various protection strategies like PID (process identifier) replacement and PIE+STRIP+ASLR. Meanwhile, the loop function triggers overflows efficiently based on context-based calculations and anti-ETM loop. Our evaluation demonstrates that the overhead of Armor is 77.31% lower than that of OLLVM [1] on SPEC2006. Armor effectively hides 54.51% of basic blocks across 16 real-world applications, triggering 113× more overflows. Moreover, we showcase two practical applications of Armor. Firstly, we conduct a cryptographic and cross-world attack on GnuPG 1.4.13 RSA private keys using ETM, which can steal entire keys from a program in the Secure world with a single run. Armor successfully reduces leaked bits by 84.5%. Secondly, Armor impedes hardware-assisted fuzzing by reducing throughput by 89.71% and branch coverage by 47.99%.
Tai Yue, Fengwei Zhang, Zhenyu Ning, Pengfei Wang 0010, Xu Zhou 0004, Kai Lu 0001, Lei Zhou 0023
IEEE Trans. Inf. Forensics Secur.5
2023 VulHawk: Cross-architecture Vulnerability Detection with Entropy-based Binary Code Search
Zhenhao Luo, Pengfei Wang 0010, Yong Tang 0005, Wei Xie 0007, Xu Zhou 0004, Danjun Liu, Kai Lu 0001
NDSS6
2023 Leveraging Free Labels to Power up Heterophilic Graph Learning in Weakly-Supervised Settings: An Empirical Study
Xugang Wu, Huijun Wu 0001, Ruibo Wang, Duanyu Li, Xu Zhou 0004, Kai Lu 0001
ECML/PKDD (3)5
2023 From Release to Rebirth: Exploiting Thanos Objects in Linux Kernel
abstract
Vulnerability fixing is time-consuming, hence, not all of the discovered vulnerabilities can be fixed timely. In reality, developers prioritize vulnerability fixing based on exploitability. Large numbers of vulnerabilities are delayed to patch or even ignored as they are regarded as “unexploitable” or underestimated owing to the difficulty in exploiting the weak primitives. However, exploits may have been in the wild. In this paper, to exploit the weak primitives that traditional approaches fail to exploit, we propose a versatile exploitation strategy that can transform weak exploit primitives into strong exploit primitives. Based on a special object in the kernel named Thanos object, our approach can exploit a UAF vulnerability that does not have function pointer dereference and an OOB write vulnerability that has limited write length and value. Our approach overcomes the shortage that traditional exploitation strategies heavily rely on the capability of the vulnerability. To facilitate using Thanos objects, we devise a tool namedTAODEto automatically search for eligible Thanos objects from the kernel. Then, it evaluates the usability of the identified Thanos objects by the complexity of the constraints. Finally, it pairs vulnerabilities with eligible Thanos objects. We have evaluated our approach with real-world kernels.TAODEsuccessfully identified numerous Thanos objects from Linux. Using the identified Thanos objects, we proved the feasibility of our approach with 20 real-world vulnerabilities, most of which traditional techniques failed to exploit. Through the experiments, we find that in addition to exploiting weak primitives, our approach can sometimes bypass the kernel SMAP mechanism (CVE-2016-10150, CVE-2016-0728), better utilize the leaked heap pointer address (CVE-2022-25636), and even theoretically break certain vulnerability patches (e.g., double-free).
Danjun Liu, Pengfei Wang 0010, Xu Zhou 0004, Wei Xie 0007, Gen Zhang, Zhenhao Luo, Tai Yue
IEEE Trans. Inf. Forensics Secur.3
2023 UltraFuzz: Towards Resource-Saving in Distributed Fuzzing
abstract
Recent research has sought to improve fuzzing performance via parallel computing. However, researchers focus on improving efficiency while ignoring the increasing cost of testing resources. Parallel fuzzing in the distributed environment amplifies the resource-wasting problem caused by the random nature of fuzzing. In the parallel mode, owing to the lack of an appropriate task dispatching scheme and timely fuzzing status synchronization among different fuzzing instances, task conflicts and workload imbalance occur, making the resource-wasting problem severe. In this paper, we design UltraFuzz, a fuzzer for resource-saving in distributed fuzzing. Based on centralized dynamic scheduling, UltraFuzz can dispatch tasks and schedule power globally and reasonably to avoid resource-wasting. Besides, UltraFuzz can elastically allocate computing power for fuzzing and seed evaluation, thereby avoiding the potential bottleneck of seed evaluation that blocks the fuzzing process. UltraFuzz was evaluated using real-world programs, and the results show that with the same testing resource, UltraFuzz outperforms state-of-the-art tools, such as AFL, AFL-P, PAFL, and EnFuzz. Most importantly, the experiment reveals certain results that seem counter-intuitive, namely that parallel fuzzing can achieve “super-linear acceleration” when compared with single-core fuzzing. We conduct additional experiments to reveal the deep reasons behind this phenomenon and dig deep into the inherent advantages of parallel fuzzing over serial fuzzing, including the global optimization of seed energy scheduling and the escape of local optimal seed. Additionally, 24 real-world vulnerabilities were discovered using UltraFuzz.
Xu Zhou 0004, Pengfei Wang 0010, Chenyifan Liu, Tai Yue, Congxi Song, Kai Lu 0001, Qidi Yin
IEEE Trans. Software Eng.1
2022 MobFuzz: Adaptive Multi-objective Optimization in Gray-box Fuzzing
Gen Zhang, Pengfei Wang 0010, Tai Yue, Shan Huang 0002, Xu Zhou 0004, Kai Lu 0001
NDSS6
2022 Towards Defense Against Adversarial Attacks on Graph Neural Networks via Calibrated Co-Training
Xugang Wu, Huijun Wu 0001, Xu Zhou 0004, Kai Lu 0001
J. Comput. Sci. Technol.3
2022 ovAFLow: Detecting Memory Corruption Bugs with Fuzzing-Based Taint Inference
Gen Zhang, Pengfei Wang 0010, Tai Yue, Xu Zhou 0004, Kai Lu 0001
J. Comput. Sci. Technol.5
2021 MEBS: Uncovering Memory Life-Cycle Bugs in Operating System Kernels
Gen Zhang, Pengfei Wang 0010, Tai Yue, Xu Zhou 0004, Kai Lu 0001
J. Comput. Sci. Technol.4
2020 EcoFuzz: Adaptive Energy-Saving Greybox Fuzzing as a Variant of the Adversarial Multi-Armed Bandit
Tai Yue, Pengfei Wang 0010, Yong Tang 0005, Enze Wang, Bo Yu 0008, Kai Lu 0001, Xu Zhou 0004
USENIX Security Symposium7
2020 Sabotaging the system boundary: A study of the inter-boundary vulnerability
Pengfei Wang 0010, Xu Zhou 0004, Kai Lu 0001
J. Inf. Secur. Appl.2
2019 AVPredictor: Comprehensive prediction and detection of atomicity violations
abstract
Summary Concurrency bugs, such as atomicity‐violation bugs, are difficult to detect due to the uncertainty of thread‐scheduling. It is particularly difficult to conduct a thorough bug fix when an atomicity‐violation bug can be triggered by different buggy interleavings. This paper proposes a prediction‐based approach to comprehensively detect atomicity‐violation bugs. A bug fix can be incomplete when the developer cannot have all the buggy interleavings. Based on the candidate interleavings, this approach can predict unmanifested atomicity‐violation bugs from a non‐buggy execution and comprehensively display all the buggy interleavings for the same bug to assist a thorough fix. We use a monitored execution to record execution traces and predict potential buggy interleavings based on the candidate interleavings identified from the trace. Then, we use controlled executions to verify the predicted buggy interleavings by controlling the thread‐scheduling. We implemented a prototype tool called AVPredictor and evaluated it with real‐world tests. Experiments show that AVPredictor can effectively find all the known atomicity‐violation bugs as well as a previously unknown bug together with all the buggy interleavings for each bug. The runtime overhead is 13x for the monitored execution and 18x for the controlled execution.
Pengfei Wang 0010, Jens Krinke, Xu Zhou 0004, Kai Lu 0001
Concurr. Comput. Pract. Exp.3
2019 DFTracker: detecting double-fetch bugs by multi-taint parallel tracking
Pengfei Wang 0010, Kai Lu 0001, Gen Li 0002, Xu Zhou 0004
Frontiers Comput. Sci.4
2018 DFTinker: Detecting and Fixing Double-Fetch Bugs in an Automated Way
Yingqi Luo, Pengfei Wang 0010, Xu Zhou 0004, Kai Lu 0001
WASA3
2018 A survey of the double-fetch vulnerabilities
abstract
Summary Race conditions widely exist in concurrent programs, and concurrency errors caused by harmful races could lead to severe system failures. A double fetch is a typical situation when the system kernel inevitably accesses user space data multiple times, and it turns into a vulnerability when the data consistency is violated under a special race condition between kernel and user space. In this survey, we present the first (to the best of our knowledge) comprehensive study on double‐fetch vulnerabilities in the real world. Our study is based on the investigation of 91 real‐world double‐fetch vulnerabilities collected from the CVE database and other relevant reports, which covers a period of recent 12 years. Our work reveals some interesting findings on the double‐fetch vulnerabilities, ranging from the various occurrences across different kernels and system levels to the involvement of specific patterns. We also divide the consequences that are usually caused by the double‐fetch vulnerabilities into four categories and discuss each, summarize viable exploitation techniques from existing works, provide useful guidances to detect and practical strategies to prevent double‐fetch vulnerabilities.
Pengfei Wang 0010, Kai Lu 0001, Gen Li 0002, Xu Zhou 0004
Concurr. Comput. Pract. Exp.4
2018 Untrusted Hardware Causes Double-Fetch Problems in the I/O Memory
Kai Lu 0001, Pengfei Wang 0010, Gen Li 0002, Xu Zhou 0004
J. Comput. Sci. Technol.4
2017 Fine-grained checkpoint based on non-volatile memory
abstract
New non-volatile memory (e.g., phase-change memory) provides fast access, large capacity, byte-addressability, and non-volatility features. These features, fast-byte-persistency, will bring new opportunities to fault tolerance. We propose a fine-grained checkpoint based on non-volatile memory. We extend the current virtual memory manager to manage non-volatile memory, and design a persistent heap with support for fast allocation and checkpointing of persistent objects. To achieve a fine-grained checkpoint, we scatter objects across virtual pages and rely on hardware page-protection to monitor the modifications. In our system, two objects in different virtual pages may reside on the same physical page. Modifying one object would not interfere with the other object. This allows us to monitor and checkpoint objects smaller than 4096 bytes in a fine-grained way. Compared with previous page-grained based checkpoint mechanisms, our new checkpoint method can greatly reduce the data copied at checkpoint time and better leverage the limited bandwidth of non-volatile memory.
Kai Lu 0001, Mikel Luján, Xu Zhou 0004
Frontiers Inf. Technol. Electron. Eng.5
2015 RaceChecker: Efficient Identification of Harmful Data Races
abstract
Data races hidden in concurrent programs have caused severe failures. To improve the reliability, many race detectors are proposed. However, most of the reported races are not harmful, which consumes manual effort to identify the harmful races. This paper proposes RaceChecker that can detect the potential races and identify the harmful races effectively and efficiently. Unlike previous detectors, RaceChecker combines happens-before relation and ad-hoc synchronization to prune the infeasible races so that fewer potential races are required to be verified. Before verification, RaceChecker groups the remaining potential races, guaranteeing the potential races in one group do not interfere with each other. Therefore, multiple potential races in one group can be verified together in one execution. To our knowledge, this is the first effective technique that groups the potential races to improve the efficiency. Unlike previous detectors that verify one potential race in one execution, RaceChecker dynamically controls thread scheduler to create real race conditions to verify multiple potential races in one execution, identifying the harmful races that cause program failures. We have implemented RaceChecker as a prototype tool and have experimented on a number of real-world concurrent programs. Results show that 66% of the potential races are infeasible and nearly 48% of the executions are reduced by the grouping strategy. The known harmful races are also identified effectively. By pruning and grouping, RaceChecker identifies the harmful races more efficiently. Comparing with RaceMob and RaceFuzzer, the time is reduced significantly, with an average of 45% and 81% respectively.
Kai Lu 0001, Zhendong Wu, Chen Chen 0016, Xu Zhou 0004
PDP5
2015 An Efficient and Flexible Deterministic Framework for Multithreaded Programs
Kai Lu 0001, Xu Zhou 0004, Tom Bergan, Chen Chen 0016
J. Comput. Sci. Technol.2
2015 Detecting harmful data races through parallel verification
Zhendong Wu, Kai Lu 0001, Xu Zhou 0004, Chen Chen 0016
J. Supercomput.4
2014 Enhancing the Security of Parallel Programs via Reducing Scheduling Space
abstract
Parallel programs face a new security problem - concurrency vulnerability, which is caused by a special thread scheduling instead of inputs. In this paper, we propose to automatically fix concurrency vulnerabilities by reducing thread scheduling space. Our method is based on two observations. First, most concurrency vulnerabilities are caused by atomicity violation errors. Second, reducing thread scheduling space does not harm the correctness of the original program. We designed a prototype runtime system shield using deterministic multithreading techniques. Shield is designed to transparently run parallel programs and schedule threads in large instruction blocks to prevent atomicity violation at best effort. In case some concurrency vulnerabilities cannot be fixed by shield's scheduling reducing scheme, we also provide a remedy strategy by integrating shield with record&replay function, so that it can help programmers to analyze attacker's behavior for manually fixing.
Xu Zhou 0004, Gen Li 0002, Kai Lu 0001, Shuangxi Wang
DASC1
2014 Efficient deterministic multithreading without global barriers
abstract
Multithreaded programs execute nondeterministically on conventional architectures and operating systems. This complicates many tasks, including debugging and testing. Deterministic multithreading (DMT) makes the output of a multithreaded program depend on its inputs only, which can totally solve the above problem. However, current DMT implementations suffer from a common inefficiency: they use frequent global barriers to enforce a deterministic ordering on memory accesses. In this paper, we eliminate that inefficiency using an execution model we call deterministic lazy release consistency (DLRC). Our execution model uses the Kendo algorithm to enforce a deterministic ordering on synchronization, and it uses a deterministic version of the lazy release consistency memory model to propagate memory updates across threads. Our approach guarantees that programs execute deterministically even when they contain data races. We implemented a DMT system based on these ideas (RFDet) and evaluated it using 16 parallel applications. Our implementation targets C/C++ programs that use POSIX threads. Results show that RFDet gains nearly 2x speedup compared with DThreads-a start-of-the-art DMT system.
Kai Lu 0001, Xu Zhou 0004, Tom Bergan
PPoPP2
2013 Pruning False Positives of Static Data-Race Detection via Thread Specialization
Chen Chen 0016, Kai Lu 0001, Xu Zhou 0004
APPT4
2013 RaceFree: an efficient multi-threading model for determinism
abstract
Current deterministic systems generally incur large overhead due to the difficulty of detecting and eliminating data races. This paper presents RaceFree, a novel multi-threading runtime that adopts a relaxed deterministic model to provide a data-race-free environment for parallel programs. This model cuts off unnecessary shared-memory communication by isolating threads in separated memories, which eliminates direct data races. Meanwhile, we leverage the happen-before relation defined by applications themselves as one-way communication pipes to perform necessary thread communication. Shared-memory communication is transparently converted to message-passing style communication by our Memory Modification Propagation (MMP) mechanism, which propagates local memory modifications to other threads through the happen-before relation pipes. The overhead of RaceFree is 67.2% according to our tests on parallel benchmarks.
Kai Lu 0001, Xu Zhou 0004, Gen Li 0002
PPoPP2
2012 dMPI: Facilitating Debugging of MPI Programs via Deterministic Message Passing
Xu Zhou 0004, Kai Lu 0001, Xicheng Lu, Baohua Fan
NPC1
2012 Exploiting parallelism in deterministic shared memory multiprocessing
Xu Zhou 0004, Kai Lu 0001
J. Parallel Distributed Comput.1