VLDB 2026 Research / reviewers in the wild / expert
Jie Lu 0009
dblp:39/2936-9
· DBLP profile ↗
29ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0002-4162-0404ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 21 · 4 first-author · 18 since 2021Security and privacy · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LifeFuzz: Lifecycle-Guided Fuzzing for Windows Driver Cross-Handler VulnerabilitiesabstractThird-party Windows drivers expose a critical attack surface. However, vulnerabilities that require cross-handler I/O Control (IOCTL) sequences remain hard to find, despite their prevalence, and often lead to privilege escalation. Static analysis suffers from high false positives, path explosion, and complex resource modeling. Meanwhile, dynamic fuzzers often exercise handlers in isolation or combine them randomly, leaving implicit state dependencies unchecked. To address the gap, we present LifeFuzz, a lifecycle-guided fuzzing framework that models global-variable lifecycles to construct dependency-respecting IOCTL sequences. Specifically, it identifies variable operations across handlers, preserves seeds that affect driver state, and then combines them into meaningful sequences. Consequently, LifeFuzz explores deep paths unreachable for existing fuzzers. We evaluate LifeFuzz on 26 Windows WDM drivers. It discovers 86 vulnerabilities, including 32 cross-handler cases, with six assigned CVE IDs. Moreover, it finds 357% more cross-handler vulnerabilities than msFuzz and achieves 19.2% higher average coverage. Overall, 37% of discovered vulnerabilities require cross-handler interactions, thereby validating lifecycle-aware, cross-handler fuzzing for driver security. Chendong Yu, Yuekang Li, Yang Xiao 0011, Jie Lu 0009, Yeting Li, Defang Bo, Wei Huo 0005 |
EuroSys | 4 |
| 2026 | Context-Free Language Reachability via Efficient Relation ChainingabstractContext-free language (CFL) reachability is a fundamental framework widely used to model a variety of program analysis tasks, though it often suffers from inherent inefficiency due to its (sub)cubic time complexity. In this paper, we propose a novel perspective, relation chaining, which interprets CFL-reachability solving as the process of chaining labeled edges representing binary relations. This formulation exposes substantial derivation redundancy (in terms of frequent and repetitive chaining operations) arising from inefficient chaining strategies employed by existing approaches. To address this, we introduce Squid , a new algorithm that incorporates two simple yet effective chaining techniques—adaptive chaining and differential chaining—built upon an enhanced graph representation. We have implemented Squid as a standalone tool and evaluated it against two state-of-the-art CFL-reachability solvers and a leading Datalog solver across three key program analyses: field-sensitive alias analysis and context-sensitive value-flow analysis for C/C++, and field-sensitive points-to analysis for Java. Experimental results show that Squid substantially improves the scalability of CFL-reachability solving by effectively reducing a large portion of redundant chaining operations. Chenghang Shi, Haofeng Li, Jie Lu 0009, Lian Li 0002 |
Proc. ACM Program. Lang. | 3 |
| 2025 | Reviving Discarded Vulnerabilities: Exploiting Previously Unexploitable Linux Kernel Bugs Through Control Metadata Fields
Jian Liu 0008, Jie Lu 0009, Shaomin Chen, Tianshuo Han, Xiaorui Gong |
CCS | 3 |
| 2025 | Module-Aware Context Sensitive Pointer AnalysisabstractThe Java Platform Module System (JPMS) has found widespread applications since introduced in Java 9. However, existing pointer analyses fail to leverage the semantics of JPMS. This paper presents a novel module-aware approach to improving the performance of pointer analysis. We model the semantics of keywords provides and uses in JPMS to recover missing points-to relations. We design a module-aware context-sensitive analysis, which can propagate and apply critical contexts (by exploiting modularity) to balance precision and efficiency better. We have implemented our module-aware pointer analysis named MPA in TAI - E and conducted extensive experiments to compare it with standard object-sensitivity. The evaluation results demonstrate that MPA finds more reachable methods and enhances existing context-sensitive approaches, striking a good balance between efficiency and precision. MPA can increase the number of reachable methods up to 90.9× (lombok) under the same analysis. Performance-wise, MPA is nearly as fast as context-insensitivity for most benchmarks, while its precision is superior to that of 1-object-sensitivity on average. Haofeng Li, Chenghang Shi, Jie Lu 0009, Lian Li 0002 |
ICSE | 3 |
| 2025 | SLVHound: Static Detection of Session Lingering Vulnerabilities in Modern Java Web ApplicationsabstractSession Lingering Vulnerability (SLV) is an often overlooked authentication flaw that allows sessions to persist after authenticationsensitive operations.Despite its widespread occurrence and severe impact, SLVs have received little attention.To address this gap, we present the first comprehensive study of SLV in Web applications, introducing a novel detection tool called SLVHound.Our approach employs static analysis of both code and SQL queries to identify authentication-sensitive operations and session expiration.SLVHound then detects SLVs by verifying whether authenticationsensitive operations are consistently followed by session expiration.We evaluated SLVHound on 15 popular Web applications, uncovering 46 potential vulnerabilities.Further analysis confirmed 44 of them as true SLVs, including 30 previously unreported vulnerabilities, with 16 CVE IDs granted. CCS Concepts• Security and privacy → Web application security. Haining Meng, Jie Lu 0009, Yongheng Huang, Lian Li 0002 |
Internetware | 2 |
| 2025 | Understanding Resource Injection Vulnerabilities in Kubernetes EcosystemsabstractCloud-native technologies have revolutionized application development, with Kubernetes emerging as the de facto standard platform for containerization and orchestration. Kubernetes manages applications through API objects called resources, where users declare desired states via resource definitions that are processed by controllers to reconcile system discrepancies. However, this resource-based architecture introduces resource injection vulnerabilities, where controllers perform privileged operations using user-controllable fields without adequate validation. Attackers can exploit these weaknesses by injecting malicious content into resource fields to achieve unauthorized access and privilege escalation.In this paper, we conduct the first comprehensive study on 125 resource injection vulnerabilities from 8,306 Kubernetes-related vulnerabilities across common databases. For all studied vulnerabilities, we investigate their vulnerable fields, root causes, privileged operations, exploitation conditions, and fixing strategies. Our study reveals many interesting findings that can guide the detection and mitigation of resource injection vulnerabilities, as well as the development of more secure cloud-native applications. Defang Bo, Jie Lu 0009, Feng Li 0045, Jingting Chen, Jinchen Wang, Chendong Yu, Yeting Li, Wei Huo 0005 |
ASE | 2 |
| 2025 | ZIPPER: Static Taint Analysis for PHP Applications with Precision and Efficiency
Xinyi Wang 0013, Yeting Li, Jie Lu 0009, Shizhe Cui, Chenghang Shi, Qin Mai, Yunpei Zhang, Yang Xiao 0011, Feng Li 0045, Wei Huo 0005 |
USENIX Security Symposium | 3 |
| 2025 | Fast Client-Driven CFL-Reachability via Regularization-Based Graph SimplificationabstractContext-free language (CFL) reachability is a critical framework for various program analyses, widely adopted despite its computational challenges due to cubic or near-cubic time complexity. This often leads to significant performance degradation in client applications. Notably, in real-world scenarios, clients typically require reachability information only for specific source-to-sink pairs, offering opportunities for targeted optimization. We introduce MoYe, an effective regularization-based graph simplification technique designed to enhance the performance of client-driven CFL-reachability analyses by pruning non-contributing edges—those that do not participate in any specified CFL-reachable paths. MoYe employs a regular approximation to ensure exact reachability results for all designated node pairs and operates linearly with respect to the number of edges in the graph. This lightweight efficiency makes MoYe a valuable pre-processing step that substantially reduces both computational time and memory requirements for CFL-reachability analysis, outperforming a recent leading graph simplification approach. Our evaluations with two prominent CFL-reachability client applications demonstrate that MoYe can substantially improve performance and reduce resource consumption. Chenghang Shi, Dongjie He, Haofeng Li, Jie Lu 0009, Lian Li 0002, Jingling Xue |
Proc. ACM Program. Lang. | 4 |
| 2024 | Detecting Broken Object-Level Authorization Vulnerabilities in Database-Backed ApplicationsabstractBroken object-level authorization (BOLA) vulnerabilities are among the most critical security risks facing database-backed applications. However, there is still a significant gap in our systematic understanding of these vulnerabilities. To bridge this gap, we conducted an in-depth study of 101 real-world BOLA vulnerabilities from opensource applications. Our study revealed the four most common object-level authorization models in database-backed application. Yongheng Huang, Chenghang Shi, Jie Lu 0009, Haofeng Li, Haining Meng, Lian Li 0002 |
CCS | 3 |
| 2024 | Boosting the Performance of Multi-Solver IFDS Algorithms with Flow-Sensitivity OptimizationsabstractThe IFDS (Inter-procedural, Finite, Distributive, Subset) algorithms are popularly used to solve a wide range of analysis problems. In particular, many interesting problems are formulated as multi-solver IFDS problems which expect multiple interleaved IFDS solvers to work together. For instance, taint analysis requires two IFDS solvers, one forward solver to propagate tainted data-flow facts, and one backward solver to solve alias relations at the same time. For such problems, large amount of additional data-flow facts need to be introduced for flow-sensitivity. This often leads to poor performance and scalability, as evident in our experiments and previous work. In this paper, we propose a novel approach to reduce the number of introduced additional data-flow facts while preserving flow-sensitivity and soundness. We have developed a new taint analysis tool, SADROID, and evaluated it on 1,228 open-source Android APPs. Evaluation results show that SADROID significantly outperforms FLowDROID (the state-of-the-art multi-solver IFDS taint analysis tool) without affecting precision and soundness: the run time performance is sped up by up to 17.89X and memory usage is optimized by up to 9X. Haofeng Li, Jie Lu 0009, Haining Meng, Liqing Cao, Lian Li 0002, Lin Gao 0002 |
CGO | 2 |
| 2024 | AutoWeb: Automatically Inferring Web Framework Semantics via Configuration Mutation
Haining Meng, Haofeng Li, Jie Lu 0009, Chenghang Shi, Liqing Cao, Lian Li 0002, Lin Gao 0002 |
ICECCS | 3 |
| 2024 | Better Not Together: Staged Solving for Context-Free Language ReachabilityabstractContext-free language reachability (CFL-reachability) is a fundamental formulation for program analysis with many applications. CFL-reachability analysis is computationally expensive, with a slightly subcubic time complexity concerning the number of nodes in the input graph. This paper proposes staged solving: a new perspective on solving CFL-reachability. Our key observation is that the context-free grammar (CFG) of a CFL-based program analysis can be decomposed into (1) a smaller CFG, L, for matching parentheses, such as procedure calls/returns, field stores/loads, and (2) a regular grammar, R, capturing control/data flows. Instead of solving these two parts monolithically (as in standard algorithms), staged solving solves L-reachability and R-reachability in two distinct stages. In practice, L-reachability, though still context-free, involves only a small subset of edges, while R-reachability can be computed efficiently with close to quadratic complexity relative to the node size of the input graph. We implement our staged CFL-reachability solver, STG, and evaluate it using two clients: context-sensitive value-flow analysis and field-sensitive alias analysis. The empirical results demonstrate that STG achieves speedups of 861.59x and 4.1x for value-flow analysis and alias analysis on average, respectively, over the standard subcubic algorithm. Moreover, we also showcase that staged solving can help to significantly improve the performance of two state-of-the-art solvers, POCR and PEARL, by 74.82x (1.78x) and 37.66x (1.7x) for value-flow (alias) analysis, respectively. Chenghang Shi, Haofeng Li, Jie Lu 0009, Lian Li 0002 |
ISSTA | 3 |
| 2024 | File Hijacking Vulnerability: The Elephant in the Room
Chendong Yu, Yang Xiao 0011, Jie Lu 0009, Yuekang Li, Yeting Li, Lian Li 0002, Jian Wang 0067, Defang Bo, Wei Huo 0005 |
NDSS | 3 |
| 2024 | An anomaly aware network embedding framework for unsupervised anomalous link detection
Dongsheng Duan, Lingling Tong, Jie Lu 0009, Cunchi Lv, Yangxi Li |
Data Min. Knowl. Discov. | 4 |
| 2024 | Boosting the Performance of Alias-Aware IFDS Analysis with CFL-Based Environment TransformersabstractThe IFDS algorithm is pivotal in solving field-sensitive data-flow problems. However, its conventional use of access paths for field sensitivity leads to the generation of a large number of data-flow facts. This causes scalability challenges in larger programs, limiting its practical application in extensive codebases. In response, we propose a new field-sensitive technique that reinterprets the generation of access paths as a Context-Free Language (CFL) for field-sensitivity and formulates it as an IDE problem. This approach significantly reduces the number of data-flow facts generated and handled during the analysis, which is a major factor in performance degradation. To demonstrate the effectiveness of this approach, we developed a taint analysis tool, IDEDroid, in the IFDS/IDE framework. IDEDroid outperforms FlowDroid, an established IFDS-based taint analysis tool, in the analysis of 24 major Android apps while improving its precision (guaranteed theoretically). The speed improvement ranges from 2.1 × to 2,368.4 × , averaging at 222.0 × , with precision gains reaching up to 20.0 % (in terms of false positives reduced). This performance indicates that IDEDroid is substantially more effective in detecting information-flow leaks, making it a potentially superior tool for mobile app vetting in the market. Haofeng Li, Chenghang Shi, Jie Lu 0009, Lian Li 0002, Jingling Xue |
Proc. ACM Program. Lang. | 3 |
| 2024 | Generic Sensitivity: Generics-Guided Context Sensitivity for Pointer AnalysisabstractGeneric programming has found widespread application in object-oriented languages like Java. However, existing context-sensitive pointer analyses fail to leverage the benefits of generic programming. This paper introducesgeneric sensitivity, a new context customization scheme targeting generics. We design our context customization scheme in such a way that generic instantiation sites, i.e., locations instantiating generic classes/methods with concrete types, are always preserved as key context elements. This is realized by augmenting contexts with a type variable lookup map, which is efficiently generated in a context-sensitive manner throughout the analysis process. We have implemented various variants of generic-sensitive analysis in WALA and conducted extensive experiments to compare it with state-of-the-art approaches, including both traditional and selective context-sensitivity methods. The evaluation results demonstrate that generic sensitivity effectively enhances existing context-sensitivity approaches, striking a new balance between efficiency and precision. For instance, it enables a 1-object-sensitive analysis to achieve overall better precision compared to a 2-object-sensitive analysis, with an average speedup of 12.6 times (up to 62 times). Haofeng Li, Tian Tan 0001, Yue Li 0006, Jie Lu 0009, Haining Meng, Liqing Cao, Yongheng Huang, Lian Li 0002, Lin Gao 0002, Peng Di, ChenXi Cui |
IEEE Trans. Software Eng. | 4 |
| 2024 | Pearl: A Multi-Derivation Approach to Efficient CFL-Reachability SolvingabstractContext-free language (CFL) reachability is a fundamental framework for formulating program analyses. CFL-reachability analysis works on top of an edge-labeled graph by deriving reachability relations and adding them as labeled edges to the graph. Existing CFL-reachability algorithms typically adopt a single-reachability relation derivation (SRD) strategy, i.e., one reachability relation is derived at a time. Unfortunately, this strategy can lead to redundancy, hindering the efficiency of the analysis. To address this problem, this paper proposesPearl, amulti-derivationapproach that reduces derivation redundancy for CFL-reachability solving, which significantly improves the efficiency of CFL-reachability analysis. Our key insight is that multiple edges can be simultaneously derived via batch propagation of reachability relations. We also tailor our multi-derivation approach to tackle transitive relations that frequently arise when solving CFL-reachability. Specifically, we present a highly efficient transitive-aware variant,PearlPG, which enhancesPearlwithpropagation graphs, a lightweight but effective graph representation, to further diminish redundant derivations. We evaluate the performance of our approach on two clients, i.e., context-sensitive value-flow analysis and field-sensitive alias analysis for C/C++. By eliminating a large amount of redundancy, our approach outperforms two baselines including the standard CFL-reachability algorithm and a state-of-the-art solverPocrspecialized for fast transitivity solving. In particular, the empirical results demonstrate that, for value-flow analysis and alias analysis respectively,PearlPGruns 3.09$\times$faster on average (up to 4.44$\times$) and 2.25$\times$faster on average (up to 3.31$\times$) thanPocr, while also consuming less memory. Chenghang Shi, Haofeng Li, Yulei Sui, Jie Lu 0009, Lian Li 0002, Jingling Xue |
IEEE Trans. Software Eng. | 4 |
| 2023 | Two Birds with One Stone: Multi-Derivation for Fast Context-Free Language Reachability AnalysisabstractContext-free language (CFL) reachability is a fundamental framework for formulating program analyses. CFL-reachability analysis works on top of an edge-labeled graph by deriving reachability relations and adding them as labeled edges to the graph. Existing CFL-reachability algorithms typically adopt a single-reachability relation derivation (SRD) strategy, i.e., one reachability relation is derived at a time. Unfortunately, this strategy can lead to redundancy, hindering the efficiency of the analysis. To address this problem, this paper proposes Pearl, a multi-derivation approach that reduces derivation redundancy for transitive relations that frequently arise when solving reachability relations, significantly improving the efficiency of CFL-reachability analysis. Our key insight is that multiple edges involving transitivity can be simultaneously derived via batch propagation of reachability relations on the transitivity-aware subgraphs that are induced from the original edge-labeled graph. We evaluate the performance of Pearl on two clients, i.e., context-sensitive value-flow analysis and field-sensitive alias analysis for C/C++. By eliminating a large amount of redundancy, Pearl achieves average speedups of 82.73x for value-flow analysis and 155.26x for alias analysis over the standard CFL-reachability algorithm. The comparison with Pocr, a state-of-the-art CFL-reachability solver, shows that Pearl runs 10.1x (up to 29.2x) and 2.37x (up to 4.22x) faster on average respectively for value-flow analysis and alias analysis with less consumed memory. Chenghang Shi, Haofeng Li, Yulei Sui, Jie Lu 0009, Lian Li 0002, Jingling Xue |
ASE | 4 |
| 2022 | Detecting Missing-Permission-Check Vulnerabilities in Distributed Cloud SystemsabstractMissing- Permission-Check (MPC) vulnerability is a type of bug where permission checks are not enforced for privileged operations. MPC vulnerability is prevalent and can cause severe security impacts. This paper proposes the first tool to detect MPC vulnerabilities in distributed cloud systems. We conduct an in-depth study of 95 real-world MPC vulnerabilities and our findings motivate a new tool named MPChecker. The tool introduces a combined log-static analysis to automatically identify privileged operations by inferring variables representing user owned data and critical system states, whose accesses need to be protected. We have evaluated MPChecker with 6 popular distributed systems. The tool reports 44 new vulnerabilities, and 43 of them have been confirmed and labeled as critical bugs. Moreover, 1 bug is particular dangerous and the developers requested to keep it undisclosed. Jie Lu 0009, Haofeng Li, Lian Li 0002 |
CCS | 1 |
| 2022 | Generic sensitivity: customizing context-sensitive pointer analysis for genericsabstractGeneric programming has been extensively used in object-oriented programs such as Java. However, existing context-sensitive pointer analyses perform poorly in analyzing generics. This paper introduces generic sensitivity, a new context customization scheme targeting generics. We design our context customization scheme in such a way that generic instantiation sites, i.e., locations instantiating generic classes/methods with concrete types, are always preserved as key context elements. This is realized by augmenting contexts with a type variable lookup map, which is efficiently updated during the analysis in a context-sensitive manner. Haofeng Li, Jie Lu 0009, Haining Meng, Liqing Cao, Yongheng Huang, Lian Li 0002, Lin Gao 0002 |
ESEC/SIGSOFT FSE | 2 |
| 2022 | CloudRaid: Detecting Distributed Concurrency Bugs via Log Mining and EnhancementabstractCloud systems suffer from distributed concurrency bugs, which often lead to data loss and service outage. This paper presentsCloudRaid, a new automatical tool for finding distributed concurrency bugs efficiently and effectively. Distributed concurrency bugs are notoriously difficult to find as they are triggered by untimely interaction among nodes, i.e., unexpected message orderings. To detect concurrency bugs in cloud systems efficiently and effectively,CloudRaidanalyzes and tests automatically only the message orderings that are likely to expose errors. Specifically,CloudRaidmines the logs from previous executions to uncover the message orderings that are feasible but inadequately tested. In addition, we also propose a log enhancing technique to introduce new logs automatically in the system being tested. These extra logs added improve further the effectiveness ofCloudRaidwithout introducing any noticeable performance overhead. Our log-based approach makes it well-suited for live systems. We have appliedCloudRaidto analyze six representative distributed systems: Hadoop2/Yarn, HBase, HDFS, Cassandra, Zookeeper, and Flink.CloudRaidhas succeeded in testing 60 different versions of these six systems (10 versions per system) in 35 hours, uncovering 31 concurrency bugs, including nine new bugs that have never been reported before. For these nine new bugs detected, which have all been confirmed by their original developers, three are critical and have already been fixed. Jie Lu 0009, Feng Li 0045, Lian Li 0002, Xiaobing Feng 0002, Jingling Xue |
IEEE Trans. Software Eng. | 1 |
| 2021 | Exposing Vulnerable Paths: Enhance Static Analysis with Lightweight Symbolic ExecutionabstractStatic analysis tools, although widely adopted in industry, suffer from a high false positive rate. This paper aims to refine the results of static analysis tools, by automatically searching for a vulnerable path from given defect report. To realize this goal, we develop SATRACER, a novel tool which integrates symbolic execution techniques with static analysis. SATRACER selectively skips those program parts which can be consistently updated by static analysis, thus drastically improving performance. We have applied SATRACER to a set of 21 real-world applications. Evaluation results show that SATRACER can successfully remove 71.4% false alarms reported by a commercial static analysis tool in 10 hours, and confirmed 29 real use-after-free bugs and 895 real null-pointer-dereference bugs. Guangwei Li, Jie Lu 0009, Lian Li 0002, Xu Song |
APSEC | 3 |
| 2021 | Scaling Up the IFDS Algorithm with Efficient Disk-Assisted ComputingabstractThe IFDS algorithm can be memory-intensive, requiring a memory budget of more than 100 GB of RAM for some applications. The large memory requirements significantly restrict the deployment of IFDS-based tools in practise. To improve this, we propose a disk-assisted solution that drastically reduces the memory requirements of traditional IFDS solvers. Our solution saves memory by 1) recomputing instead of memorizing intermediate analysis data, and 2) swapping in-memory data to disk when memory usages reach a threshold. We implement sophisticated scheduling schemes to swap data between memory and disks efficiently. We have developed a new taint analysis tool, DiskDroid, based on our disk-assisted IFDS solver. Compared to FlowDroid, a state-of-the-art IFDS-based taint analysis tool, for a set of 19 apps which take from 10 to 128 GB of RAM by FlowDroid, DiskDroid can analyze them with less than 10GB of RAM at a slight performance improvement of 8.6%. In addition, for 21 apps requiring more than 128GB of RAM by FlowDroid, DiskDroid can analyze each app in 3 hours, under the same memory budget of 10GB. This makes the tool deployable to normal desktop environments. We make the tool publicly available at https://github.com/HaofLi/DiskDroid. Haofeng Li, Haining Meng, Hengjie Zheng, Liqing Cao, Jie Lu 0009, Lian Li 0002, Lin Gao 0002 |
CGO | 5 |
| 2021 | GoBench: A Benchmark Suite of Real-World Go Concurrency BugsabstractGo, a fast growing programming language, is often considered as “the programming language of the cloud”. The language provides a rich set of synchronization primitives, making it easy to write concurrent programs with great parallelism. However. the rich set of primitives also introduces many bugs. We build Gobench, the first benchmark suite for Go concurrency bugs. Currently, Gobench consists of 82 real bugs from 9 popular open source applications and 103 bug kernels. The bug kernels are carefully extracted and simplified from 67 out of these 82 bugs and 36 additional bugs reported in a recent study to preserve their bug-inducing complexities as much as possible. These bugs cover a variety of concurrency issues, both traditional and Go-specific. We believe Gobench will be instrumental in helping researchers understand concurrency bugs in Go and develop effective tools for their detection. We have therefore evaluated a range of representative concurrency error detection tools using Gobench. Our evaluation has revealed their limitations and provided insights for making further improvements. Guangwei Li, Jie Lu 0009, Lian Li 0002, Jingling Xue |
CGO | 3 |
| 2021 | Detecting TensorFlow Program Bugs in Real-World Industrial EnvironmentabstractDeep learning has been widely adopted in industry and has achieved great success in a wide range of application areas. Bugs in deep learning programs can cause catastrophic failures, in addition to a serious waste of resources and time.This paper aims at detecting industrial TensorFlow program bugs. We report an extensive empirical study on 12,289 failed TensorFlow jobs, showing that existing static tools can effectively detect 72.55% of the top three types of Python bugs in industrial TensorFlow programs. In addition, we propose (for the first time) a constraint-based approach for detecting TensorFlow shape-related errors (one of the most common TensorFlow-specific bugs), together with an associated tool, ShapeTracer. Our evaluation on a set of 60 industrial TensorFlow programs shows that ShapeTracer is efficient and effective: it analyzes each program in at most 3 seconds and detects effectively 40 out of 60 industrial TensorFlow program bugs, with no false positives. ShapeTracer has been deployed in the platform-X platform and will be released soon. Jie Lu 0009, Guangwei Li, Lian Li 0002, Liang You, Jingling Xue |
ASE | 2 |
| 2020 | AANE: Anomaly Aware Network Embedding For Anomalous Link DetectionabstractExisting network embedding models regard all the links in a network as normal and model them without distinction. In real networks, there may be anomalous links like noise or adversarial links. We explicitly consider the existence of anomalous links in a network and propose anomaly aware network embedding (AANE) model. The key of AANE is the design of a new loss, which consists of anomaly aware loss and adjusted fitting loss. We adopt an anomaly indicator to iteratively select significant anomalous links from the network during model training, and removal loss and deviation loss are designed to model the reconstruction errors of selected anomalous and normal links respectively. To instantiate AANE, AAGAE and AAGCN are implemented on graph auto-encoder (GAE) and graph convolution based auto-encoder (GCNAE) respectively. For the purpose of evaluation, a heuristic anomalous link generation algorithm is proposed and by using the algorithm we generate anomalous links into six real world network datasets. Experimental results show that AANE outperforms both basic and competitive network embedding models in terms of anomalous link detection performance in most cases. Dongsheng Duan, Lingling Tong, Yangxi Li, Jie Lu 0009 |
ICDM | 4 |
| 2019 | CrashTuner: detecting crash-recovery bugs in cloud systems via meta-info analysisabstractCrash-recovery bugs (bugs in crash-recovery-related mechanisms) are among the most severe bugs in cloud systems and can easily cause system failures. It is notoriously difficult to detect crash-recovery bugs since these bugs can only be exposed when nodes crash under special timing conditions. This paper presents CrashTuner, a novel fault-injection testing approach to combat crash-recovery bugs. The novelty of CrashTuner lies in how we identify fault-injection points (crash points) that are likely to expose errors. We observe that if a node crashes while accessing meta-info variables, i.e., variables referencing high-level system state information (e.g., an instance of node or task), it often triggers crash-recovery bugs. Hence, we identify crash points by automatically inferring meta-info variables via a log-based static program analysis. Our approach is automatic and no manual specification is required. Jie Lu 0009, Lian Li 0002, Xiaobing Feng 0002, Liang You |
SOSP | 1 |
| 2019 | Understanding Node Change Bugs for Distributed SystemsabstractDistributed systems are the fundamental infrastructure for modern cloud applications and the reliability of these systems directly impacts service availability. Distributed systems run on clusters of nodes. When the system is running, nodes can join or leave the cluster at anytime, due to unexpected failure or system maintenance. It is essential for distributed systems to tolerate such node changes. However, it is also notoriously difficult and challenging to handle node changes right. There are widely existing node change bugs which can lead to catastrophic failures. We believe that a comprehensive study on node change bugs is necessary to better prevent and diagnose node change bugs. In this paper, we perform an extensive empirical study on node change bugs. We manually went through 6,660 bug issues of 5 representative distributed systems, where 620 issues were identified as node change bugs. We studied 120 bug examples in detail to understand the root causes, the impacts, the trigger conditions and fixing strategies of node change bugs. Our findings shed lights on new detection and diagnosis techniques for node change bugs. In our empirical study, we develop two useful tools, NCTrigger and NPEDetector. NCTrigger helps users to automatically reproduce a node change bug by injecting node change events based on user specification. It largely reduces the manual efforts to reproduce a bug (from 2 days to less than half a day). NPEDetector is a static analysis tool to detect null pointer exception errors. We develop this tool based on our findings that node operations often lead to null pointer exception errors, and these errors share a simple common pattern. Experimental results show that this tool can detect 60 new null pointer errors, including 7 node change bugs. 23 bugs have already been patched and fixed. Jie Lu 0009, Liu Chen, Lian Li 0002, Xiaobing Feng 0002 |
SANER | 1 |
| 2018 | CloudRaid: hunting concurrency bugs in the cloud via log-miningabstractCloud systems suffer from distributed concurrency bugs, which are notoriously difficult to detect and often lead to data loss and service outage. This paper presents CloudRaid, a new effective tool to battle distributed concurrency bugs. CloudRaid automatically detects concurrency bugs in cloud systems, by analyzing and testing those message orderings that are likely to expose errors. We observe that large-scale online cloud applications process millions of user requests per second, exercising many permutations of message orderings extensively. Those already sufficiently-tested message orderings are unlikely to expose errors. Hence, CloudRaid mines logs from previous executions to uncover those message orderings which are feasible, but not sufficiently tested. Specifically, CloudRaid tries to flip the order of a pair of messages if they may happen in parallel, but S always arrives before P from existing logs, i.e., excercising the order P ↣ S. The log-based approach makes it suitable to live systems. Jie Lu 0009, Feng Li 0045, Lian Li 0002, Xiaobing Feng 0002 |
ESEC/SIGSOFT FSE | 1 |