EDBT 2026 Demo / reviewers in the wild / expert
Xiang Gao 0012
dblp:14/3881-12
· DBLP profile ↗
43ranked-venue papers
8as first author
34since 2021 · last 2026
0000-0001-9895-4600ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 35 · 7 first-author · 28 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Bug Detection by Inferring Implicit API Contract of Pointer State Transition
Xingjing Deng, Xiang Gao 0012, Hailong Sun 0001 |
DSN | 3 |
| 2026 | Fed4Fed: A Privacy-Preserving Federated Statistical Approach for Evaluating Federated Learning ModelsabstractWith the widespread application of federated learning in healthcare scenarios, ensuring performance fairness of disease diagnosis models across different medical institutions (clients) has attracted increasing attention. However, accurately evaluating whether models achieve this goal is equally critical yet faces numerous challenges: on one hand, each client can only evaluate the global model based on its own limited private data, which easily leads to performance estimation bias; on the other hand, due to data privacy constraints, clients cannot know the model's performance at other institutions, making it difficult to determine whether the global model truly achieves cross-client performance fairness. To address this, this paper proposes theFed4Fedfederated evaluation framework, which can more accurately evaluate the global model's performance in actual deployment while protecting data privacy by combining private data from multiple clients, and rigorously infer model performance fairness based on statistical hypothesis testing. Specifically,Fed4Feddraws inspiration from federated learning principles to collaboratively utilize multi-party private data while protecting data privacy. Second, it innovatively introduces Bootstrap methods and statistical inference strategies to construct and analyze the statistical distribution of model performance, reducing the randomness of performance evaluation. Third, based on statistical homogeneity testing theory, two fairness testing methods are designed to provide theoretical guarantees for evaluating performance fairness. Finally, experiments on synthetic datasets as well as four types of multi-modal real datasets including CIFAR-10, MNIST, Fashion-MNIST, and SST demonstrate that:Fed4Fedeffectively overcomes the limitations of existing evaluation methods, with a fairness misjudgment rate below 5%, an average confidence interval coverage rate of 94.28% for performance, and robust performance across different degrees of non-independent and identically distributed (non-IID) scenarios. Zhongchi Wang, Hailong Sun 0001, Wei Ni 0001, Xiang Gao 0012 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN ModelsabstractWith the growing incorporation of deep neural network (DNN) models into modern software systems, the prohibitive construction costs of DNN models have become a significant challenge in software development. To address this challenge, model reuse has been widely applied to reduce model training costs; however, indiscriminately reusing an entire model may incur significant inference overhead. Consequently, DNN modularization—borrowing the idea of modularization in software engineering—has increasingly gained attention, enabling module reuse by decomposing a DNN model into modules. In particular, the emerging modularizing-while-training (MwT) paradigm, which outperforms modularizing-after-training by incorporating modularization into the model’s training process, has been demonstrated as a more effective approach for DNN modularization. However, existing MwT approaches focus on small-scale convolutional neural network (CNN) models at the convolutional kernel level. They struggle to handle diverse DNNs and large-scale models, particularly Transformer-based models, which consistently achieve state-of-the-art results across various tasks. To address these limitations, we propose NeMo, a scalable and more generalizable MwT approach. NeMo operates at the neuron level—a fundamental component common to all DNNs—thereby ensuring applicability to Transformers and various DNN architectures. Moreover, we design a contrastive learning-based modular training method, equipped with an effective composite loss function, hence being scalable to large-scale models. Comprehensive experiments on two Transformer-based models and four CNN models across two widely used classification datasets demonstrate NeMo’s superiority over the state-of-the-art MwT method. Results show average performance gains of 1.72% in module classification accuracy and a 58.10% reduction in module size. Our findings demonstrate that NeMo exhibits efficacy across both CNN and large-scale Transformer-based models. Moreover, a case study based on open source projects demonstrates the potential benefits of NeMo in practical scenarios, offering a promising approach for achieving scalable and generalizable DNN modularization. Xiaohan Bi, Binhang Qi, Hailong Sun 0001, Xiang Gao 0012, Yue Yu 0001, Xiaojun Liang |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2026 | Reference-Based Retrieval-Augmented Unit Test GenerationabstractAutomated unit test generation has been widely studied, with Large Language Models (LLMs) recently showing significant potential. LLMs like GPT-4, trained in vast text and code data, excel in various code-related tasks, including unit test generation. However, existing LLM-based approaches often focus solely on the context within the code itself, such as referenced variables, while neglecting broader task-specific contexts, such as the utility of referring to existing tests of relevant methods in unit test generation. Moreover, in the context of unit test generation, these tools prioritize high code coverage, often at the expense of practical usability, correctness, and maintainability. In response, we propose Reference-Based Retrieval Augmentation , a novel mechanism that extends LLM-based Retrieval-Augmented Generation (RAG) to retrieve relevant information by considering task-specific context. In the unit test generation task, for a given focal method, the reference relationships is defined as the reusability or referentiality of tests between the focal method and other methods. To generate high-quality unit tests for the focal method, the test reference relationships are then used to retrieve relevant methods and their existing unit tests. Specifically, we account for the unique structure of unit tests by dividing the test generation process into Given , When , and Then phases. When generating unit tests for a focal method, we retrieve pre-existing tests of other relevant methods, which can provide valuable insights for any of the Given , When , and Then phases. We implement this approach in a tool called RefTest , which sequentially performs preprocessing, test reference retrieval, and unit test generation, using an incremental strategy in which newly generated tests guide the creation of subsequent ones. We evaluated RefTest on 12 open source projects with 1,515 methods, and the results demonstrate that RefTest consistently outperforms existing tools in terms of correctness, completeness, and maintainability of the generated tests. Yuanzhang Lin, Xiang Gao 0012, Hailong Sun 0001, Yuan Yuan 0004 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2026 | Scalable Large-Scale Multi-Granularity Code Clone Detection via Clustering Search and Pre-Trained ModelsabstractCode cloning is a common phenomenon in software development, which reduces developers’ programming efforts but also poses risks of defect inheritance. Clone detection locates exact or similar pieces of code within or between software systems. With the amount of source code increasing steadily, efficient and large-scale clone detection has become a necessity. Moreover, code clones may occur at various levels of code granularity, e.g., file, function, and block level, which pose more challenges for efficient clone detection. Although numerous methods have been proposed to detect code clones at different granularities, they often suffer from low detection efficiency, false positive results and are typically limited to identifying clones at a specific granularity. In this paper, we introduce an efficient clone detection, named MGCD, to detect code clones among large-scale codebases. Specifically, we embed function-level code into vectors using a pre-trained model and perform clustering search with the IVF Flat algorithm to identify clone candidates. These candidates are then filtered through an entropy-based method to enhance accuracy and avoid false positive results. Moreover, we leverage the information from function-level clone detection results to further conduct file and block level clone detection. We evaluate our approach on the BigCloneBench benchmark. Experimental results show that our approach only takes 0.23 ms to search clone results among 800,000 functions and achieves high precision and recall. Yifan An, Xiang Gao 0012, Hailong Sun 0001 |
IEEE Trans. Software Eng. | 3 |
| 2025 | UICOMPASS: UI Map Guided Mobile Task Automation via Adaptive Action GenerationabstractMobile task automation is an emerging technology that leverages AI to automatically execute routine tasks by users' commands on mobile devices like Android, thus enhancing efficiency and productivity.While large language models (LLMs) excel at general mobile tasks through training on massive datasets, they struggle with app-specific workflows.To solve this problem, we designed UI Map, a structured representation of target app's UI information.We further propose a UI Map-guided LLM-based approach UICOMPASS to automate mobile tasks.Specifically, UICOMPASS first leverages static analysis and LLMs to automatically build UI Map from either source codes of apps or byte codes (i.e., APK packages).During task execution, UICOMPASS mines the task-relevant information from UI Map to feed into the LLMs, generates a planned path, and adaptively adjusts the path based on the actual app state and action history.Experimental results demonstrate that UICOMPASS achieves a 14.52% higher task executing success rate than SOTA approaches.Even when only APK is available, UICOMPASS maintains superior performance, demonstrating its applicability to closed-source apps. Yuanzhang Lin, He Rui, Qingao Dong, Mingyi Zhou, Xiang Gao 0012, Hailong Sun 0001 |
EMNLP | 7 |
| 2025 | CABS: Conflict-Aware and Balanced Sparsification for Enhancing Model MergingabstractModel merging based on task vectors, i.e., the parameter differences between fine-tuned models and a shared base model, provides an efficient way to integrate multiple task-specific models into a multitask model without retraining. Recent works have endeavored to address the conflicts between task vectors, one of the significant challenges faced by model merging, through sparsification; however, two issues significantly limit their performance: high parameter overlap and unbalanced weight distribution. To address these issues, we propose a simple yet effective framework called CABS (Conflict-Aware and Balanced Sparsification), consisting of Conflict-Aware Sparsification (CA) and Balanced Sparsification (BS). CA reduces parameter overlap by applying masks during sequential pruning, ensuring that each task vector retains distinct, non-overlapping parameters. BS leverages $n$:$m$ pruning to preserve critical weights while maintaining an even distribution across layers. Our comprehensive experiments demonstrate that CABS outperforms state-of-the-art methods across diverse tasks and model sizes. Zongzhen Yang, Binhang Qi, Hailong Sun 0001, Wenrui Long, Ruobing Zhao, Xiang Gao 0012 |
ICML | 6 |
| 2025 | Code Property Graph Meets Typestate: A Scalable Framework to Behavioral Bug DetectionabstractBehavioral bugs caused by incorrect state changes are particularly challenging to identify because they depend on specific code execution paths. While code property graph (CPG) combine multiple code views through abstract syntax trees (AST), their built-in redundancy from syntax details and fixed connection rules make them hard to scale-a major problem when analyzing large software systems. We introduce QVoG, a new framework that improves CPG by combining graphbased code analysis with state behavior checking. Our main innovation lies in simplifying the CPG at the statement level by consolidating control and data flows into meaningful code blocks and optimizing the edges. This approach reduces the graph size by more than 10 times compared to AST-based methods while maintaining accuracy. This lightweight design allows easy integration of state tracking, where we match object lifecycle rules to simplified CPG connections using replaceable patterns. The combination of streamlined graphs and state-aware analysis helps QVoG effectively find difficult-to-identify behavioral bugs, successfully detecting 25 issues (including 17 confirmed cases and 2 official CVE) in real-world projects. Importantly, QVoG analyzes raw source code without requiring compilation and supports projects exceeding 1 million lines of code. Xingjing Deng, Zhengyao Liu, Xitong Zhong, Shuo Hong, Yixin Yang 0006, Xiang Gao 0012, Xuhui Yan, Hailong Sun 0001 |
ICSME | 6 |
| 2025 | Enhanced Vulnerability Localization: Harmonizing Task-Specific Tuning and General LLM PromptingabstractLarge Language Models (LLMs) have shown significant potential for vulnerability localization in software security. However, current LLM-based approaches face a critical dilemma: direct application of general-purpose LLMs lacks crucial domainspecific expertise, while fine-tuning suffers from limited robustness when faced with unfamiliar data. These problems result in subpar performance in vulnerability localization and weak generalization capabilities. To address these limitations, we introduce ENVUL, a novel domain adaptation framework for vulnerability localization. ENVUL improves vulnerability localization by synergizing enhanced task-specific tuning with prompt engineering of general-purpose LLMs. ENVUL incorporates three key innovations for addressing two problems: (1) how to optimize fine-tuning for localization task, and (2) when to wisely choose tuning and prompting. To solve the first problem, we introduce: (a). a context Consolidator that captures rich statement-level code semantic, improving the model's understanding of code context; (b). a semantic Indicator employing attention rectification to highlight patterns indicative of vulnerabilities, focusing the model on critical security signals. To solve the second problem, we introduce a dynamic routing mechanism based on joint-representation similarity analysis that strategically delegates tasks between the fine-tuned model and the general LLM. It ensures ENVUL's robust performance across diverse real-world vulnerability types. Real-world evaluations demonstrate ENVUL's robust expertise in outperforming state-of-the-art vulnerability localization baselines, achieving absolute improvements of$\mathbf{2 2. 7 \% - 3 0. 3 \%}$in top-1 accuracy. Notably, ENVul exhibits exceptional generalization, achieving$\mathbf{4 3. 6 \% - 5 0 \%}$higher accuracy on unfamiliar vulnerability types. Wentong Tian, Yuanzhang Lin, Xiang Gao 0012, Hailong Sun 0001 |
ICSME | 3 |
| 2025 | Enhancing Automated Vulnerability Repair Through Dependency Embedding and Pattern StoreabstractIn recent years, the proliferation of software vulnerabilities has significantly increased the complexities and costs associated with manual remediation efforts. Although AI-based methods for automated vulnerability repair are gaining traction, many existing approaches have two limitations: 1) treat code as a sequence of tokens, neglecting critical structural information like control flow and data flow, and 2) do not fully utilize the repair patterns of vulnerabilities. To address these limitations, we introduce FAVOR, an innovative tool that utilizes both the vulnerable function's code and its control flow graph (CFG) as inputs. FAVOR incorporates a dependency embedding module to capture structural and dependency information and leverages CodeT5, a state-of-the-art model pre-trained for code generation tasks. To further enhance the repair process, we introduce a pattern store that uses KNN search to retrieve similar past repair patterns, which helps guide the model toward generating more contextually accurate patches. In our experiments, FAVOR, trained on a dataset of 6548 faulty C/C++ functions, repaired 45 more vulnerabilities compared to VULREPAIR, demonstrating improved accuracy and efficiency in automated vulnerability repair. Qingao Dong, Yuanzhang Lin, Hailong Sun 0001, Xiang Gao 0012 |
SANER | 4 |
| 2025 | Understanding vulnerabilities in software supply chains
Xiang Gao 0012, Hailong Sun 0001 |
Empir. Softw. Eng. | 2 |
| 2025 | OCLVerifer: Automated verification of OCL contracts in requirements models
Peiye Yang, Li Zhang 0029, Xiang Gao 0012, Yilong Yang 0001 |
Sci. Comput. Program. | 4 |
| 2024 | Modularizing while Training: A New Paradigm for Modularizing DNN ModelsabstractDeep neural network (DNN) models have become increasingly crucial components of intelligent software systems. However, training a DNN model is typically expensive in terms of both time and computational resources. To address this issue, recent research has focused on reusing existing DNN models - borrowing the concept of software reuse in software engineering. However, reusing an entire model could cause extra overhead or inherit the weaknesses from the undesired functionalities. Hence, existing work proposes to decompose an already trained model into modules, i.e., modularizing-after-training, to enable module reuse. Since the trained models are not built for modularization, modularizing-after-training may incur huge overhead and model accuracy loss. In this paper, we propose a novel approach that incorporates modularization into the model training process, i.e., modularizing-while-training (MwT). We train a model to be structurally modular through two loss functions that optimize intra-module cohesion and inter-module coupling. We have implemented the proposed approach for modularizing Convolutional Neural Network (CNN) models. The evaluation results on representative models demonstrate that MwT outperforms the existing state-of-the-art modularizing-after-training approach. Specifically, the accuracy loss caused by MwT is only 1.13 percentage points, which is less than that of the existing approach. The kernel retention rate of the modules generated by MwT is only 14.58%, with a reduction of 74.31% over the existing approach. Furthermore, the total time cost required for training and modularizing is only 108 minutes, which is half the time required by the existing approach. Our work demonstrates that MwT is a new and more effective paradigm for realizing DNN model modularization, offering a fresh perspective on achieving model reuse. Binhang Qi, Hailong Sun 0001, Hongyu Zhang 0002, Ruobing Zhao, Xiang Gao 0012 |
ICSE | 5 |
| 2024 | Investigating White-Box Attacks for On-Device ModelsabstractNumerous mobile apps have leveraged deep learning capabilities. However, on-device models are vulnerable to attacks as they can be easily extracted from their corresponding mobile apps. Although the structure and parameters information of these models can be accessed, existing on-device attacking approaches only generate black-box attacks (i.e., indirect white-box attacks), which are less effective and efficient than white-box strategies. This is because mobile deep learning (DL) frameworks like TensorFlow Lite (TFLite) do not support gradient computing (referred to as non-debuggable models), which is necessary for white-box attacking algorithms. Thus, we argue that existing findings may underestimate the harm-fulness of on-device attacks. To validate this, we systematically analyze the difficulties of transforming the on-device model to its debuggable version and propose a Reverse Engineering framework for On-device Models (REOM), which automatically reverses the compiled on-device TFLite model to its debuggable version, enabling attackers to launch white-box attacks. Our empirical results show that our approach is effective in achieving automated transformation (i.e., 92.6%) among 244 TFLite models. Compared with previous attacks using surrogate models, REOM enables attackers to achieve higher attack success rates (10.23%→89.03%) with a hundred times smaller attack perturbations (1.0→0.01). Our findings emphasize the need for developers to carefully consider their model deployment strategies, and use white-box methods to evaluate the vulnerability of on-device models. Our artifacts 1 are available. Mingyi Zhou, Xiang Gao 0012, Jing Wu 0021, Kui Liu 0001, Hailong Sun 0001, Li Li 0029 |
ICSE | 2 |
| 2024 | API Misuse Detection via Probabilistic Graphical ModelabstractAPI misuses can cause a range of issues in software development, including program crashes, bugs, and vulnerabilities. Different approaches have been developed to automatically detect API misuses by checking the program against usage rules extracted from extensive codebase or API documents. However, these mined rules may not be precise or complete, leading to high false positive/negative rates. In this paper, we propose a novel solution to this problem by representing the mined API usage rules as a probabilistic graphical model, where each rule's probability value represents its trustworthiness of being correct. Our approach automatically constructs probabilistic usage rules by mining codebase and documents, and aggregating knowledge from different sources. Here, the usage rules obtained from the codebase initialize the probabilistic model, while the knowledge from the documents serves as a supplement for adjusting and complementing the probabilities accordingly. We evaluate our approach on the MuBench benchmark. Experimental results show that our approach achieves 42.0% precision and 54.5% recall, significantly outperforming state-of-the-art approaches. Wentong Tian, Xiang Gao 0012, Hailong Sun 0001, Li Li 0029 |
ISSTA | 3 |
| 2024 | Model-less Is the Best Model: Generating Pure Code Implementations to Replace On-Device DL ModelsabstractRecent studies show that on-device deployed deep learning (DL) models, such as those of Tensor Flow Lite (TFLite), can be easily extracted from real-world applications and devices by attackers to generate many kinds of adversarial and other attacks. Although securing deployed on-device DL models has gained increasing attention, no existing methods can fully prevent these attacks. Traditional software protection techniques have been widely explored. If on-device models can be implemented using pure code, such as C++, it will open the possibility of reusing existing robust software protection techniques. However, due to the complexity of DL models, there is no automatic method that can translate DL models to pure code. To fill this gap, we propose a novel method, CustomDLCoder, to automatically extract on-device DL model information and synthesize a customized executable program for a wide range of DL models. CustomDLCoder first parses the DL model, extracts its backend computing codes, configures the extracted codes, and then generates a customized program to implement and deploy the DL model without explicit model representation. The synthesized program hides model information for DL deployment environments since it does not need to retain explicit model representation, preventing many attacks on the DL model. In addition, it improves ML performance because the customized code removes model parsing and preprocessing steps and only retains the data computing process. Our experimental results show that CustomDLCoder improves model security by disabling on-device model sniffing. Compared with the original on-device platform (i.e., TFLite), our method can accelerate model inference by 21.0% and 24.3% on x86-64 and ARM64 platforms, respectively. Most importantly, it can significantly reduce memory consumption by 68.8% and 36.0% on x86-64 and ARM64 platforms, respectively. Mingyi Zhou, Xiang Gao 0012, John C. Grundy, Chunyang Chen 0001, Xiao Chen 0002, Li Li 0029 |
ISSTA | 2 |
| 2024 | LLM-Based Java Concurrent Program to ArkTS ConverterabstractHarmonyOS NEXT is a distributed operating system developed to support HarmonyOS native apps. To support the new and independent Harmony ecosystem, developers are required to migrate their applications from Android to HarmonyOS. However, HarmonyOS utilizes ArkTS, a superset of TypeScript, as the programming language for application development. Hence, migrating applications to HarmonyOS requires translating programs across different program languages, e.g., Java, which is known to be very challenging, especially for concurrency programs. Java utilizes shared memory to implement concurrency programs, while ArkTS relies on message passing (i.e., Actor model). This paper presents an LLM-based concurrent Java program to ArkTS converter. Runlin Liu, Yunge Hu, Xiang Gao 0012 |
ASE | 5 |
| 2024 | DynaMO: Protecting Mobile DL Models through Coupling Obfuscated DL OperatorsabstractDeploying deep learning (DL) models on mobile applications (Apps) has become ever-more popular. However, existing studies show attackers can easily reverse-engineer mobile DL models in Apps to steal intellectual property or generate effective attacks. A recent approach, Model Obfuscation, has been proposed to defend against such reverse engineering by obfuscating DL model representations, such as weights and computational graphs, without affecting model performance. These existing model obfuscation methods use static methods to obfuscate the model representation, or they use half-dynamic methods but require users to restore the model information through additional input arguments. However, these static methods or half-dynamic methods cannot provide enough protection for on-device DL models. Attackers can use dynamic analysis to mine the sensitive information in the inference codes as the correct model information and intermediate results must be recovered at runtime for static and half-dynamic obfuscation methods. We assess the vulnerability of the existing obfuscation strategies using an instrumentation method and tool, DLModelExplorer, that dynamically extracts correct sensitive model information (i.e., weights, computational graph) at runtime. Experiments show it achieves very high attack performance (e.g., 98.76% of weights extraction rate and 99.89% of obfuscating operator classification rate). To defend against such attacks based on dynamic instrumentation, we propose DynaMO, a Dynamic Model Obfuscation strategy similar to Homomorphic Encryption. The obfuscation and recovery process can be done through simple linear transformation for the weights of randomly coupled eligible operators, which is a fully dynamic obfuscation strategy. Experiments show that our proposed strategy can dramatically improve model security compared with the existing obfuscation strategies, with only negligible overheads for on-device models. Our prototype tool is publicly available at https://github.com/zhoumingyi/DynaMO. Mingyi Zhou, Xiang Gao 0012, Xiao Chen 0002, Chunyang Chen 0001, John C. Grundy, Li Li 0029 |
ASE | 2 |
| 2024 | ModelGalaxy: A Versatile Model Retrieval PlatformabstractWith the growing number of available machine learning models and the emergence of model-sharing platforms, model reuse has become a significant approach to harnessing the power of artificial intelligence. One of the key issues to realizing model reuse resides in efficiently and accurately finding the target models that meet user needs from a model repository. However, the existing popular model-sharing platforms (e.g., Hugging Face) mainly support model retrieval based on model name matching and task filtering. If not familiar with the platform or specific models, users may suffer from low retrieval efficiency and a less user-friendly interaction experience. To address these issues, we have developed ModelGalaxy, a versatile model retrieval platform supporting multiple model retrieval methods, including keyword-based search, dataset-based search, and user-task-centric search. Moreover, ModelGalaxy leverages the power of large language models to provide users with easily retrieving and using models. Our source code is available at https://github.com/zwl906711886/ModelGalaxy. Wenling Zhang, Zhaotian Li, Hailong Sun 0001, Xiang Gao 0012, Xudong Liu 0001 |
SIGIR | 5 |
| 2024 | Investigating and Detecting Silent Bugs in PyTorch ProgramsabstractDeep Learning (DL) has been widely applied in various fields. Unlike traditional software, DL programs possess the “black box” characteristic that can make it challenging for developers to debug when anomalous behaviors arise. In particular, silent bugs, a type of bugs in DL programs, can lead to erroneous behaviors without causing system crashes or suspensions, and they do not display error messages to users. This makes silent bugs more difficult for developers to discover, locate, and fix. In this paper, we present the first detailed study of silent bugs in PyTorch programs. We collect 14,523 posts from the official PyTorch forum and use a LLM-based semi-automated approach to filter the silent bugs. By analyzing the symptoms, root causes, and patterns of silent bugs, we have derived several important findings and implications: (1) most silent bugs cause abnormal outputs, which requires the design of more flexible test oracles to detect them, (2) the wide range of symptoms and root causes do not necessarily have one-to-one correspondences, which makes detecting and debugging silent bugs more challenging, (3) silent bugs exhibit common bug patterns, such as redundant, missing, or misplaced operations. Building upon these findings, we design and implement an extensible rule-based tool PYSIASSIST to help developer debug and resolve silent bugs. Evaluation results show that Pysiassist achieves 92.4% precision and 85.3% recall, outperforming existing techniques. Shuo Hong, Hailong Sun 0001, Xiang Gao 0012, Shin Hwei Tan |
SANER | 3 |
| 2024 | Reducing False Positives of Static Bug Detectors Through Code Representation LearningabstractWith the increasing significance of software correctness and security, automatic static analysis tools (ASATs) play a more and more important role in software development due to their ability and scalability. However, compared to dynamic analysis methods, static tools often suffer from the severe problem of generating high false positive rates, due to their analysis mechanisms. To alleviate the false positive problem, many approaches have been proposed, which focus on manually extracted features from code snippets and then prioritize real warnings by means of statistics or machine learning techniques. However, manual encoded features are insufficient to achieve satisfactory performance across different datasets. In this study, we focus on exploring the effectiveness of various code representation learning (CRL) techniques in understanding the semantics of warnings generated by ASATs. In particular, our large-scale empirical study not only reveals that CRL models can effectively differentiate buggy code snippets (i.e., containing warnings detected by ASATs) from clean ones (the median of F1-score reaches 87.3 % for binary classification, and reaches 77.4 % for multi-class classification), they are also promising in identifying false positive warnings (the F1-score of best performer is 75.6%). Such findings drive us to further design a novel approach named PRI SM, to PRIoritize Static warnings based on aggregating multiple CRL Models to reduce the false positives generated by existing ASATs. Extensive evaluations demonstrate that our designed approach can outperform existing baselines significantly. Yixin Yang 0006, Ming Wen 0001, Xiang Gao 0012, Hailong Sun 0001 |
SANER | 3 |
| 2024 | Reusing Convolutional Neural Network Models through Modularization and CompositionabstractWith the widespread success of deep learning technologies, many trained deep neural network (DNN) models are now publicly available. However, directly reusing the public DNN models for new tasks often fails due to mismatching functionality or performance. Inspired by the notion of modularization and composition in software reuse, we investigate the possibility of improving the reusability of DNN models in a more fine-grained manner. Specifically, we propose two modularization approaches named CNNSplitter and GradSplitter, which can decompose a trained convolutional neural network (CNN) model for N -class classification into N small reusable modules. Each module recognizes one of the N classes and contains a part of the convolution kernels of the trained CNN model. Then, the resulting modules can be reused to patch existing CNN models or build new CNN models through composition. The main difference between CNNSplitter and GradSplitter lies in their search methods: the former relies on a genetic algorithm to explore search space, while the latter utilizes a gradient-based search method. Our experiments with three representative CNNs on three widely used public datasets demonstrate the effectiveness of the proposed approaches. Compared with CNNSplitter, GradSplitter incurs less accuracy loss, produces much smaller modules (19.88% fewer kernels), and achieves better results on patching weak models. In particular, experiments on GradSplitter show that (1) by patching weak models, the average improvement in terms of precision, recall, and F1-score is 17.13%, 4.95%, and 11.47%, respectively, and (2) for a new task, compared with the models trained from scratch, reusing modules achieves similar accuracy (the average loss of accuracy is only 2.46%) without a costly training process. Our approaches provide a viable solution to the rapid development and improvement of CNN models. Binhang Qi, Hailong Sun 0001, Hongyu Zhang 0002, Xiang Gao 0012 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | AutoMRM: A Model Retrieval Method Based on Multimodal Query and Meta-learningabstractWith more and more Deep Neural Network (DNN) models are publicly available on model sharing platforms (e.g., HuggingFace), model reuse has become a promising way in practice to improve the efficiency of DNN model construction by avoiding the costs of model training. To that end, a pivotal step for model reuse is model retrieval, which facilitates discovering suitable models from a model hub that match the requirements of users. However, the existing model retrieval methods have inadequate performance and efficiency, since they focus on matching user requirements with the model names, and thus cannot work well for high-dimensional data such as images. In this paper, we propose a user-task-centric multimodal model retrieval method named AutoMRM. AutoMRM can retrieve DNN models suitable for the user's task according to both the dataset and description of the task. Moreover, AutoMRM utilizes meta-learning to retrieve models for previously unseen task queries. Specifically, given a task, AutoMRM extracts the latent meta-features from the dataset and description for training meta-learners offline and obtaining the representation of user task queries online. Experimental results demonstrate that AutoMRM outperforms existing model retrieval methods including the state-of-the-art method in both effectiveness and efficiency. Zhaotian Li, Binhang Qi, Hailong Sun 0001, Xiang Gao 0012 |
CIKM | 4 |
| 2023 | Automated Repair of Programs from Large Language ModelsabstractLarge language models such as Codex, have shown the capability to produce code for many programming tasks. However, the success rate of existing models is low, especially for complex programming tasks. One of the reasons is that language models lack awareness of program semantics, resulting in incorrect programs, or even programs which do not compile. In this paper, we systematically study whether automated program repair (APR) techniques can fix the incorrect solutions produced by language models in LeetCode contests. The goal is to study whether APR techniques can enhance reliability in the code produced by large language models. Our study revealed that: (1) automatically generated code shares common programming mistakes with human-crafted solutions, indicating APR techniques may have potential to fix auto-generated code; (2) given bug location information provided by a statistical fault localization approach, the newly released Codex edit mode, which supports editing code, is similar to or better than existing Java repair tools TBar and Recoder in fixing incorrect solutions. By analyzing the experimental results generated by these tools, we provide several suggestions: (1) enhancing APR tools to surpass limitations in patch space (e.g., introducing more flexible fault localization) is desirable; (2) as large language models can derive more fix patterns by training on more data, future APR tools could shift focus from adding more fix patterns to synthesis/semantics based approaches, (3) combination of language models with APR to curate patch ingredients, is worth studying. Zhiyu Fan, Xiang Gao 0012, Martin Mirchev, Abhik Roychoudhury, Shin Hwei Tan |
ICSE | 2 |
| 2023 | Reusing Deep Neural Network Models through Model Re-engineeringabstractTraining deep neural network (DNN) models, which has become an important task in today's software development, is often costly in terms of computational resources and time. With the inspiration of software reuse, building DNN models through reusing existing ones has gained increasing attention recently. Prior approaches to DNN model reuse have two main limitations: 1) reusing the entire model, while only a small part of the model's functionalities (labels) are required, would cause much overhead (e.g., computational and time costs for inference), and 2) model reuse would inherit the defects and weaknesses of the reused model, and hence put the new system under threats of security attack. To solve the above problem, we propose SeaM, a tool that re-engineers a trained DNN model to improve its reusability. Specifically, given a target problem and a trained model, SeaM utilizes a gradient-based search method to search for the model's weights that are relevant to the target problem. The re-engineered model that only retains the relevant weights is then reused to solve the target problem. Evaluation results on widely-used models show that the re-engineered models produced by SeaM only contain 10.11% weights of the original models, resulting 42.41% reduction in terms of inference time. For the target problem, the re-engineered models even outperform the original models in classification accuracy by 5.85%. Moreover, reusing the re-engineered models inherits an average of 57% fewer defects than reusing the entire model. We believe our approach to reducing reuse overhead and defect inheritance is one important step forward for practical model reuse. Binhang Qi, Hailong Sun 0001, Xiang Gao 0012, Hongyu Zhang 0002, Zhaotian Li, Xudong Liu 0001 |
ICSE | 3 |
| 2023 | ModelObfuscator: Obfuscating Model Information to Protect Deployed ML-Based SystemsabstractMore and more edge devices and mobile apps are leveraging deep learning (DL) capabilities. Deploying such models on devices – referred to as on-device models – rather than as remote cloud-hosted services, has gained popularity because it avoids transmitting user’s data off of the device and achieves high response time. However, on-device models can be easily attacked, as they can be accessed by unpacking corresponding apps and the model is fully exposed to attackers. Recent studies show that attackers can easily generate white-box-like attacks for an on-device model or even inverse its training data. To protect on-device models from white-box attacks, we propose a novel technique called model obfuscation. Specifically, model obfuscation hides and obfuscates the key information – structure, parameters and attributes – of models by renaming, parameter encapsulation, neural structure obfuscation, shortcut injection, and extra layer injection. We have developed a prototype tool ModelObfuscator to automatically obfuscate on-device TFLite models. Our experiments show that this proposed approach can dramatically improve model security by significantly increasing the difficulty of parsing models’ inner information, without increasing the latency of DL models. Our proposed on-device model obfuscation has the potential to be a fundamental technique for on-device model deployment. Our prototype tool is publicly available at https://github.com/zhoumingyi/ModelObfuscator. Mingyi Zhou, Xiang Gao 0012, Jing Wu 0021, John C. Grundy, Xiao Chen 0002, Chunyang Chen 0001, Li Li 0029 |
ISSTA | 2 |
| 2023 | Automated Fixing of Web UI Tests via Iterative Element MatchingabstractWeb UI test cases are used for the automatic testing of web applications. When a web application is updated, these UI tests should also be updated for regression testing of the new version of web application. With the rapid evolution, updating UI tests is a tedious and time-consuming task. To solve these problems, automatically repairing web UI tests has gained increasing attention recently. To repair web UI tests, the most important step is to match the UI elements before and after the web page update. Existing work matches UI elements according to visual information, attributes value, or Document Object Model (DOM) structures. However, they either achieve low element matching accuracy or only work on simple UI tests. To solve these problems, we proposed UITestFix, an approach based on a novel iterative matching algorithm for improving the accuracy of matching UI elements. UITestFix is designed based on two main insights: (1) beyond attribute and DOM structures, the relations between different elements can also guide the matching process, and (2) the matching results of previous iterations could guide the matching of the current iteration. Our evaluation of publicly available datasets and two industrial apps shows that UITestFix outperforms four existing approaches by achieving more accurate element matching and producing more correct fixes. Yuanzhang Lin, Guoyao Wen, Xiang Gao 0012 |
ASE | 3 |
| 2022 | Trust Enhancement Issues in Program RepairabstractAutomated program repair is an emerging technology that seeks to automatically rectify bugs and vulnerabilities using learning, search, and semantic analysis. Trust in automatically generated patches is necessary for achieving greater adoption of program repair. Towards this goal, we survey more than 100 software practitioners to understand the artifacts and setups needed to enhance trust in automatically generated patches. Based on the feedback from the survey on developer preferences, we quantitatively evaluate existing test-suite based program repair tools. We find that they cannot produce high-quality patches within a top-10 ranking and an acceptable time period of 1 hour. The developer feedback from our qualitative study and the observations from our quantitative examination of existing repair tools point to actionable insights to drive program repair research. Specifically, we note that producing repairs within an acceptable time-bound is very much dependent on leveraging an abstract search space representation of a rich enough search space. Moreover, while additional developer inputs are valuable for generating or ranking patches, developers do not seem to be interested in a significant human-in-the-loop interaction. Yannic Noller, Ridwan Salihin Shariffdeen, Xiang Gao 0012, Abhik Roychoudhury |
ICSE | 3 |
| 2022 | Program vulnerability repair via inductive inferenceabstractProgram vulnerabilities, even when detected and reported, are not fixed immediately. The time lag between the reporting and fixing of a vulnerability causes open-source software systems to suffer from significant exposure to possible attacks. In this paper, we propose a counter-example guided inductive inference procedure over program states to define likely invariants at possible fix locations. The likely invariants are constructed via mutation over states at the fix location, which turns out to be more effective for inductive property inference, as compared to the usual greybox fuzzing over program inputs. Once such likely invariants, which we call patch invariants, are identified, we can use them to construct patches via simple patch templates. Our work assumes that only one failing input (representing the exploit) is available to start the repair process. Experiments on the VulnLoc data-set of 39 vulnerabilities, which has been curated in previous works on vulnerability repair, show the effectiveness of our repair procedure. As compared to proposed approaches for vulnerability repair such as CPR or SenX which are based on concolic and symbolic execution respectively, we can repair significantly more vulnerabilities. Our results show the potential for program repair via inductive constraint inference, as opposed to generating repair constraints via deductive/symbolic analysis of a given test-suite. Yuntong Zhang 0002, Xiang Gao 0012, Gregory J. Duck, Abhik Roychoudhury |
ISSTA | 2 |
| 2022 | Patching Weak Convolutional Neural Network Models through Modularization and CompositionabstractDespite great success in many applications, deep neural networks are not always robust in practice. For instance, a convolutional neuron network (CNN) model for classification tasks often performs unsatisfactorily in classifying some particular classes of objects. In this work, we are concerned with patching the weak part of a CNN model instead of improving it through the costly retraining of the entire model. Inspired by the fundamental concepts of modularization and composition in software engineering, we propose a compressed modularization approach, CNNSplitter, which decomposes a strong CNN model for N-class classification into N smaller CNN modules. Each module is a sub-model containing a part of the convolution kernels of the strong model. To patch a weak CNN model that performs unsatisfactorily on a target class (TC), we compose the weak CNN model with the corresponding module obtained from a strong CNN model. The ability of the weak CNN model to recognize the TC can thus be improved through patching. Moreover, the ability to recognize non-TCs is also improved, as the samples misclassified as TC could be classified as non-TCs correctly. Experimental results with two representative CNNs on three widely-used datasets show that the averaged improvement on the TC in terms of precision and recall are 12.54% and 2.14%, respectively. Moreover, patching improves the accuracy of non-TCs by 1.18%. The results demonstrate that CNNSplitter can patch a weak CNN model through modularization and composition, thus providing a new solution for developing robust CNN models. Binhang Qi, Hailong Sun 0001, Xiang Gao 0012, Hongyu Zhang 0002 |
ASE | 3 |
| 2021 | Automated patch backporting in Linux (experience paper)abstractWhenever a bug or vulnerability is detected in the Linux kernel, the kernel developers will endeavour to fix it by introducing a patch into the mainline version of the Linux kernel source tree. However, many users run older “stable” versions of Linux, meaning that the patch should also be “backported” to one or more of these older kernel versions. This process is error-prone and there is usually along delay in publishing the backported patch. Based on an empirical study, we show that around 8% of all commits submitted to Linux mainline are backported to older versions,but often more than one month elapses before the backport is available. Hence, we propose a patch backporting technique that can automatically transfer patches from the mainline version of Linux into older stable versions. Our approach first synthesizes a partial transformation rule based on a Linux mainline patch. This rule can then be generalized by analysing the alignment between the mainline and target versions. The generalized rule is then applied to the target version to produce a backported patch. We have implemented our transformation technique in a tool called FixMorph and evaluated it on 350 Linux mainline patches. FixMorph correctly backports 75.1% of them. Compared to existing techniques, FixMorph improves both the precision and recall in backporting patches. Apart from automation of software maintenance tasks, patch backporting helps in reducing the exposure to known security vulnerabilities in stable versions of the Linux kernel. Ridwan Salihin Shariffdeen, Xiang Gao 0012, Gregory J. Duck, Shin Hwei Tan, Julia Lawall, Abhik Roychoudhury |
ISSTA | 2 |
| 2021 | Scalable Fuzzing of Program Binaries with E9AFLabstractGreybox fuzzing is an effective method for software testing. Greybox fuzzers, such as AFL, use instrumentation that collects path coverage information in order to guide the fuzzing process. The instrumentation is usually inserted by a modified compiler toolchain, meaning that the program must be recompiled in order to be compatible with greybox fuzzing. When source code is unavailable, or for projects with complex build systems, recompilation is not always feasible. In this paper, we present E9AFL, a fast and scalable tool that automatically inserts AFL instrumentation to program binaries. E9AFL is built on top of the E9Patch static binary rewriting tool. To combat the overhead caused by binary instrumentation, E9AFL develops a set of optimization strategies. Our evaluation results show that E9AFL outperforms existing binary instrumentation tools and achieves comparable performance with the compile time instrumentation. Xiang Gao 0012, Gregory J. Duck, Abhik Roychoudhury |
ASE | 1 |
| 2021 | APIfix: output-oriented program synthesis for combating breaking changes in librariesabstractUse of third-party libraries is extremely common in application software. The libraries evolve to accommodate new features or mitigate security vulnerabilities, thereby breaking the Application Programming Interface(API) used by the software. Such breaking changes in the libraries may discourage client code from using the new library versions thereby keeping the application vulnerable and not up-to-date. We propose a novel output-oriented program synthesis algorithm to automate API usage adaptations via program transformation. Our aim is not only to rely on the few example human adaptations of the clients from the old library version to the new library version, since this can lead to over-fitting transformation rules. Instead, we also rely on example usages of the new updated library in clients, which provide valuable context for synthesizing and applying the transformation rules. Our tool APIFix provides an automated mechanism to transform application code using the old library versions to code using the new library versions - thereby achieving automated API usage adaptation to fix the effect of breaking changes. Our evaluation shows that the transformation rules inferred by APIFix achieve 98.7% precision and 91.5% recall. By comparing our approach to state-of-the-art program synthesis approaches, we show that our approach significantly reduces over-fitting while synthesizing transformation rules for API usage adaptations. Xiang Gao 0012, Arjun Radhakrishna, Gustavo Soares, Ridwan Salihin Shariffdeen, Sumit Gulwani, Abhik Roychoudhury |
Proc. ACM Program. Lang. | 1 |
| 2021 | Beyond Tests: Program Vulnerability Repair via Crash Constraint ExtractionabstractAutomated program repair is an emerging technology that seeks to automatically rectify program errors and vulnerabilities. Repair techniques are driven by a correctness criterion that is often in the form of a test suite. Such test-based repair may produce overfitting patches, where the patches produced fail on tests outside the test suite driving the repair. In this work, we present a repair method that fixes program vulnerabilities without the need for a voluminous test suite. Given a vulnerability as evidenced by an exploit, the technique extracts a constraint representing the vulnerability with the help of sanitizers. The extracted constraint serves as a proof obligation that our synthesized patch should satisfy. The proof obligation is met by propagating the extracted constraint to locations that are deemed to be “suitable” fix locations. An implementation of our approach (E xtract F ix ) on top of the KLEE symbolic execution engine shows its efficacy in fixing a wide range of vulnerabilities taken from the ManyBugs benchmark, real-world CVEs and Google’s OSS-Fuzz framework. We believe that our work presents a way forward for the overfitting problem in program repair by generalizing observable hazards/vulnerabilities (as constraint) from a single failing test or exploit. Xiang Gao 0012, Bo Wang 0050, Gregory J. Duck, Ruyi Ji, Yingfei Xiong 0001, Abhik Roychoudhury |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2020 | Fuzz testing based data augmentation to improve robustness of deep neural networksabstractDeep neural networks (DNN) have been shown to be notoriously brittle to small perturbations in their input data. This problem is analogous to the over-fitting problem in test-based program synthesis and automatic program repair, which is a consequence of the incomplete specification, i.e., the limited tests or training examples, that the program synthesis or repair algorithm has to learn from. Recently, test generation techniques have been successfully employed to augment existing specifications of intended program behavior, to improve the generalizability of program synthesis and repair. Inspired by these approaches, in this paper, we propose a technique that re-purposes software testing methods, specifically mutation-based fuzzing, to augment the training data of DNNs, with the objective of enhancing their robustness. Our technique casts the DNN data augmentation problem as an optimization problem. It uses genetic search to generate the most suitable variant of an input data to use for training the DNN, while simultaneously identifying opportunities to accelerate training by skipping augmentation in many instances. We instantiate this technique in two tools, Sensei and Sensei-SA, and evaluate them on 15 DNN models spanning 5 popular image data-sets. Our evaluation shows that Sensei can improve the robust accuracy of the DNN, compared to the state of the art, on each of the 15 models, by upto 11.9% and 5.5% on average. Further, Sensei-SA can reduce the average DNN training time by 25%, while still improving robust accuracy. Xiang Gao 0012, Ripon K. Saha, Mukul R. Prasad, Abhik Roychoudhury |
ICSE | 1 |
| 2020 | Binary rewriting without control flow recoveryabstractStatic binary rewriting has many important applications in software security and systems, such as hardening, repair, patching, instrumentation, and debugging. While many different static binary rewriting tools have been proposed, most rely on recovering control flow information from the input binary. The recovery step is necessary since the rewriting process may move instructions, meaning that the set of jump targets in the rewritten binary needs to be adjusted accordingly. Since the static recovery of control flow information is a hard problem in general, most tools rely on a set of simplifying heuristics or assumptions, such as specific compilers, specific source languages, or binary file meta information. However, the reliance on assumptions or heuristics tends to scale poorly in practice, and most state-of-the-art static binary rewriting tools cannot handle very large/complex programs such as web browsers. Gregory J. Duck, Xiang Gao 0012, Abhik Roychoudhury |
PLDI | 2 |
| 2020 | Feedback-driven semi-supervised synthesis of program transformationsabstractWhile editing code, it is common for developers to make multiple related repeated edits that are all instances of a more general program transformation. Since this process can be tedious and error-prone, we study the problem of automatically learning program transformations from past edits, which can then be used to predict future edits. We take a novel view of the problem as a semi-supervised learning problem: apart from the concrete edits that are instances of the general transformation, the learning procedure also exploits access to additional inputs (program subtrees) that are marked as positive or negative depending on whether the transformation applies on those inputs. We present a procedure to solve the semi-supervised transformation learning problem using anti-unification and programming-by-example synthesis technology. To eliminate reliance on access to marked additional inputs, we generalize the semi-supervised learning procedure to a feedback-driven procedure that also generates the marked additional inputs in an iterative loop. We apply these ideas to build and evaluate three applications that use different mechanisms for generating feedback. Compared to existing tools that learn program transformations from edits, our feedback-driven semi-supervised approach is vastly more effective in successfully predicting edits with significantly lesser amounts of past edit data. Xiang Gao 0012, Shraddha Barke, Arjun Radhakrishna, Gustavo Soares, Sumit Gulwani, Alan Leung, Nachiappan Nagappan, Ashish Tiwari 0001 |
Proc. ACM Program. Lang. | 1 |
| 2019 | Crash-avoiding program repairabstractExisting program repair systems modify a buggy program so that the modified program passes given tests. The repaired program may not satisfy even the most basic notion of correctness, namely crash-freedom. In other words, repair tools might generate patches which over-fit the test data driving the repair, and the automatically repaired programs may even introduce crashes or vulnerabilities. We propose an integrated approach for detecting and discarding crashing patches. Our approach fuses test and patch generation into a single process, in which patches are generated with the objective of passing existing tests, and new tests are generated with the objective of filtering out over-fitted patches by distinguishing candidate patches in terms of behavior. We use crash-freedom as the oracle to discard patch candidates which crash on the new tests. In its core, our approach defines a grey-box fuzzing strategy that gives higher priority to new tests that separate patches behaving equivalently on existing tests. This test generation strategy identifies semantic differences between patch candidates, and reduces over-fitting in program repair. We evaluated our approach on real-world vulnerabilities and open-source subjects from the Google OSS-Fuzz infrastructure. We found that our tool Fix2Fit (implementing patch space directed test generation), produces crash-avoiding patches. While we do not give formal guarantees about crash-freedom, cross-validation with fuzzing tools and their sanitizers provides greater confidence about the crash-freedom of our suggested patches. Xiang Gao 0012, Sergey Mechtaev, Abhik Roychoudhury |
ISSTA | 1 |
| 2019 | EROFS: A Compression-friendly Readonly File System for Resource-scarce Devices
Xiang Gao 0012, Mingkai Dong 0002, Xie Miao, Haibo Chen 0001 |
USENIX ATC | 1 |
| 2018 | Repairing crashes in Android appsabstractAndroid apps are omnipresent, and frequently suffer from crashes --- leading to poor user experience and economic loss. Past work focused on automated test generation to detect crashes in Android apps. However, automated repair of crashes has not been studied. In this paper, we propose the first approach to automatically repair Android apps, specifically we propose a technique for fixing crashes in Android apps. Unlike most test-based repair approaches, we do not need a test-suite; instead a single failing test is meticulously analyzed for crash locations and reasons behind these crashes. Our approach hinges on a careful empirical study which seeks to establish common root-causes for crashes in Android apps, and then distills the remedy of these root-causes in the form of eight generic transformation operators. These operators are applied using a search-based repair framework embodied in our repair tool Droix. We also prepare a benchmark DroixBench capturing reproducible crashes in Android apps. Our evaluation of Droix on DroixBench reveals that the automatically produced patches are often syntactically identical to the human patch, and on some rare occasion even better than the human patch (in terms of avoiding regressions). These results confirm our intuition that our proposed transformations form a sufficient set of operators to patch crashes in Android. Shin Hwei Tan, Xiang Gao 0012, Abhik Roychoudhury |
ICSE | 3 |
| 2018 | Android testing via synthetic symbolic executionabstractSymbolic execution of Android applications is challenging as it involves either building a customized VM for Android or modeling the Android libraries. Since the Android Runtime evolves from one version to another, building a high-fidelity symbolic execution engine involves modeling the effect of the libraries and their evolved versions. Without simulating the behavior of Android libraries, path divergence may occur due to constraint loss when the symbolic values flow into Android framework and these values later affect the subsequent path taken. Previous works such as JPF-Android have relied on the modeling of execution environment such as libraries. In this work, we build a dynamic symbolic execution engine for Android apps, without any manual modeling of execution environment. Environment (or library) dependent control flow decisions in the application will trigger an on-demand program synthesis step to automatically deduce a representation of the library.This representation is refined on-the-fly by running the corresponding library multiple times.The overarching goal of the refinement is to enhance behavioral coverage and to alleviate the path divergence problem during symbolic execution. Moreover, our library synthesis can be made context-specific. Compared to traditional synthesis approaches which aim to synthesize the complete library code, our context-specific synthesis engine can generate more precise expressions for a given context. The evaluation of our dynamic symbolic execution engine, built on top of JDART, shows that the library models obtained from program synthesis are often more accurate than the semi-manual models in JPF-Android. Furthermore, our symbolic execution engine could reach more branch targets, as compared to using the JPF-Android models. Xiang Gao 0012, Shin Hwei Tan, Abhik Roychoudhury |
ASE | 1 |
| 2018 | Test-Equivalence Analysis for Automatic Patch GenerationabstractAutomated program repair is a problem of finding a transformation (called a patch) of a given incorrect program that eliminates the observable failures. It has important applications such as providing debugging aids, automatically grading student assignments, and patching security vulnerabilities. A common challenge faced by existing repair techniques is scalability to large patch spaces, since there are many candidate patches that these techniques explicitly or implicitly consider. The correctness criteria for program repair is often given as a suite of tests. Current repair techniques do not scale due to the large number of test executions performed by the underlying search algorithms. In this work, we address this problem by introducing a methodology of patch generation based on a test-equivalence relation (if two programs are “test-equivalent” for a given test, they produce indistinguishable results on this test). We propose two test-equivalence relations based on runtime values and dependencies, respectively, and present an algorithm that performs on-the-fly partitioning of patches into test-equivalence classes. Our experiments on real-world programs reveal that the proposed methodology drastically reduces the number of test executions and therefore provides an order of magnitude efficiency improvement over existing repair techniques, without sacrificing patch quality. Sergey Mechtaev, Xiang Gao 0012, Shin Hwei Tan, Abhik Roychoudhury |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2016 | Write-back aware shared last-level cache management for hybrid main memoryabstractHybrid main memory with both DRAM and emerging non-volatile memory (NVM) becomes a promising solution for high performance and energy-efficient embedded systems. Cache plays an important role and highly affects the number of write backs to NVM and DRAM blocks. However, existing cache policies fail to fully address the significant asymmetry between NVM operations (especially writes) and DRAM operations, leading to non-optimal system designs. We propose a write-back aware last-level cache management scheme for the hybrid main memory, which improves the cache hit ratio of NVM memory blocks and minimizes write-backs to NVM. Experimental results show that our proposed framework leads to better performance and energy saving compared with the state-of-the-art cache management scheme for hybrid main memory architecture. Deshan Zhang, Lei Ju 0001, Mengying Zhao, Xiang Gao 0012, Zhiping Jia |
DAC | 4 |