VLDB 2026 Research / reviewers in the wild / expert
Wei Yang 0013
dblp:03/1094-13
· DBLP profile ↗
60ranked-venue papers
4as first author
40since 2021 · last 2026
0000-0002-5338-7347ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 33 · 3 first-author · 22 since 2021Artificial intelligence and machine learning · 11 · 9 since 2021Security and privacy · 9 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation MetricsabstractWeb applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and commercial potential.However, building a benchmark for LLM-generated web apps remains challenging due to the need for realworld user requirements, generalized evaluation metrics without relying on ground-truth implementations or test cases, and interpretable evaluation results.To address these challenges, we introduce WebCoderBench, the first realworld, generalized, and interpretable benchmark for web app generation.WebCoderBench comprises 1,572 real user requirements, covering diverse modalities and expression styles that reflect realistic user intentions.Web-CoderBench provides 24 fine-grained evaluation metrics across 9 perspectives, combining the rule-based and LLM-as-a-judge paradigms for fully automated, objective, and general evaluation.Moreover, WebCoderBench adopts human-preference-aligned weights over metrics to yield interpretable overall scores.Experiments across 12 representative LLMs and 2 LLM-based agents show that there exists no dominant model across all evaluation metrics, offering an opportunity for LLM developers to optimize their models in a targeted manner for a more powerful version. Yingjie Fu, Wei Yang 0013, Ying Zhang 0012, Tao Xie 0001 |
ACL (1) | 3 |
| 2026 | PARD: Enhancing Goodput for Inference Pipeline via Proactive Request DroppingabstractModern deep neural network (DNN) and large language model (LLM) applications integrate multiple models into inference pipelines with stringent latency requirements for customized tasks. To mitigate extensive request timeouts caused by accumulation, systems for inference pipelines commonly drop a subset of requests so the remaining ones can satisfy latency constraints. Since it is commonly believed that request dropping adversely affects goodput, existing systems only drop requests when they have to, which we call reactive dropping. However, this reactive policy can not maintain high goodput, as it neither makes timely dropping decisions nor identifies the proper set of requests to drop, leading to issues of dropping requests too late or dropping the wrong set of requests. Yitao Hu, Mingfang Ji, Wei Yang 0013, Yuhao Zhang 0006, Laiping Zhao, Wenxin Li 0001, Xiulong Liu 0001, Wenyu Qu, Hao Wang 0022 |
EuroSys | 5 |
| 2026 | From User Operations to Agentic Automation: Toward Intent-Oriented Software in the LLM Era
Tao Xie 0001, Dezhi Ran, Mengzhou Wu, Yuzhe Guo, Wei Yang 0013 |
J. Comput. Sci. Technol. | 6 |
| 2026 | Judge: Effective State Abstraction for Guiding Automated Web GUI TestingabstractAutomated web GUI testing approaches aim to maximize the code coverage of a web app within a specific time budget. However, due to the highly dynamic characteristics of web apps, testing approaches often get stuck in loops or repeatedly explore the same app areas. To address this issue, existing approaches conduct state abstraction, grouping similar pages into the same state in an effort to approximate the ideal state (i.e., a state that encompasses all-and-only those pages exhibiting the same behavior from a testing perspective) to reduce repetitive explorations. Typically, these approaches rely on the Document Object Model (DOM) or visual similarity, using predefined thresholds or learning-based classifiers to determine which pages should belong to the same state. However, pages within the same ideal state still exhibit discrepancies, caused by factors such as dynamically loaded data and dynamically expanded UI elements. The varying page complexities and design styles among apps bring even more challenges. These phenomena present substantial obstacles to existing approaches in determining desirable classification thresholds or training desirable classifiers, preventing them from conducting satisfactory state abstraction to guide the testing process. To address the preceding challenges, in this article, we propose Judge, a novel approach based on structure merging and contrastive learning for state abstraction. Judge includes a “merge-and-classify” strategy. In the “merge” phase, Judge iterates through the DOM tree of each given page and merges web element siblings that share the same subtree structure into a single one to abstract and simplify the page, while discarding text contents and HTML attributes of web elements in the process. In this way, Judge mitigates the negative effects introduced by dynamically loaded data and dynamically expanded UI elements, substantially reducing discrepancies between pages in the same ideal state. In the “classify” phase, Judge uses a dedicated contrastive learning model to embed simplified page DOMs into vectors and further conducts classification with a Support Vector Machine (SVM), enabling classification in high-dimensional vector space and improving generalizability across diverse web apps. We evaluate Judge against 13 widely used baseline approaches. The results highlight that Judge outperforms these baseline approaches in classifying page pairs, with an average margin ranging from 8.95% to 28.90% in the F1 score across three manually labeled datasets. Additionally, when compared to the five most effective baseline approaches, Judge demonstrates superiority in guiding the exploration of automated web GUI testing in six widely studied apps, with code coverage improved by an average of 2.62–14.12%. The code and data of Judge are publicly accessible. Junheng Wang, Wei Yang 0013, Ying Zhang 0012, Tao Xie 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | TaOPT: Tool-Agnostic Optimization of Parallelized Automated Mobile UI TestingabstractThe emergence of modern testing clouds, equipped with a vast array of real testing devices and high-fidelity emulators, has significantly increased the need for parallel automated mobile testing to optimally utilize the resources of testing clouds. Parallel testing aligns perfectly with the characteristic of rapid iteration cycles for mobile app development, where testing time is limited. While numerous tools have been proposed for optimizing the testing effectiveness on a single testing device, it remains an open problem to optimize the parallelization of automated mobile UI testing in terms of resource and time utilization. To optimize the parallelization of automated mobile UI testing, in this paper, we propose TaOPT, a fully automated, tool-agnostic approach, which improves the parallelization effectiveness of any given testing tool without modifying the tool's internal workflow. In particular, TaOPT conducts online analysis to infer loosely coupled UI subspaces in the App Under Test (AUT). TaOPT then manages access to these subspaces across various testing devices, guiding automated UI testing toward distinct subspaces on different devices without knowing the testing tool's internal workflow. We apply TaOPT on 18 highly popular mobile apps with three state-of-the-art automated UI testing tools for Android. Evaluation results show that TaOPT helps the tools reach comparable code coverage using 60% less testing duration and 62% less machine time than the baseline on average. In addition, TaOPT consistently enhances automated UI testing tools to detect 1.2 to 2.1 times more unique crashes given the same testing resources. Dezhi Ran, Wei Yang 0013, Tao Xie 0001 |
ASPLOS (2) | 4 |
| 2025 | RVISmith: Fuzzing Compilers for RVV IntrinsicsabstractModern processors are equipped with single instruction multiple data (SIMD) instructions for fine-grained data parallelism. Compiler auto-vectorization techniques that target SIMD instructions face performance limitations due to insufficient information available at compile time, requiring programmers to manually manipulate SIMD instructions. SIMD intrinsics, a type of built-in function provided by modern compilers, enable programmers to manipulate SIMD instructions within high-level programming languages. Bugs in compilers for SIMD intrinsics can introduce potential threats to software security, producing unintended calculation results, data loss, program crashes, etc. Yibo He, Cunjian Huang, Xianmiao Qu, Hongdeng Chen, Wei Yang 0013, Tao Xie 0001 |
CCS | 5 |
| 2025 | Automatic Mathematic In-Context Example Generation for LLM Using Multi-Modal ConsistencyabstractLarge Language Models (LLMs) have advanced Natural Language Processing (NLP) tasks but are limited in mathematical reasoning. To address this, few-shot examples are used in prompts for in-context learning. However, existing methods require annotated datasets, resulting in higher computational costs and lower quality examples. To mitigate these limitations, we propose AutoMathIC, a framework that automatically generates high-quality in-context examples to enhance LLMs’ mathematical reasoning. AutoMathIC ensures consistency across different modalities (e.g., Chain-of-Thought (CoT), code snippets, and equations) by generating and selecting mutations that improve response consistency. Evaluated on four math problem datasets, AutoMathIC outperforms six baselines, with LLM accuracy ranging from 87.0% to 99.3% for GPT-3.5 and 93.1% to 98.7% for GPT-4o-mini. It surpasses the state-of-the-art in-context example retrieval method in three of the four datasets by 0.3% to 11.8%, without relying on an annotated dataset. Wei Yang 0013, Gopal Gupta 0001, Shiyi Wei |
COLING | 2 |
| 2025 | StdGEN: Semantic-Decomposed 3D Character Generation from Single ImagesabstractWe present StdGEN, an innovative pipeline for generating semantically decomposed high-quality 3D characters from single images, enabling broad applications in virtual reality, gaming, and filmmaking, etc. Unlike previous methods which struggle with limited decomposability, unsatisfactory quality, and long optimization times, StdGEN features decomposability, effectiveness and efficiency; i.e., it generates intricately detailed 3D characters with separated semantic components such as the body, clothes, and hair, in three minutes. At the core of StdGEN is our proposed Semantic-aware Large Reconstruction Model (S-LRM), a transformer-based generalizable model that jointly reconstructs geometry, color and semantics from multi-view images in a feed-forward manner. A differentiable multilayer semantic surface extraction scheme is introduced to acquire meshes from hybrid implicit fields reconstructed by our S-LRM. Additionally, a specialized efficient multi-view diffusion model and an iterative multi-layer surface refinement module are integrated into the pipeline to facilitate high-quality, decomposable 3D character generation. Extensive experiments demonstrate our state-of-theart performance in 3D anime character generation, surpassing existing baselines by a significant margin in geometry, texture and decomposability. StdGEN offers ready-touse semantic-decomposed 3D characters and enables flexible customization for a wide range of applications. Project page: https://stdgen.github.io Yanning Zhou 0003, Wang Zhao 0001, Zhongkai Wu, Kaiwen Xiao, Wei Yang 0013, Yong-Jin Liu 0001, Xiao Han 0011 |
CVPR | 6 |
| 2025 | WAF: An Efficient WebAssembly-Based Execution Environment for User-Defined FunctionsabstractUser-Defined Functions (UDFs) have long served as the standard method for extending the capabilities of data management systems. With the advent of WebAssembly (WASM), UDFs' dependencies, such as language runtimes and libraries, can be compiled into a WASM module, which is then instantiated to execute the UDF. This approach offers several key advantages: 1) it allows developers to write UDFs in their preferred programming language, rather than being limited to those natively supported by the database engine; 2) it isolates UDFs' dependencies within the WASM module, mitigating the risk of errors caused by conflicting dependencies on the same host; and 3) it promotes cross-platform compatibility, enabling seamless execution of UDFs across different engines, operating systems, and architectures. However, our analysis reveals that executing a WASM-based UDF incurs overhead due to data transfer between the database engine and the WASM runtime. This process involves data copying and data layout adjustments, which can significantly impact performance. To address these challenges, we present WAF, a WASM-based UDF execution environment. WAF leverages shared memory to eliminate data copying and shifts data layout adjustments from the execution phase to the compilation phase. Experimental results show that WAF reduces the execution overhead of WASM-based UDFs by 3.1x and achieves an 18.1x speedup compared to the container-based approach, eliminating nearly all data transfer delays. Hao Fan 0006, Junhui Peng, Song Wu 0001, Chen Yu 0003, Hai Jin 0001, Wei Yang 0013 |
ICDE | 9 |
| 2025 | CodeImprove: Program Adaptation for Deep Code ModelsabstractLeveraging deep learning (DL)-based code analysis tools to solve software engineering tasks is becoming increasingly popular. Code models often suffer performance degradation due to various reasons (e.g., code data shifts). Retraining is often required to address these issues, but frequent model updates are costly in labeling and deployment. In this paper, we explore an alternative solution: Adapting the program inputs to the code models. This can be achieved by two steps: 1) input validation that focuses on identifying whether an input is an out-of-scope input program that are beyond a model's handling capability, and 2) input adaptation that adapts out-of-scope inputs to become in-scope inputs. Validating program input is challenging, as current techniques focus on continuous inputs such as image data and fail with discrete inputs like code data, which have unique characteristics and are processed differently by deep learning models. Adapting out-of-scope programs is also challenging due to their vast search spaces. Therefore, in this paper, we propose CodeImprove, which distinguishes out-of-scope from normal inputs and converts such out-of-scope inputs back to in-scope inputs through program transformation. In particular, we propose a validity score metric to identify out-of-scope inputs and leverage genetics algorithms to apply semantic preserving program transformation to convert out-ofscope inputs to in-scope inputs. Our experimental results show CodeImprove can enhance upto 8.78% of accuracy, and 51.28% of relative improvements in three code models on two SE tasks. Additionally, our input validation is promising in detecting out-of-scope inputs (AUC score of 0.924). Ravishka Rathnasuriya, Wei Yang 0013 |
ICSE | 3 |
| 2025 | Element-Aware Fine-Tuning of Vision-Language Models for Cost-Efficient GUI Testing in an Industrial SettingabstractUser Interface (UI) testing is crucial for quality assurance of industrial mobile applications, and yet it remains labor-intensive and challenging to automate effectively. Recent advances in Vision-Language Models (VLMs) present a promising solution for automating GUI testing by mapping natural language instructions to pixel-level actions, significantly reducing the manual effort required for writing test scripts and even designing test cases. While numerous VLMs have been proposed and evaluated for GUI testing, they often fail to meet two critical industrial requirements: (1) effectiveness when handling complex, multi-step workflows in industrial applications, and (2) efficiency for large-scale, high-frequency testing environments typical in industrial settings. Toward addressing the preceding industrial requirements, in this paper, we report our experiences in developing and deploying RePeek, a novel approach employing a unified three-stage pipeline for both training and inference, enables a VLM to explicitly detect and reason over discrete GUI elements, thereby overcoming the limitations of pixel-based reasoning for both efficiency and effectiveness improvements. In the first stage, RePeek integrates a lightweight UI-element detector named OmniParser to decompose UI screenshots into a structured element list. In the second stage, RePeek adopts the vision encoder of the VLM to generate the embedding for each element. In the third stage, RePeek fuses these element embeddings with the textual instruction to reason and perform classification directly on the UI elements, empowering efficient small models to achieve superior performance against expensive large models. Comprehensive evaluations on public benchmarks and deployment at WeChat show that RePeek consistently achieves superior accuracy and efficiency compared to state-of-the-art VLMs. Specifically, RePeek enables a fine-tuned Qwen2.5-VL-3B model to outperform a 72B model with 75% less training data, validating the effectiveness of incorporating domain knowledge into VLM-based GUI testing. We conclude by summarizing three key lessons from developing and deploying RePeek, offering insights for both researchers and practitioners working on industrial-strength UI testing. Mengzhou Wu, Yuzhe Guo, Haochuan Lu, Xia Zeng, Liangchao Yao, Yuetang Deng, Dezhi Ran, Wei Yang 0013, Tao Xie 0001 |
ASE | 10 |
| 2025 | SoK: Efficiency Robustness of Dynamic Deep Learning Systems
Ravishka Rathnasuriya, Tingxi Li, Zexin Xu, Mirazul Haque, Wei Yang 0013 |
USENIX Security Symposium | 7 |
| 2025 | DPEfficR: a data and parameter efficient approach for training neural API recommendation model
Xiaohong Han, Xiaoning Feng, Guangzhao Sun, Wei Yang 0013 |
Autom. Softw. Eng. | 6 |
| 2025 | MalScan: Android Malware Detection Based on Social-Network Centrality AnalysisabstractMalware scanning of an app market is expected to be scalable and effective. However, existing approaches use syntax-based features that can be evaded by transformation attacks or semantic-based features which are usually extracted by expensive program analysis. Therefore, to address the scalability challenges of traditional heavyweight static analysis, we propose a graph-based lightweight approachMalScanfor Android malware detection.MalScanconsiders the function call graph as a complex social network and employs centrality analysis on sensitiveapplication program interfaces(APIs) to express the semantic characteristics of the graph. On this basis, machine learning algorithms and ensemble learning algorithms are applied to classify the extracted features. We evaluateMalScanon datasets of 104,892 benign apps and 108,640 malwares, and the results of experiments indicate thatMalScanoutperforms six state-of-the-art detectors and can quickly detect Android malware with an f-value as high as 99%. In addition, there are also significant improvements in the robustness of Android app evolution and robustness to obfuscation. Finally, we conduct an exhaustive statistical study of over one million applications in the Google-Play app market and successfully identify 498 zero-day malware, which further validates the feasibility ofMalScanon market-wide malware scanning. Yueming Wu 0001, Wenqi Suo, Siyue Feng, Deqing Zou, Wei Yang 0013, Yang Liu 0003, Hai Jin 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | Foundation Model Engineering: Engineering Foundation Models Just as Engineering SoftwareabstractBy treating data and models as source code, Foundation Models (FMs) become a new type of software. Mirroring the concept of software crisis, the increasing complexity of FMs makes FM crisis a tangible concern in the coming decade, appealing for new theories and methodologies from the field of software engineering. In this article, we outline our vision of introducing FM engineering, a strategic response to the anticipated FM crisis with principled engineering methodologies. FM engineering aims to mitigate potential issues in FM development and application through the introduction of declarative, automated, and unified programming interfaces for both data and model management, reducing the complexities involved in working with FMs by providing a more structured and intuitive process for developers. Through the establishment of FM engineering, we aim to provide a robust, automated, and extensible framework that addresses the imminent challenges, and discover new research opportunities for the software engineering field. Dezhi Ran, Mengzhou Wu, Wei Yang 0013, Tao Xie 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | LogLabeler: Towards Effective Acquisition of Log Labels in Industrial Log-Based AnalysisabstractLog-based AIOps is a widely researched topic aiming at reducing the developer burden in system maintenance. Since industrial developers prefer lightweight supervised solutions for log-based AIOps, the strong dependence of these solutions on labeled data creates significant challenges for teams new to building log-based AIOps capabilities, such as high labeling costs, inconsistent annotations, and manual management issues. Log-labeling faces challenges in integrating existing artifacts to reduce labeling costs and manage labels effectively. To the best of our knowledge, no prior research addresses of assisting log-labeling problem. In this article, we propose a new approach called LogLabeler to assist developers in annotating and managing log labels. LogLabeler leverages existing artifacts for initial-label-acquisition, minimizes labeling costs by automatically generating all log labels, and shields developers from manual label management through a human-in-the-loop refinement approach. Evaluations on real-world datasets from Alibaba and open-source datasets show that LogLabeler can effectively supplement log labels, achieving comparable accuracy to existing baselines while operating more efficiently. Furthermore, we demonstrate LogLabeler's practical effectiveness at Alibaba through a case study, highlighting its benefits to developers. Zongyang Li, Qinglong Wang 0003, Shangming Cai, Zheng Liu 0022, Tao Ma 0006, Wei Yang 0013, Ying Li 0012, Tao Xie 0001 |
IEEE Trans. Serv. Comput. | 6 |
| 2024 | MENDNet: Just-in-time Fault Detection and Mitigation in AI Systems with Uncertainty Quantification and Multi-Exit NetworksabstractHardware faults in AI accelerators, particularly in accelerator memory, can alter pre-trained deep neural network parameters, leading to errors that compromise performance. To address this, just-intime (JIT) fault detection and mitigation are crucial. However, existing fault detection/mitigation approaches, either interrupt continuous execution or introduce significant latency, making them less ideal for JIT implementation. To circumvent this issue, this paper explores uncertainty quantification in deep neural networks as a means of facilitating an efficient and novel fault detection approach in AI accelerators. Furthermore, in order to mitigate the impact of such faults, we propose MENDNet, which leverages the properties of multi-exit neural networks, coupled with the proposed uncertainty quantification framework. By tuning the confidence threshold for inference in each exit and leveraging the energy-based uncertainty quantification metric, MENDNet can make accurate predictions even in the presence of faults in the accelerator. When evaluated on state-of-the-art network-dataset configurations and with multiple fault rate-fault position combinations, our proposed approach furnishes up to 80.42% improvement in accuracy over a traditional DNN implementation, thereby instilling the reliability of the AI accelerator in mission mode. Shamik Kundu, Mirazul Haque, Sanjay Das, Wei Yang 0013, Kanad Basu |
DAC | 4 |
| 2024 | Guardian: A Runtime Framework for LLM-Based UI ExplorationabstractTests for feature-based UI testing have been indispensable for ensuring the quality of mobile applications (apps for short). The high manual labor costs to create such tests have led to a strong interest in automated feature-based UI testing, where an approach automatically explores the App under Test (AUT) to find correct sequences of UI events achieving the target test objective, given only a high-level test objective description. Given that the task of automated feature-based UI testing resembles conventional AI planning problems, large language models (LLMs), known for their effectiveness in AI planning, could be ideal for this task. However, our study reveals that LLMs struggle with following specific instructions for UI testing and replanning based on new information. This limitation results in reduced effectiveness of LLM-driven solutions for automated feature-based UI testing, despite the use of advanced prompting techniques. Toward addressing the preceding limitation, we propose Guardian, a runtime system framework to improve the effectiveness of automated feature-based UI testing by offloading computational tasks from LLMs with two major strategies. First, Guardian refines UI action space that the LLM can plan over, enforcing the instruction following of the LLM by construction. Second, Guardian deliberately checks whether the gradually enriched information invalidates previous planning by the LLM. Guardian removes the invalidated UI actions from the UI action space that the LLM can plan over, restores the state of the AUT to the state before the execution of the invalidated UI actions, and prompts the LLM to re-plan with the new UI action space. We instantiate Guardian with ChatGPT and construct a benchmark named FestiVal with 58 tasks from 23 highly popular apps. Evaluation results on FestiVal show that Guardian achieves 48.3 Dezhi Ran, Hao Wang 0112, Mengzhou Wu, Ying Zhang 0012, Wei Yang 0013, Tao Xie 0001 |
ISSTA | 7 |
| 2024 | WEFix: Intelligent Automatic Generation of Explicit Waits for Efficient Web End-to-End Flaky TestsabstractWeb end-to-end (e2e) testing evaluates the workflow of a web application. It simulates real-world user scenarios to ensure the application flows behave as expected. However, web e2e tests are notorious for being flaky, i.e., the tests can produce inconsistent results despite no changes to the code. One common type of flakiness is caused by nondeterministic execution orders between the test code and the client-side code under test. In particular, UI-based flakiness emerges as a notably prevalent and challenging issue to fix because the test code has limited knowledge about the client-side code execution. In this paper, we propose WEFix, a technique that can automatically generate fix code for UI-based flakiness in web e2e testing. The core of our approach is to leverage browser UI changes to predict the client-side code execution and generate proper wait oracles. We evaluate the effectiveness and efficiency of WEFix against 122 web e2e flaky tests from seven popular real-world projects. Our results show that WEFix dramatically reduces the overhead (from 3.7$\times$ to 1.25$\times$) while achieving a high correctness (98%). Xinyue Liu 0005, Weike Fang, Wei Yang 0013, Weihang Wang 0001 |
WWW | 4 |
| 2024 | LLMEffiChecker: Understanding and Testing Efficiency Degradation of Large Language ModelsabstractLarge Language Models (LLMs) have received much recent attention due to their human-level accuracy. While existing works mostly focus on either improving accuracy or testing accuracy robustness, the computation efficiency of LLMs, which is of paramount importance due to often vast generation demands and real-time requirements, has surprisingly received little attention. In this article, we make the first attempt to understand and test potential computation efficiency robustness in state-of-the-art LLMs. By analyzing the working mechanism and implementation of 20,543 public-accessible LLMs, we observe a fundamental property in LLMs that could be manipulated in an adversarial manner to reduce computation efficiency significantly. Our interesting observation is that the output length determines the computation efficiency of LLMs instead of the input, where the output length depends on two factors: an often sufficiently large yet pessimistic pre-configured threshold controlling the max number of iterations and a runtime-generated end of sentence (EOS) token. Our key motivation is to generate test inputs that could sufficiently delay the generation of EOS such that LLMs would have to go through enough iterations to satisfy the pre-configured threshold. We present LLMEffiChecker , which can work under both white-box setting and black-box setting. In the white-box scenario, LLMEffiChecker develops a gradient-guided technique that searches for a minimal and unnoticeable perturbation at character-level, token-level, and structure-level. In the black-box scenario, LLMEffiChecker employs a causal inference-based approach to find critical tokens and similarly applies three levels of imperceptible perturbation to them. Both the white-box and black-box settings effectively delay the appearance of EOS, compelling these inputs to reach the naturally unreachable threshold. To demonstrate the effectiveness of LLMEffiChecker , we conduct a systematic evaluation on nine publicly available LLMs: Google T5, AllenAI WMT14, Helsinki-NLP translator, Facebook FairSeq, UNICAMP-DL translator, MarianMT, Google FLAN-T5, MBZUAI LaMini-GPT, and Salesforce CodeGen. Experimental results show that LLMEffiChecker can increase on average LLMs’ response latency and energy consumption by 325% to 3,244% and 344% to 3,616%, respectively, by perturbing just one character or token in the input sentence. Our case study shows that inputs generated by LLMEffiChecker significantly affect the battery power in real-world mobile devices (i.e., drain more than 30 times battery power than normal inputs). Xiaoning Feng, Xiaohong Han, Wei Yang 0013 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | Automated Testing Linguistic Capabilities of NLP ModelsabstractNatural language processing (NLP) has gained widespread adoption in the development of real-world applications. However, the black-box nature of neural networks in NLP applications poses a challenge when evaluating their performance, let alone ensuring it. Recent research has proposed testing techniques to enhance the trustworthiness of NLP-based applications. However, most existing works use a single, aggregated metric (i.e., accuracy) which is difficult for users to assess NLP model performance on fine-grained aspects, such as LCs. To address this limitation, we present ALiCT, an automated testing technique for validating NLP applications based on their LCs. ALiCT takes user-specified LCs as inputs and produces diverse test suite with test oracles for each of given LC. We evaluate ALiCT on two widely adopted NLP tasks, sentiment analysis and hate speech detection, in terms of diversity, effectiveness, and consistency. Using Self-BLEU and syntactic diversity metrics, our findings reveal that ALiCT generates test cases that are 190% and 2213% more diverse in semantics and syntax, respectively, compared to those produced by state-of-the-art techniques. In addition, ALiCT is capable of producing a larger number of NLP model failures in 22 out of 25 LCs over the two NLP applications. Austin Mordahl, Cong Liu 0005, Wei Yang 0013, Shiyi Wei |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2023 | Dynamic Transformers Provide a False Sense of EfficiencyabstractDespite much success in natural language processing (NLP), pre-trained language models typically lead to a high computational cost during inference.Multi-exit is a mainstream approach to address this issue by making a tradeoff between efficiency and accuracy, where the saving of computation comes from an early exit.However, whether such saving from earlyexiting is robust remains unknown.Motivated by this, we first show that directly adapting existing adversarial attack approaches targeting model accuracy cannot significantly reduce inference efficiency.To this end, we propose a simple yet effective attacking framework, SAME, a novel slowdown attack framework on multi-exit models, which is specially tailored to reduce the efficiency of the multi-exit models.By leveraging the multi-exit models' design characteristics, we utilize all internal predictions to guide the adversarial sample generation instead of merely considering the final prediction.Experiments on the GLUE benchmark show that SAME can effectively diminish the efficiency gain of various multi-exit models by 80% on average, convincingly validating its effectiveness and generalization ability. 1 Yiming Chen 0010, Zexin Li 0001, Wei Yang 0013, Cong Liu 0005, Robby T. Tan, Haizhou Li 0001 |
ACL (1) | 4 |
| 2023 | The Dark Side of Dynamic Routing Neural Networks: Towards Efficiency Backdoor InjectionabstractRecent advancements in deploying deep neural networks (DNNs) on resource-constrained devices have generated interest in input-adaptive dynamic neural networks (DyNNs). DyNNs offer more efficient inferences and enable the deployment of DNNs on devices with limited resources, such as mobile devices. However, we have discovered a new vulnerability in DyNNs that could potentially compromise their efficiency. Specifically, we investigate whether adversaries can manipulate DyNNs' computational costs to create a false sense of efficiency. To address this question, we propose EfficFrog, an adversarial attack that injects universal efficiency backdoors in DyNNs. To inject a backdoor trigger into DyNNs, EfficFrog poisons only a minimal percentage of the DyNNs' training data. During the inference phase, EfficFrog can slow down the backdoored DyNNs and abuse the computational resources of systems running DyNNs by adding the trigger to any input. To evaluate EfficFrog, we tested it on three DNN backbone architectures (based on VGG16, MobileNet, and ResNet56) using two popular datasets (CIFAR-10 and Tiny ImageNet). Our results demonstrate that EfficFrog reduces the efficiency of DyNNs on triggered input samples while keeping the efficiency of clean samples almost the same. Mirazul Haque, Cong Liu 0005, Wei Yang 0013 |
CVPR | 5 |
| 2023 | SlothSpeech: Denial-of-service Attack Against Speech Recognition Models
Mirazul Haque, Rutvij Shah, Berrak Sisman, Cong Liu 0005, Wei Yang 0013 |
INTERSPEECH | 6 |
| 2023 | DyCL: Dynamic Neural Network Compilation Via Program Rewriting and Graph OptimizationabstractThe deep learning (DL) compiler serves as a vital infrastructure component to enable the deployment of deep neural networks on diverse hardware platforms such as mobile devices and Raspberry Pi. DL compiler’s primary function is to translate DNN programs written in high-level DL frameworks such as PyTorch and TensorFlow into portable executables. These executables can then be flexibly executed by the deployed host programs. However, existing DL compilers rely on a tracing mechanism, which involves feeding a runtime input to a neural network program and tracing the program execution paths to generate the computational graph necessary for compilation. Unfortunately, this mechanism falls short when dealing with modern dynamic neural networks (DyNNs) that possess varying computational graphs depending on the inputs. Consequently, conventional DL compilers struggle to accurately compile DyNNs into executable code. To address this limitation, we propose DyCL, a general approach that enables any existing DL compiler to successfully compile DyNNs. DyCL tackles the dynamic nature of DyNNs by introducing a compilation mechanism that redistributes the control and data flow of the original DNN programs during the compilation process. Specifically, DyCL develops program analysis and program transformation techniques to convert a dynamic neural network into multiple sub-neural networks. Each sub-neural network is devoid of conditional statements and is compiled independently. Furthermore, DyCL synthesizes a host module that models the control flow of the DyNNs and facilitates the invocation of the sub-neural networks. Our evaluation demonstrates the effectiveness of DyCL, achieving a 100% success rate in compiling all dynamic neural networks. Moreover, the compiled executables generated by DyCL exhibit significantly improved performance, running between 1.12× and 20.21× faster than the original DyNNs executed on general-purpose DL frameworks. Shiyi Wei, Cong Liu 0005, Wei Yang 0013 |
ISSTA | 4 |
| 2023 | RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language ModelsabstractRecent advancements in language models (LMs) have gained substantial attentions on their capability to generate human-like responses. Though exhibiting a promising future for various applications such as conversation AI, these LMs face deployment challenges on various devices due to their extreme computational cost and unpredictable inference latency. Such varied inference latency, identified as a consequence of uncertainty intrinsic to the nature of language, can lead to computational inefficiency and degrade the overall performance of LMs, especially under high-traffic workloads. Unfortunately, the bandwidth of these uncertainty sources is extensive, complicating the prediction of latency and the effects emanating from such uncertainties. To understand and mitigate the impact of uncertainty on real-time response-demanding systems, we take the first step to comprehend, quantify and optimize these uncertainty-induced latency performance variations in LMs. Specifically, we present RT-LM, an uncertainty-aware resource management ecosystem for real-time inference of LMs. RT-LM innovatively quantifies how specific input uncertainties, recognized within the NLP community, adversely affect latency, often leading to an increased output length. Exploiting these insights, we devise a lightweight yet effective method to dynamically correlate input text uncertainties with output length at runtime. Utilizing this quantification as a latency heuristic, we integrate the uncertainty information into a system-level scheduler which explores several uncertainty-induced optimization opportunities, including uncertainty-aware prioritization, dynamic consolidation, and strategic CPU offloading. Quantitative experiments across five state-of-the-art LMs on two hardware platforms demonstrates that RT-LM can significantly reduce the average response time and improve throughput while incurring a rather small runtime overhead. Yufei Li 0001, Zexin Li 0001, Wei Yang 0013, Cong Liu 0005 |
RTSS | 3 |
| 2022 | TestAug: A Framework for Augmenting Capability-based NLP TestsabstractThe recently proposed capability-based NLP testing allows model developers to test the functional capabilities of NLP models, revealing functional failures for models with good held-out evaluation scores. However, existing work on capability-based testing requires the developer to compose each individual test template from scratch. Such approach thus requires extensive manual efforts and is less scalable. In this paper, we investigate a different approach that requires the developer to only annotate a few test templates, while leveraging the GPT-3 engine to generate the majority of test cases. While our approach saves the manual efforts by design, it guarantees the correctness of the generated suites with a validity checker. Moreover, our experimental results show that the test suites generated by GPT-3 are more diverse than the manually created ones; they can also be used to detect more errors compared to manually created counterparts. Our test suites can be downloaded at https://anonymous-researcher-nlp.github.io/testaug/. Guanqun Yang, Mirazul Haque, Qiaochu Song, Wei Yang 0013, Xueqing Liu 0001 |
COLING | 4 |
| 2022 | NICGSlowDown: Evaluating the Efficiency Robustness of Neural Image Caption Generation ModelsabstractNeural image caption generation (NICG) models have received massive attention from the research community due to their excellent performance in visual understanding. Existing work focuses on improving NICG model ac-curacy while efficiency is less explored. However, many real-world applications require real-time feedback, which highly relies on the efficiency of NICG models. Recent re-search observed that the efficiency of NICG models could vary for different inputs. This observation brings in a new attack surface of NICG models, i.e., An adversary might be able to slightly change inputs to cause the NICG mod-els to consume more computational resources. To further understand such efficiency-oriented threats, we propose a new attack approach, NICGSlowDown, to evaluate the ef-ficiency robustness of NICG models. Our experimental re-sults show that NICGSlowDown can generate images with human-unnoticeable perturbations that will increase the NICG model latency up to 483.86%. We hope this research could raise the community's concern about the efficiency robustness of NICG models. Mirazul Haque, Cong Liu 0005, Wei Yang 0013 |
CVPR | 5 |
| 2022 | EREBA: Black-box Energy Testing of Adaptive Neural NetworksabstractRecently, various Deep Neural Network (DNN) models have been proposed for environments like embedded systems with stringent energy constraints. The fundamental problem of determining the robustness of a DNN with respect to its energy consumption (energy robustness) is relatively unexplored compared to accuracy-based robustness. This work investigates the energy robustness of Adaptive Neural Networks (AdNNs), a type of energy-saving DNNs proposed for many energy-sensitive domains and have recently gained traction. We propose EREBA, the first black-box testing method for determining the energy robustness of an AdNN. EREBA explores and infers the relationship between inputs and the energy consumption of AdNNs to generate energy surging samples. Extensive implementation and evaluation using three state-of-the-art AdNNs demonstrate that test inputs generated by EREBA could degrade the performance of the system substantially. The test inputs generated by EREBA can increase the energy consumption of AdNNs by 2,000% compared to the original inputs. Our results also show that test inputs generated via EREBA are valuable in detecting energy surging inputs. Mirazul Haque, Yaswanth Yadlapalli, Wei Yang 0013, Cong Liu 0005 |
ICSE | 3 |
| 2022 | VulCNN: An Image-inspired Scalable Vulnerability Detection SystemabstractSince deep learning (DL) can automatically learn features from source code, it has been widely used to detect source code vulnerability. To achieve scalable vulnerability scanning, some prior studies intend to process the source code directly by treating them as text. To achieve accurate vulnerability detection, other approaches consider distilling the program semantics into graph representations and using them to detect vulnerability. In practice, text-based techniques are scalable but not accurate due to the lack of program semantics. Graph-based methods are accurate but not scalable since graph analysis is typically time-consuming. Yueming Wu 0001, Deqing Zou, Shihan Dou, Wei Yang 0013, Hai Jin 0001 |
ICSE | 4 |
| 2022 | Learn to Reverse DNNs from AI Programs AutomaticallyabstractWith the privatization deployment of DNNs on edge devices, the security of on-device DNNs has raised significant concern. To quantify the model leakage risk of on-device DNNs automatically, we propose NNReverse, the first learning-based method which can reverse DNNs from AI programs without domain knowledge. NNReverse trains a representation model to represent the semantics of binary code for DNN layers. By searching the most similar function in our database, NNReverse infers the layer type of a given function’s binary code. To represent assembly instructions semantics precisely, NNReverse proposes a more fine-grained embedding model to represent the textual and structural-semantic of assembly functions. Hamed Khanpour, Cong Liu 0005, Wei Yang 0013 |
IJCAI | 4 |
| 2022 | An Empirical Analysis of Compatibility Issues for Industrial Mobile Games (Practical Experience Report)abstractDetecting and fixing compatibility issues become increasingly important for mobile game development. The constant evolution of mobile operating systems and the severe fragmentation of mobile devices makes it challenging for game developers to detect and fix compatibility issues in time for various device models. The undetected compatibility issues can ruin the user experience, and cause financial loss to game companies and players. Unfortunately, up to the present, mobile game testing is still rather challenging in general. The pressing compatibility issue of mobile games is largely untouched in the research community so far. To bridge the gap, in this experience paper, we perform an empirical study on common compatibility issues of popular commercial mobile games. In particular, we select four active and representative mobile games with well-documented bug reports, containing over seven million lines of code and over 20,000 commits over the past several years. We successfully create a dataset with complete information about bugs and bug fixing details, to investigate the common compatibility issues and fixing strategies. We performed an in-depth manual inspection of the most common symptoms and root causes of these compatibility issues, and analyzed the common fixing strategies of issues under each root cause category. We believe our findings and implications are useful for developers in addressing compatibility hurdles during the developing process. Our results also provide insights for future research on compatibility issue testing and bug fixing for mobile games. Lei Ma 0003, Shangjie Lu, Changjie Fan, Wei Yang 0013 |
ISSRE | 7 |
| 2022 | DeepPerform: An Efficient Approach for Performance Testing of Resource-Constrained Neural NetworksabstractToday, an increasing number of Adaptive Deep Neural Networks (AdNNs) are being used on resource-constrained embedded devices. We observe that, similar to traditional software, redundant computation exists in AdNNs, resulting in considerable performance degradation. The performance degradation is dependent on the input and is referred to as input-dependent performance bottlenecks (IDPBs). To ensure an AdNN satisfies the performance requirements of resource-constrained applications, it is essential to conduct performance testing to detect IDPBs in the AdNN. Existing neural network testing methods are primarily concerned with correctness testing, which does not involve performance testing. To fill this gap, we propose DeepPerform, a scalable approach to generate test samples to detect the IDPBs in AdNNs. We first demonstrate how the problem of generating performance test samples detecting IDPBs can be formulated as an optimization problem. Following that, we demonstrate how DeepPerform efficiently handles the optimization problem by learning and estimating the distribution of AdNNs’ computational consumption. We evaluate DeepPerform on three widely used datasets against five popular AdNN models. The results show that DeepPerform generates test samples that cause more severe performance degradation (FLOPs: increase up to 552%). Furthermore, DeepPerform is substantially more efficient than the baseline methods in generating test inputs (runtime overhead: only 6–10 milliseconds). Mirazul Haque, Cong Liu 0005, Wei Yang 0013 |
ASE | 4 |
| 2022 | NMTSloth: understanding and testing efficiency degradation of neural machine translation systemsabstractNeural Machine Translation (NMT) systems have received much recent attention due to their human-level accuracy. While existing works mostly focus on either improving accuracy or testing accuracy robustness, the computation efficiency of NMT systems, which is of paramount importance due to often vast translation demands and real-time requirements, has surprisingly received little attention. In this paper, we make the first attempt to understand and test potential computation efficiency robustness in state-of-the-art NMT systems. By analyzing the working mechanism and implementation of 1455 public-accessible NMT systems, we observe a fundamental property in NMT systems that could be manipulated in an adversarial manner to reduce computation efficiency significantly. Our interesting observation is that the output length determines the computation efficiency of NMT systems instead of the input, where the output length depends on two factors: an often sufficiently large yet pessimistic pre-configured threshold controlling the max number of iterations and a runtime generated end of sentence (EOS) token. Our key motivation is to generate test inputs that could sufficiently delay the generation of EOS such that NMT systems would have to go through enough iterations to satisfy the pre-configured threshold. We present NMTSloth, which develops a gradient-guided technique that searches for a minimal and unnoticeable perturbation at character-level, token-level, and structure-level, which sufficiently delays the appearance of EOS and forces these inputs to reach the naturally-unreachable threshold. To demonstrate the effectiveness of NMTSloth, we conduct a systematic evaluation on three public-available NMT systems: Google T5, AllenAI WMT14, and Helsinki-NLP translators. Experimental results show that NMTSloth can increase NMT systems' response latency and energy consumption by 85% to 3153% and 86% to 3052%, respectively, by perturbing just one character or token in the input sentence. Our case study shows that inputs generated by NMTSloth significantly affect the battery power in real-world mobile devices (i.e., drain more than 30 times battery power than normal inputs). Cong Liu 0005, Mirazul Haque, Wei Yang 0013 |
ESEC/SIGSOFT FSE | 5 |
| 2021 | An Empirical Analysis of UI-based Flaky TestsabstractFlaky tests have gained attention from the research community in recent years and with good reason. These tests lead to wasted time and resources, and they reduce the reliability of the test suites and build systems they affect. However, most of the existing work on flaky tests focus exclusively on traditional unit tests. This work ignores UI tests that have larger input spaces and more diverse running conditions than traditional unit tests. In addition, UI tests tend to be more complex and resource-heavy, making them unsuited for detection techniques involving rerunning test suites multiple times. In this paper, we perform a study on flaky UI tests. We analyze 235 flaky UI test samples found in 62 projects from both web and Android environments. We identify the common underlying root causes of flakiness in the UI tests, the strategies used to manifest the flaky behavior, and the fixing strategies used to remedy flaky UI tests. The findings made in this work can provide a foundation for the development of detection and prevention techniques for flakiness arising in UI tests. Alan Romano, Sampath Grandhi, Wei Yang 0013, Weihang Wang 0001 |
ICSE | 4 |
| 2021 | WebEvo: taming web application evolution via detecting semantic structure changesabstractThe development of Web technology and the beginning of the Big Data era have led to the development of technologies for extracting data from websites, such as information retrieval (IR) and robotic process automation (RPA) tools. As websites are constantly evolving, to prevent these tools from functioning improperly due to website evolution, it is important to monitor the changes in websites and report them to the developers and testers. Existing monitoring tools mainly use DOM-tree based techniques to detect changes in the new web pages. However, these monitoring tools incorrectly report content-based changes (i.e., web content refreshed every time a web page is retrieved) as the changes that will adversely affect the performance of the IR and RPA tools. This results in false warnings since the IR and RPA tools typically consider these changes as expected and retrieve dynamic data from them. Moreover, these monitoring tools cannot identify GUI widget evolution (e.g., moving a button), and thus cannot help the IR and RPA tools adapt to the evolved widgets (e.g., automatic repair of locators for the evolved widgets). To address the limitations of the existing monitoring tools, we propose an approach, WebEvo, that leverages historic pages to identify the DOM elements whose changes are content-based changes, which can be safely ignored when reporting changes in the new web pages. Furthermore, to identify refactoring changes that preserve semantics and appearances of GUI widgets, WebEvo adapts computer vision (CV) techniques to identify the mappings of the GUI widgets from the old web page to the new web page on an element-by-element basis. Empirical evaluations on 13 real-world websites from 9 popular categories demonstrate the superiority of WebEvo over the existing DOM-tree based detection or whole-page visual comparison in terms of both effectiveness and efficiency. Fei Shao, Wasif Arman Haque, Jingwei Xu 0004, Ying Zhang 0012, Wei Yang 0013, Yanfang Ye 0001, Xusheng Xiao |
ISSTA | 6 |
| 2021 | HomDroid: detecting Android covert malware by social-network homophily analysisabstractAndroid has become the most popular mobile operating system. Correspondingly, an increasing number of Android malware has been developed and spread to steal users’ private information. There exists one type of malware whose benign behaviors are developed to camouflage malicious behaviors. The malicious component occupies a small part of the entire code of the application (app for short), and the malicious part is strongly coupled with the benign part. In this case, the malware may cause false negatives when malware detectors extract features from the entire apps to conduct classification because the malicious features of these apps may be hidden among benign features. Moreover, some previous work aims to divide the entire app into several parts to discover the malicious part. However, the premise of these methods to commence app partition is that the connections between the normal part and the malicious part are weak (repackaged malware). Yueming Wu 0001, Deqing Zou, Wei Yang 0013, Hai Jin 0001 |
ISSTA | 3 |
| 2021 | GLIB: towards automated test oracle for graphically-rich applicationsabstractGraphically-rich applications such as games are ubiquitous with attractive visual effects of Graphical User Interface (GUI) that offers a bridge between software applications and end-users. However, various types of graphical glitches may arise from such GUI complexity and have become one of the main component of software compatibility issues. Our study on bug reports from game development teams in NetEase Inc. indicates that graphical glitches frequently occur during the GUI rendering and severely degrade the quality of graphically-rich applications such as video games. Existing automated testing techniques for such applications focus mainly on generating various GUI test sequences and check whether the test sequences can cause crashes. These techniques require constant human attention to captures non-crashing bugs such as bugs causing graphical glitches. In this paper, we present the first step in automating the test oracle for detecting non-crashing bugs in graphically-rich applications. Specifically, we propose GLIB based on a code-based data augmentation technique to detect game GUI glitches. We perform an evaluation of GLIB on 20 real-world game apps (with bug reports available) and the result shows that GLIB can achieve 100% precision and 99.5% recall in detecting non-crashing bugs such as game GUI glitches. Practical application of GLIB on another 14 real-world games (without bug reports) further demonstrates that GLIB can effectively uncover GUI glitches, with 48 of 53 bugs reported by GLIB having been confirmed and fixed so far. Ke Chen 0005, Yufei Li 0001, Changjie Fan, Zhipeng Hu, Wei Yang 0013 |
ESEC/SIGSOFT FSE | 6 |
| 2021 | Vet: identifying and avoiding UI exploration tarpitsabstractDespite over a decade of research, it is still challenging for mobile UI testing tools to achieve satisfactory effectiveness, especially on industrial apps with rich features and large code bases. Our experiences suggest that existing mobile UI testing tools are prone to exploration tarpits, where the tools get stuck with a small fraction of app functionalities for an extensive amount of time. For example, a tool logs out an app at early stages without being able to log back in, and since then the tool gets stuck with exploring the app’s pre-login functionalities (i.e., exploration tarpits) instead of its main functionalities. While tool vendors/users can manually hardcode rules for the tools to avoid specific exploration tarpits, these rules can hardly generalize, being fragile in face of diverted testing environments, fast app iterations, and the demand of batch testing product lines. To identify and resolve exploration tarpits, we propose VET, a general approach including a supporting system for the given specific Android UI testing tool on the given specific app under test (AUT). VET runs the tool on the AUT for some time and records UI traces, based on which VET identifies exploration tarpits by recognizing their patterns in the UI traces. VET then pinpoints the actions (e.g., clicking logout) or the screens that lead to or exhibit exploration tarpits. In subsequent test runs, VET guides the testing tool to prevent or recover from exploration tarpits. From our evaluation with state-of-the-art Android UI testing tools on popular industrial apps, VET identifies exploration tarpits that cost up to 98.6% testing time budget. These exploration tarpits reveal not only limitations in UI exploration strategies but also defects in tool implementations. VET automatically addresses the identified exploration tarpits, enabling each evaluated tool to achieve higher code coverage and improve crash-triggering capabilities. Wei Yang 0013, Tianyin Xu, Tao Xie 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2021 | IntDroid: Android Malware Detection Based on API Intimacy AnalysisabstractAndroid, the most popular mobile operating system, has attracted millions of users around the world. Meanwhile, the number of new Android malware instances has grown exponentially in recent years. On the one hand, existing Android malware detection systems have shown that distilling the program semantics into a graph representation and detecting malicious programs by conducting graph matching are able to achieve high accuracy on detecting Android malware. However, these traditional graph-based approaches always perform expensive program analysis and suffer from low scalability on malware detection. On the other hand, because of the high scalability of social network analysis, it has been applied to complete large-scale malware detection. However, the social-network-analysis-based method only considers simple semantic information (i.e., centrality) for achieving market-wide mobile malware scanning, which may limit the detection effectiveness when benign apps show some similar behaviors as malware. In this article, we aim to combine the high accuracy of traditional graph-based method with the high scalability of social-network-analysis--based method for Android malware detection. Instead of using traditional heavyweight static analysis, we treat function call graphs of apps as complex social networks and apply social-network--based centrality analysis to unearth the central nodes within call graphs. After obtaining the central nodes, the average intimacies between sensitive API calls and central nodes are computed to represent the semantic features of the graphs. We implement our approach in a tool called IntDroid and evaluate it on a dataset of 3,988 benign samples and 4,265 malicious samples. Experimental results show that IntDroid is capable of detecting Android malware with an F-measure of 97.1% while maintaining a True-positive Rate of 99.1%. Although the scalability is not as fast as a social-network-analysis--based method (i.e., MalScan ), compared to a traditional graph-based method, IntDroid is more than six times faster than MaMaDroid . Moreover, in a corpus of apps collected from GooglePlay market, IntDroid is able to identify 28 zero-day malware that can evade detection of existing tools, one of which has been downloaded and installed by more than ten million users. This app has also been flagged as malware by six anti-virus scanners in VirusTotal, one of which is Symantec Mobile Insight . Deqing Zou, Yueming Wu 0001, Siru Yang, Anki Chauhan, Wei Yang 0013, Jiangying Zhong, Shihan Dou, Hai Jin 0001 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2020 | ILFO: Adversarial Attack on Adaptive Neural NetworksabstractWith the increasing number of layers and parameters in neural networks, the energy consumption of neural networks has become a great concern to society, especially to users of handheld or embedded devices. In this paper, we investigate the robustness of neural networks against energy-oriented attacks. Specifically, we propose ILFO (Intermediate Output-Based Loss Function Optimization) attack against a common type of energy-saving neural networks, Adaptive Neural Networks (AdNN). AdNNs save energy consumption by dynamically deactivating part of its model based on the need of the inputs. ILFO leverages intermediate output as a proxy to infer the relation between input and its corresponding energy consumption. ILFO has shown an increase up to 100 % of the FLOPs (floating-point operations per second) reduced by AdNNs with minimum noise added to input images. To our knowledge, this is the first attempt to attack the energy consumption of an AdNN. Mirazul Haque, Anki Chauhan, Cong Liu 0005, Wei Yang 0013 |
CVPR | 4 |
| 2020 | Database-Access Performance Antipatterns in Database-Backed Web ApplicationsabstractDatabase-backed web applications are prone to performance bugs related to database accesses. While much work has been conducted on database-access antipatterns with some recent work focusing on performance impact, there still lacks a comprehensive view of database-access performance antipatterns in database-backed web applications. To date, no existing work systematically reports known antipatterns in the literature, and no existing work has studied database-access performance bugs in major types of web applications that access databases differently.To address this issue, we first summarize all known database-access performance antipatterns found through our literature survey, and we report all of them in this paper. We further collect database-access performance bugs from web applications that access databases through language-provided SQL interfaces, which have been largely ignored by recent work, to check how extensively the known antipatterns can cover these bugs. For bugs not covered by the known antipatterns, we extract new database-access performance antipatterns based on real-world performance bugs from such web applications. Our study in total reports 24 known and 10 new database-access performance antipatterns. Our results can guide future work to develop effective tool support for different types of web applications. Shudi Shao, Zhengyi Qiu, Wei Yang 0013, Guoliang Jin, Tao Xie 0001, Xintao Wu |
ICSME | 4 |
| 2020 | SCDetector: Software Functional Clone Detection Based on Semantic Tokens AnalysisabstractCode clone detection is to find out code fragments with similar functionalities, which has been more and more important in software engineering. Many approaches have been proposed to detect code clones, in which token-based methods are the most scalable but cannot handle semantic clones because of the lack of consideration of program semantics. To address the issue, researchers conduct program analysis to distill the program semantics into a graph representation and detect clones by matching the graphs. However, such approaches suffer from low scalability since graph matching is typically time-consuming. Yueming Wu 0001, Deqing Zou, Shihan Dou, Siru Yang, Wei Yang 0013, Hai Jin 0001 |
ASE | 5 |
| 2020 | DENAS: automated rule generation by knowledge extraction from neural networksabstractDeep neural networks (DNNs) have been widely applied in the software development process to automatically learn patterns from massive data. However, many applications still make decisions based on rules that are manually crafted and verified by domain experts due to safety or security concerns. In this paper, we aim to close the gap between DNNs and rule-based systems by automating the rule generation process via extracting knowledge from well-trained DNNs. Existing techniques with similar purposes either rely on specific DNNs input instances or use inherently unstable random sampling of the input space. Therefore, these approaches either limit the exploration area to a local decision-space of the DNNs or fail to converge to a consistent set of rules. The resulting rules thus lack representativeness and stability. Soroush Bateni, Sampath Grandhi, Xiaodi Li 0002, Cong Liu 0005, Wei Yang 0013 |
ESEC/SIGSOFT FSE | 6 |
| 2020 | TextExerciser: Feedback-driven Text Input Exercising for Android ApplicationsabstractDynamic analysis of Android apps is often used together with an exerciser to increase its code coverage. One big obstacle in designing such Android app exercisers comes from the existence of text-based inputs, which are often constrained by the nature of the input field, such as the length and character restrictions.In this paper, we propose TextExerciser, an iterative, feedback-driven text input exerciser, which generates text inputs for Android apps. Our key insight is that Android apps often provide feedback, called hints, for malformed inputs so that our system can utilize such hints to improve the input generation.We implemented a prototype of TextExerciser and evaluated it by comparing TextExerciser with state-of-the-art exercisers, such as The Monkey and DroidBot. Our evaluation shows that TextExerciser can achieve significantly higher code coverage and trigger more sensitive behaviors than these tools. We also combine TextExerciser with dynamic analysis tools and show they are able to detect more privacy leaks and vulnerabilities with TextExerciser than with existing exercisers. Particularly, existing tools, under the help of TextExerciser, find several new vulnerabilities, such as one user credential leak in a popular social app with more than 10,000,000 downloads. Yuyu He 0001, Lei Zhang 0096, Zhemin Yang, Yinzhi Cao, Keke Lian, Shuai Li 0006, Wei Yang 0013, Zhibo Zhang 0006, Min Yang 0002, Yuan Zhang 0009, Hai-Xin Duan |
SP | 7 |
| 2019 | Charting the Attack Surface of Trigger-Action IoT PlatformsabstractInternet of Things (IoT) deployments are becoming increasingly automated and vastly more complex. Facilitated by programming abstractions such as trigger-action rules, end-users can now easily create new functionalities by interconnecting their devices and other online services. However, when multiple rules are simultaneously enabled, complex system behaviors arise that are difficult to understand or diagnose. While history tells us that such conditions are ripe for exploitation, at present the security states of trigger-action IoT deployments are largely unknown. In this work, we conduct a comprehensive analysis of the interactions between trigger-action rules in order to identify their security risks. Using IFTTT as an exemplar platform, we first enumerate the space of inter-rule vulnerabilities that exist within trigger-action platforms. To aid users in the identification of these dangers, we go on to present iRuler, a system that performs Satisfiability Modulo Theories (SMT) solving and model checking to discover inter-rule vulnerabilities within IoT deployments. iRuler operates over an abstracted information flow model that represents the attack surface of an IoT deployment, but we discover in practice that such models are difficult to obtain given the closed nature of IoT platforms. To address this, we develop methods that assist in inferring trigger-action information flows based on Natural Language Processing. We develop a novel evaluative methodology for approximating plausible real-world IoT deployments based on the installation counts of 315,393 IFTTT applets, determining that 66% of the synthetic deployments in the IFTTT ecosystem exhibit the potential for inter-rule vulnerabilities. Combined, these efforts provide the insight into the real-world dangers of IoT deployment misconfigurations. Qi Wang 0017, Pubali Datta, Wei Yang 0013, Si Liu 0003, Adam Bates 0001, Carl A. Gunter |
CCS | 3 |
| 2019 | MalScan: Fast Market-Wide Mobile Malware Scanning by Social-Network Centrality AnalysisabstractMalware scanning of an app market is expected to be scalable and effective. However, existing approaches use either syntax-based features which can be evaded by transformation attacks or semantic-based features which are usually extracted by performing expensive program analysis. Therefor, in this paper, we propose a lightweight graph-based approach to perform Android malware detection. Instead of traditional heavyweight static analysis, we treat function call graphs of apps as social networks and perform social-network-based centrality analysis to represent the semantic features of the graphs. Our key insight is that centrality provides a succinct and fault-tolerant representation of graph semantics, especially for graphs with certain amount of inaccurate information (e.g., inaccurate call graphs). We implement a prototype system, MalScan, and evaluate it on datasets of 15,285 benign samples and 15,430 malicious samples. Experimental results show that MalScan is capable of detecting Android malware with up to 98% accuracy under one second which is more than 100 times faster than two state-of-the-art approaches, namely MaMaDroid and Drebin. We also demonstrate the feasibility of MalScan on market-wide malware scanning by performing a statistical study on over 3 million apps. Finally, in a corpus of dataset collected from Google-Play app market, MalScan is able to identify 18 zero-day malware including malware samples that can evade detection of existing tools. Yueming Wu 0001, Xiaodi Li 0002, Deqing Zou, Wei Yang 0013, Hai Jin 0001 |
ASE | 4 |
| 2019 | REINAM: reinforcement learning for input-grammar inferenceabstractProgram input grammars (i.e., grammars encoding the language of valid program inputs) facilitate a wide range of applications in software engineering such as symbolic execution and delta debugging. Grammars synthesized by existing approaches can cover only a small part of the valid input space mainly due to unanalyzable code (e.g., native code) in programs and lacking high-quality and high-variety seed inputs. To address these challenges, we present REINAM, a reinforcement-learning approach for synthesizing probabilistic context-free program input grammars without any seed inputs. REINAM uses an industrial symbolic execution engine to generate an initial set of inputs for the given target program, and then uses an iterative process of grammar generalization to proactively generate additional inputs to infer grammars generalized from these initial seed inputs. To efficiently search for target generalizations in a huge search space of candidate generalization operators, REINAM includes a novel formulation of the search problem as a reinforcement learning problem. Our evaluation on eleven real-world benchmarks shows that REINAM outperforms an existing state-of-the-art approach on precision and recall of synthesized grammars, and fuzz testing based on REINAM substantially increases the coverage of the space of valid inputs. REINAM is able to synthesize a grammar covering the entire valid input space for some benchmarks without decreasing the accuracy of the grammar. Zhengkai Wu, Evan Johnson 0001, Wei Yang 0013, Osbert Bastani, Dawn Song, Jian Peng 0001, Tao Xie 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2018 | Property Inference Attacks on Fully Connected Neural Networks using Permutation Invariant RepresentationsabstractWith the growing adoption of machine learning, sharing of learned models is becoming popular. However, in addition to the prediction properties the model producer aims to share, there is also a risk that the model consumer can infer other properties of the training data the model producer did not intend to share. In this paper, we focus on the inference of global properties of the training data, such as the environment in which the data was produced, or the fraction of the data that comes from a certain class, as applied to white-box Fully Connected Neural Networks (FCNNs). Because of their complexity and inscrutability, FCNNs have a particularly high risk of leaking unexpected information about their training sets; at the same time, this complexity makes extracting this information challenging. We develop techniques that reduce this complexity by noting that FCNNs are invariant under permutation of nodes in each layer. We develop our techniques using representations that capture this invariance and simplify the information extraction task. We evaluate our techniques on several synthetic and standard benchmark datasets and show that they are very effective at inferring various data properties. We also perform two case studies to demonstrate the impact of our attack. In the first case study we show that a classifier that recognizes smiling faces also leaks information about the relative attractiveness of the individuals in its training set. In the second case study we show that a classifier that recognizes Bitcoin mining from performance counters also leaks information about whether the classifier was trained on logs from machines that were patched for the Meltdown and Spectre attacks. Karan Ganju, Qi Wang 0017, Wei Yang 0013, Carl A. Gunter, Nikita Borisov |
CCS | 3 |
| 2018 | SemRegex: A Semantics-Based Approach for Generating Regular Expressions from Natural Language SpecificationsabstractRecent research proposes syntax-based approaches to address the problem of generating programs from natural language specifications.These approaches typically train a sequence-to-sequence learning model using a syntax-based objective: maximum likelihood estimation (MLE).Such syntax-based approaches do not effectively address the goal of generating semantically correct programs, because these approaches fail to handle Program Aliasing, i.e., semantically equivalent programs may have many syntactically different forms.To address this issue, in this paper, we propose a semantics-based approach named SemRegex.SemRegex provides solutions for a subtask of the program-synthesis problem: generating regular expressions from natural language.Different from the existing syntax-based approaches, SemRegex trains the model by maximizing the expected semantic correctness of the generated regular expressions.The semantic correctness is measured using the DFA-equivalence oracle, random test cases, and distinguishing test cases.The experiments on three public datasets demonstrate the superiority of SemRegex over the existing state-of-the-art approaches. Zexuan Zhong, Wei Yang 0013, Jian Peng 0001, Tao Xie 0001, Jian-Guang Lou, Ting Liu 0002, Dongmei Zhang 0001 |
EMNLP | 3 |
| 2018 | EnMobile: entity-based characterization and analysis of mobile malwareabstractModern mobile malware tend to conduct their malicious exploits through sophisticated patterns of interactions that involve multiple entities, e.g., the mobile platform, human users, and network locations. Such malware often evade the detection by existing approaches due to their limited expressiveness and accuracy in characterizing and detecting these malware. To address these issues, in this paper, we recognize entities in the environment of an app, the app's interactions with such entities, and the provenance of these interactions, i.e., the intent and ownership of each interaction, as the key to comprehensively characterizing modern mobile apps, and mobile malware in particular. With this insight, we propose a novel approach named EnMobile including a new entity-based characterization of mobile-app behaviors, and corresponding static analyses, to accurately characterize an app's interactions with entities. We implement EnMobile and provide a practical application of EnMobile in a signature-based scheme for detecting mobile malware. We evaluate EnMobile on a set of 6614 apps consisting of malware from Genome and Drebin along with benign apps from Google Play. Our results show that EnMobile detects malware with substantially higher precision and recall than four state-of-the-art approaches, namely Apposcopy, Drebin, MUDFLOW, and AppContext. Wei Yang 0013, Mukul R. Prasad, Tao Xie 0001 |
ICSE | 1 |
| 2018 | An empirical study of Android test generation tools in industrial casesabstractUser Interface (UI) testing is a popular approach to ensure the quality of mobile apps. Numerous test generation tools have been developed to support UI testing on mobile apps, especially for Android apps. Previous work evaluates and compares different test generation tools using only relatively simple open-source apps, while real-world industrial apps tend to have more complex functionalities and implementations. There is no direct comparison among test generation tools with regard to effectiveness and ease-of-use on these industrial apps. To address such limitation, we study existing state-of-the-art or state-of-the-practice test generation tools on 68 widely-used industrial apps. We directly compare the tools with regard to code coverage and fault-detection ability. According to our results, Monkey, a state-of-the-practice tool from Google, achieves the highest method coverage on 22 of 41 apps whose method coverage data can be obtained. Of all 68 apps under study, Monkey also achieves the highest activity coverage on 35 apps, while Stoat, a state-of-the-art tool, is able to trigger the highest number of unique crashes on 23 apps. By analyzing the experimental results, we provide suggestions for combining different test generation tools to achieve better performance. We also report our experience in applying these tools to industrial apps under study. Our study results give insights on how Android UI test generation tools could be improved to better handle complex industrial apps. Dengfeng Li 0003, Wei Yang 0013, Yurui Cao, Zhenwen Zhang, Yuetang Deng, Tao Xie 0001 |
ASE | 3 |
| 2018 | Mining Android App Descriptions for Permission Requirements RecommendationabstractDuring the development or maintenance of an Android app, the app developer needs to determine the app's security and privacy requirements such as permission requirements. Permission requirements include two folds. First, what permissions (i.e., access to sensitive resources, e.g., location or contact list) the app needs to request. Second, how to explain the reason of permission usages to users. In this paper, we focus on the multiple challenges that developers face when creating permission-usage explanations. We propose a novel framework, CLAP, that mines potential explanations from the descriptions of similar apps. CLAP leverages information retrieval and text summarization techniques to find frequent permission usages. We evaluate CLAP on a large dataset containing 1.4 million Android apps. The evaluation results outperform existing state-of-the-art approaches, showing great promise of CLAP as a tool for assisting developers and permission requirements discovery. Xueqing Liu 0001, Yue Leng, Wei Yang 0013, ChengXiang Zhai, Tao Xie 0001 |
RE | 3 |
| 2018 | A Large-Scale Empirical Study on Android Runtime-Permission Rationale MessagesabstractAfter Android 6.0 introduces the runtime-permission system, many apps provide runtime-permission-group rationales for the users to better understand the permissions requested by the apps. To understand the patterns of rationales and to what extent the rationales can improve the users' understanding of the purposes of requesting permission groups, we conduct a large-scale measurement study on five aspects of runtime rationales. We have five main findings: (1) less than 25% apps under study provide rationales; (2) for permission-group purposes that are difficult to understand, the proportions of apps that provide rationales are even lower; (3) the purposes stated in a significant proportion of rationales are incorrect; (4) a large proportion of customized rationales do not provide more information than the default permission-requesting message of Android; (5) apps that provide rationales are more likely to explain the same permission group's purposes in their descriptions than apps that do not provide rationales. We further discuss important implications from these findings. Xueqing Liu 0001, Yue Leng, Wei Yang 0013, ChengXiang Zhai, Tao Xie 0001 |
VL/HCC | 3 |
| 2017 | Malware Detection in Adversarial Settings: Exploiting Feature Evolutions and Confusions in Android AppsabstractExisting techniques on adversarial malware generation employ feature mutations based on feature vectors extracted from malware. However, most (if not all) of these techniques suffer from a common limitation: feasibility of these attacks is unknown. The synthesized mutations may break the inherent constraints posed by code structures of the malware, causing either crashes or malfunctioning of malicious payloads. To address the limitation, we present Malware Recomposition Variation (MRV), an approach that conducts semantic analysis of existing malware to systematically construct new malware variants for malware detectors to test and strengthen their detection signatures/models. In particular, we use two variation strategies (i.e., malware evolution attack and malware confusion attack) following structures of existing malware to enhance feasibility of the attacks. Upon the given malware, we conduct semantic-feature mutation analysis and phylogenetic analysis to synthesize mutation strategies. Based on these strategies, we perform program transplantation to automatically mutate malware bytecode to generate new malware variants. We evaluate our MRV approach on actual malware variants, and our empirical evaluation on 1,935 Android benign apps and 1,917 malware shows that MRV produces malware variants that can have high likelihood to evade detection while still retaining their malicious behaviors. We also propose and evaluate three defense mechanisms to counter MRV. Wei Yang 0013, Deguang Kong, Tao Xie 0001, Carl A. Gunter |
ACSAC | 1 |
| 2016 | Free for All! Assessing User Data Exposure to Advertising Libraries on Android
Soteris Demetriou, Whitney Merrill, Wei Yang 0013, Aston Zhang, Carl A. Gunter |
NDSS | 3 |
| 2016 | Automated test input generation for Android: are we really there yet in an industrial case?abstractGiven the ever increasing number of research tools to automatically generate inputs to test Android applications (or simply apps), researchers recently asked the question "Are we there yet?" (in terms of the practicality of the tools). By conducting an empirical study of the various tools, the researchers found that Monkey (the most widely used tool of this category in industrial practices) outperformed all of the research tools that they studied. In this paper, we present two significant extensions of that study. First, we conduct the first industrial case study of applying Monkey against WeChat, a popular messenger app with over 762 million monthly active users, and report the empirical findings on Monkey's limitations in an industrial setting. Second, we develop a new approach to address major limitations of Monkey and accomplish substantial code-coverage improvements over Monkey, along with empirical insights for future enhancements to both Monkey and our approach. Xia Zeng, Dengfeng Li 0003, Wujie Zheng, Yuetang Deng, Wing Lam, Wei Yang 0013, Tao Xie 0001 |
SIGSOFT FSE | 7 |
| 2015 | AppContext: Differentiating Malicious and Benign Mobile App Behaviors Using ContextabstractMobile malware attempts to evade detection during app analysis by mimicking security-sensitive behaviors of benign apps that provide similar functionality (e.g., sending SMS messages), and suppressing their payload to reduce the chance of being observed (e.g., executing only its payload at night). Since current approaches focus their analyses on the types of security-sensitive resources being accessed (e.g., network), these evasive techniques in malware make differentiating between malicious and benign app behaviors a difficult task during app analysis. We propose that the malicious and benign behaviors within apps can be differentiated based on the contexts that trigger security-sensitive behaviors, i.e., the events and conditions that cause the security-sensitive behaviors to occur. In this work, we introduce AppContext, an approach of static program analysis that extracts the contexts of security-sensitive behaviors to assist app analysis in differentiating between malicious and benign behaviors. We implement a prototype of AppContext and evaluate AppContext on 202 malicious apps from various malware datasets, and 633 benign apps from the Google Play Store. AppContext correctly identifies 192 malicious apps with 87.7% precision and 95% recall. Our evaluation results suggest that the maliciousness of a security-sensitive behavior is more closely related to the intention of the behavior (reflected via contexts) than the type of the security-sensitive resources that the behavior accesses. Wei Yang 0013, Xusheng Xiao, Benjamin Andow, Tao Xie 0001, William Enck |
ICSE (1) | 1 |
| 2013 | A Grey-Box Approach for Automated GUI-Model Generation of Mobile Applications
Wei Yang 0013, Mukul R. Prasad, Tao Xie 0001 |
FASE | 1 |
| 2013 | WHYPER: Towards Automating Risk Assessment of Mobile Applications
Rahul Pandita, Xusheng Xiao, Wei Yang 0013, William Enck, Tao Xie 0001 |
USENIX Security Symposium | 3 |