Yu Zhao 0010

dblp:57/2056-10 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 4 first-author · 9 since 2021Computer networks · 3 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-author
YearPublicationVenuePosition
2025 Multimodal Fusion for Android Malware Detection Based on Large Pre-Trained Models
abstract
Malware detection is a critical issue in software engineering as it directly threatens user information security. Existing approaches often focus on individual modality (either source code or binary code) for the detection, but it ignores to effectively exploit the complementary information between them. This limits the detection performance, especially in complex and evasive malware scenarios. In this paper, we take Android applications written in Java as objects, and provide a novel fine-grained multimodal fusion method with large pre-trained models to combine the features from source and binary codes for the malware detection. For the source code modality, we employ the graphical user interface (GUI) as a framework to segment the source code into snippets, and use a pre-trained programming language model to extract feature representations. For the binary code modality, we convert binary code into grayscale images and fine-tune a pre-trained vision model to extract features indirectly. We then implement cross-modal attention and devise a contrastive loss to align features across modalities, supplementing this with supervised classification loss to refine the multimodal fusion process specifically for malware detection. Our experiments, conducted using the Data-MD and Data-MC benchmarks, demonstrate that our approach achieves a precision of 0.977 and a recall of 0.984 in detecting malware. This underscores the advantages of using large pre-trained models for feature representation and the fusion of information across different modalities for effective malware detection.
Lei Liu 0040, Yuzhou Liu 0001, Yu Zhao 0010, Peng Zhang 0053, Huaxiao Liu
IEEE Trans. Software Eng.4
2024 A Study of Using Multimodal LLMs for Non-Crash Functional Bug Detection in Android Apps
abstract
Numerous approaches employing various strategies have been developed to test the graphical user interfaces (GUIs) of mobile apps. However, traditional GUI testing techniques, such as random and model-based testing, primarily focus on generating test sequences that excel in achieving high code coverage but often fail to act as effective test oracles for noncrash functional (NCF) bug detection. To tackle these limitations, this study empirically investigates the capability of leveraging large language models (LLMs) to be test oracles to detect NCF bugs in Android apps. Our intuition is that the training corpora of LLMs, encompassing extensive mobile app usage and bug report descriptions, enable them with the domain knowledge relevant to NCF bug detection. We conducted a comprehensive empirical study to explore the effectiveness of LLMs as test oracles for detecting NCF bugs in Android apps on 71 welldocumented NCF bugs. The results demonstrated that LLMs achieve a 49% bug detection rate, outperforming existing tools for detecting NCF bugs in Android apps. Additionally, by leveraging LLMs to be test oracles, we successfully detected 24 previously unknown NCF bugs in 64 Android apps, with four of these bugs being confirmed or fixed. However, we also identified limitations of LLMs, primarily related to performance degradation, inherent randomness, and false positives. Our study highlights the potential of leveraging LLMs as test oracles for Android NCF bug detection and suggests directions for future research.
Bangyan Ju, Tingting Yu 0001, Tamerlan Abdullayev, Dingbang Wang, Yu Zhao 0010
APSEC7
2024 Feedback-Driven Automated Whole Bug Report Reproduction for Android Apps
abstract
In software development, bug report reproduction is a challenging task. This paper introduces ReBL, a novel feedback-driven approach that leverages GPT-4, a large-scale language model (LLM), to automatically reproduce Android bug reports. Unlike traditional methods, ReBL bypasses the use of Step to Reproduce (S2R) entities. Instead, it leverages the entire textual bug report and employs innovative prompts to enhance GPT’s contextual reasoning. This approach is more flexible and context-aware than the traditional step-by-step entity matching approach, resulting in improved accuracy and effectiveness. In addition to handling crash reports, ReBL has the capability of handling non-crash functional bug reports. Our evaluation of 96 Android bug reports (73 crash and 23 non-crash) demonstrates that ReBL successfully reproduced 90.63% of these reports, averaging only 74.98 seconds per bug report. Additionally, ReBL outperformed three existing tools in both success rate and speed.
Dingbang Wang, Yu Zhao 0010, Sidong Feng, William G. J. Halfond, Chunyang Chen 0001, Xiaoxia Sun, Jiangfan Shi, Tingting Yu 0001
ISSTA2
2024 DinoDroid: Testing Android Apps Using Deep Q-Networks
abstract
The large demand of mobile devices creates significant concerns about the quality of mobile applications (apps). Developers need to guarantee the quality of mobile apps before it is released to the market. There have been many approaches using different strategies to test the GUI of mobile apps. However, they still need improvement due to their limited effectiveness. In this article, we propose DinoDroid, an approach based on deep Q-networks to automate testing of Android apps. DinoDroid learns a behavior model from a set of existing apps and the learned model can be used to explore and generate tests for new apps. DinoDroid is able to capture the fine-grained details of GUI events (e.g., the content of GUI widgets) and use them as features that are fed into deep neural network, which acts as the agent to guide app exploration. DinoDroid automatically adapts the learned model during the exploration without the need of any modeling strategies or pre-defined rules. We conduct experiments on 64 open-source Android apps. The results showed that DinoDroid outperforms existing Android testing tools in terms of code coverage and bug detection.
Yu Zhao 0010, Brent E. Harrison, Tingting Yu 0001
ACM Trans. Softw. Eng. Methodol.1
2023 An Empirical Study of Regression Testing for Android Apps in Continuous Integration Environment
abstract
Continuous integration (CI) has become a popular method for automating code changes, testing, and software project delivery. However, sufficient testing prior to code submission is crucial to prevent build breaks. Additionally, testing must provide developers with quick feedback on code changes, which requires fast testing times. While regression test selection (RTS) has been studied to improve the cost-effectiveness of regression testing for lower-level tests (i.e., unit tests), it has not been applied to the testing of user interfaces (UI) in application domains such as mobile apps. UI testing at the UI level requires different techniques such as impact analysis and automated test execution. In this paper, we examine the use of RTS in CI settings for UI testing across various open-source mobile apps. Our analysis focuses on using Frequency Analysis to understand the need for RTS, Cost Analysis to evaluate the cost of impact analysis and test case selection algorithms, and Test Reuse Analysis to determine the reusability of UI test sequences for automation. The insights from this study will guide practitioners and researchers in developing advanced RTS techniques that can be adapted to CI environments for mobile apps.
Dingbang Wang, Yu Zhao 0010, Lu Xiao 0001, Tingting Yu 0001
ESEM2
2023 Automatically Reproducing Android Bug Reports using Natural Language Processing and Reinforcement Learning
abstract
As part of the process of resolving issues submitted by users via bug reports, Android developers attempt to reproduce and observe the crashes described by the bug reports. Due to the low-quality of bug reports and the complexity of modern apps, the reproduction process is non-trivial and time-consuming. Therefore, automatic approaches that can help reproduce Android bug reports are in great need. However, current approaches to help developers automatically reproduce bug reports are only able to handle limited forms of natural language text and struggle to successfully reproduce crashes for which the initial bug report had missing or imprecise steps. In this paper, we introduce a new fully automated approach to reproduce crashes from Android bug reports that addresses these limitations. Our approach accomplishes this by leveraging natural language processing techniques to more holistically and accurately analyze the natural language in Android bug reports and designing new techniques, based on reinforcement learning, to guide the search for successful reproducing steps. We conducted an empirical evaluation of our approach on 77 real world bug reports. Our approach achieved 67% precision and 77% recall in accurately extracting reproduction steps from bug reports, reproduced 74% of the total bug reports, and reproduced 64% of the bug reports that contained missing steps, significantly outperforming state of the art techniques.
Robert Winn, Yu Zhao 0010, Tingting Yu 0001, William G. J. Halfond
ISSTA3
2022 Effective fault localization and context-aware debugging for concurrent programs
abstract
Summary Concurrent programs are difficult to debug because concurrency faults usually occur under specific inputs and thread interleavings. Fault localization techniques for sequential programs are often ineffective because the root causes of concurrency faults involve memory accesses across multiple threads rather than single statements. Previous research has proposed techniques to analyse passing and failing executions obtained from running a set of test cases for identifying faulty memory access patterns. However, stand‐alone access patterns do not provide enough contextual information, such as the path leading to the failure, for developers to understand the bug. We present an approach, Coadec, to automatically generate interthread control flow paths that can link memory access patterns that occurred most frequently in the failing executions to better diagnose concurrency bugs. Coadec consists of two phases. In the first phase, we use feature selection techniques from machine learning to localize suspicious memory access patterns based on failing and passing executions. The patterns with maximum feature diversity information can point to the most suspicious pattern. We then apply a data mining technique and identify the memory access patterns that occurred most frequently in the failing executions. Finally, Coadec identifies faulty program paths by connecting both the frequent patterns and the suspicious pattern. We also evaluate the effectiveness of fault localization using test suites generated from different test adequacy criteria. We introduce and have evaluated Coadec on 10 real‐world multithreaded Java applications. Results indicate that Coadec outperforms state‐of‐the‐art approaches for localizing concurrency faults and that Coadec's context debugging can help developers understand concurrency fault by inspecting a small percentage of code.
Justin Chu, Tingting Yu 0001, Jane Huffman Hayes, Xue Han 0007, Yu Zhao 0010
Softw. Test. Verification Reliab.5
2022 ReCDroid+: Automated End-to-End Crash Reproduction from Bug Reports for Android Apps
abstract
The large demand of mobile devices creates significant concerns about the quality of mobile applications (apps). Developers heavily rely on bug reports in issue tracking systems to reproduce failures (e.g., crashes). However, the process of crash reproduction is often manually done by developers, making the resolution of bugs inefficient, especially given that bug reports are often written in natural language. To improve the productivity of developers in resolving bug reports, in this paper, we introduce a novel approach, called ReCDroid+, that can automatically reproduce crashes from bug reports for Android apps. ReCDroid+ uses a combination of natural language processing (NLP) , deep learning, and dynamic GUI exploration to synthesize event sequences with the goal of reproducing the reported crash. We have evaluated ReCDroid+ on 66 original bug reports from 37 Android apps. The results show that ReCDroid+ successfully reproduced 42 crashes (63.6% success rate) directly from the textual description of the manually reproduced bug reports. A user study involving 12 participants demonstrates that ReCDroid+ can improve the productivity of developers when resolving crash bug reports.
Yu Zhao 0010, Ting Su 0001, Yang Liu 0003, Wei Zheng 0006, Xiaoxue Wu 0001, Ramakanth Kavuluru, William G. J. Halfond, Tingting Yu 0001
ACM Trans. Softw. Eng. Methodol.1
2021 Improving high-impact bug report prediction with combination of interactive machine learning and active learning
Xiaoxue Wu 0001, Wei Zheng 0006, Xiang Chen 0005, Yu Zhao 0010, Tingting Yu 0001
Inf. Softw. Technol.4
2020 SICS: Secure and Dynamic Middlebox Outsourcing
abstract
There is an increasing trend that enterprises outsource their middlebox processing to a cloud for lower cost and easier management. However, outsourcing middleboxes brings threats to the enterprise's private information, including the traffic and rules of middleboxes, all of which are visible within the cloud. Existing solutions for secure middlebox outsourcing either incur significant performance overhead or do not support incremental updates. In this article, we present a secure and dynamic middlebox outsourcing framework, SICS, short for Secure In-Cloud Service. SICS encrypts each packet header and uses a label for in-cloud rule matching, which enables the cloud to perform its functionalities correctly with minimum header information leakage. Evaluation results show that SICS achieves higher throughput, faster construction and update speed, and lower resource overhead at the enterprise and in the cloud when compared with existing solutions.
Huazhe Wang, Xin Li 0057, Yang Wang 0009, Yu Zhao 0010, Ye Yu 0001, Hongkun Yang, Chen Qian 0001
IEEE/ACM Trans. Netw.4
2019 ReCDroid: automatically reproducing Android application crashes from bug reports
abstract
The large demand of mobile devices creates significant concerns about the quality of mobile applications (apps). Developers heavily rely on bug reports in issue tracking systems to reproduce failures (e.g., crashes). However, the process of crash reproduction is often manually done by developers, making the resolution of bugs inefficient, especially that bug reports are often written in natural language. To improve the productivity of developers in resolving bug reports, in this paper, we introduce a novel approach, called ReCDroid, that can automatically reproduce crashes from bug reports for Android apps. ReCDroid uses a combination of natural language processing (NLP) and dynamic GUI exploration to synthesize event sequences with the goal of reproducing the reported crash. We have evaluated ReCDroid on 51 original bug reports from 33 Android apps. The results show that ReCDroid successfully reproduced 33 crashes (63.5% success rate) directly from the textual description of bug reports. A user study involving 12 participants demonstrates that ReCDroid can improve the productivity of developers when resolving crash bug reports.
Yu Zhao 0010, Tingting Yu 0001, Ting Su 0001, Yang Liu 0003, Wei Zheng 0006, Jingzhi Zhang, William G. J. Halfond
ICSE1
2019 Automatically Extracting Bug Reproducing Steps from Android Bug Reports
Yu Zhao 0010, Kye Miller, Tingting Yu 0001, Wei Zheng 0006, Minchao Pu
ICSR1
2018 FREDI: Robust RSS-based ranging with multipath effect and radio interference
Yu Zhao 0010, Yunhuai Liu, Tingting Yu 0001, Tian He 0001, Chen Qian 0001
Comput. Networks1
2017 Channel Quality Correlation Based Channel Probing in Multiple Channels
abstract
Many wireless networks provide a large number of available channels for data transmissions. Due to the multi-path environment, channels have different channel qualities. There- fore, selecting a good channel from the multiple channels can improve the communication effectiveness. The Packet Reception Rate (PRR) has been utilized to represent the channel quality. In order to know which channel has good quality, communication systems probe each channel and measure PRR. However, probing is time and energy consuming. In this paper, we propose a wireless channel probing approach, optimal channel probing (OCP), to efficiently select high-quality wireless channels. The OCP method utilizes a MAX-separation method to reduce the number of probing channels. In an evaluation using simulations and usrp2 real devices, we found that OCP is both effective and efficient at selecting high- quality channels and more efficient than existing channel probing methods.
Yu Zhao 0010, Tingting Yu 0001
ICCCN1
2017 Pronto: Efficient Test Packet Generation for Dynamic Network Data Planes
abstract
Computer networks are becoming increasingly complex today and thus prone to various network faults. Traditional testing tools (e.g., ping, traceroute) that often involve substantial manual effort to uncover faults are inefficient. This paper focuses on fault detection of the network data plane using test packets. Existing solutions of test packet generation either take very long time (e.g., more than one hour) to complete or generate too many test packets that may hurt regular traffic. In this paper, we present Pronto, an automated test packet generation tool that generates test packets to exercise data plane rules in the entire network in a short time (e.g., several seconds) and can quickly react to rule changes due to network dynamics. In addition, Pronto minimizes the number of test packets by allowing a packet to test multiple rules at different switches. The performance evaluation using two real network data plane rule sets shows that Pronto is faster than a recently developed tool by more than two orders of magnitude. Pronto can update the probes for rule changes using less than 1ms while existing methods have no such update function.
Yu Zhao 0010, Huazhe Wang, Tingting Yu 0001, Chen Qian 0001
ICDCS1
2014 Cloud-assisted analysis for energy efficiency in intelligent video systems
Yu Zhao 0010, Yunhuai Liu, Chuanping Hu
J. Supercomput.2