Yao Zhang 0019

dblp:57/3892-19 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-7375-9152ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 12 · 3 first-author · 11 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented Generation
Zhiping Zhou, Xiaohong Li 0001, Yao Zhang 0019, Yuekang Li, Wenbu Feng, Yunqian Wang
NDSS4
2026 kAPR: A coverage-guided, context-aware agent for automated repair of Linux kernel bugs
Bingzheng Li, Xiaokang Yin 0002, Yao Zhang 0019, Shengli Liu 0003, Shouling Ji
Inf. Softw. Technol.3
2026 Determining the Unreachable: Constraint-Guided Reachability Analysis for Dependency Vulnerabilities
abstract
In software development, investigating the accessibility of dependency vulnerabilities is of great importance, as third-party libraries often contain known vulnerabilities that could be exploited in the application's business logic. The existing accessibility analysis methods encounter challenges such as undecidability, abstraction loss, and path explosion in large-scale programs, resulting in an inaccurate distinction between accessibility vulnerabilities and non-accessibility vulnerabilities. This paper introduces an approach called ConVReach for analyzing the reachability of vulnerabilities in dependencies in C/C++ programs. ConVReach overcomes the problems of high abstraction loss and potential path explosion in the current methods by combining static and dynamic approaches, particularly a constraint-guided analysis method. This approach extracts and decomposes the path constraints that trigger vulnerabilities, independently verifies the satisfiability of each constraint, and then aggregates the feasible paths. This effectively reduces unnecessary path exploration and avoids the common path explosion issues in traditional methods. Experimental results show that ConVReach outperforms existing tools in both accuracy and efficiency, effectively distinguishing between reachable and unreachable vulnerabilities, and significantly reducing false positives and false negatives. We constructed a benchmark dataset to evaluate ConVReach , which includes 53 CVEs and 347 flags artificially inserted into various open-source projects. This dataset was designed to simulate both real-world vulnerabilities and complex scenarios. Through testing on this dataset, ConVReach demonstrated exceptional performance. It successfully identified 59 out of 61 reachable vulnerabilities and all 23 unreachable ones in the CVE dataset. Within a 24-hour time budget, ConVReach detected above 50% more reachable vulnerabilities than the baseline tools in the first 6 hours and nearly completed the detection of reachable vulnerabilities by the 12-hour mark. These results highlight ConVReach 's superior ability to handle both real-world vulnerabilities and challenging cases.
Wenbu Feng, Xiaohong Li 0001, Yao Zhang 0019, Yuekang Li, Zhiping Zhou, Yunqian Wang
Proc. ACM Program. Lang.4
2026 Knowledge is Power: A Knowledge Graph-Based Approach for Mobile Malware Traceability Analysis
Yao Zhang 0019, Guangquan Xu, Xiaohong Li 0001, Sen Chen 0001, Zhenchang Xing, Yude Bai, Yongqiang Lyu 0001, Wei Gong 0001, Xibin Zhao
IEEE Trans. Mob. Comput.1
2026 LIMR: Intent-Aware Mashup API Recommendation via LLM-Augmented Multi-Scale Fusion
abstract
The increasing availability of Web APIs has amplified the complexity of mashup creation, where developers must identify compatible and functionally relevant APIs based on often ambiguous natural language descriptions. Traditional methods also fall short in capturing hierarchical semantic cues, modeling compatibility, and aligning with developer intent. Although large language models (LLMs) offer strong generalization capabilities, they remain unreliable in mashup recommendation due to hallucinated outputs, limited controllability, and token-length constraints when dealing with large-scale API repositories. To overcome these limitations, we introduceLIMR, an intent-aware mashup recommendation framework that combines LLM-augmented semantic reasoning with structured, multi-scale neural modeling.LIMRfirst prompts a LLM to extract high-level intent from user requirements, which serves as a global semantic signal. This intent is fused with low-level, multi-scale features extracted by a convolutional encoder, which are designed to capture fine-grained lexical/phrasal patterns at different granularities and provide precise semantic grounding for API matching. These heterogeneous representations are further contextually refined through a Transformer-based interaction module. To handle nonlinear semantic dependencies and compositional complexity,LIMRintegrates a Kolmogorov-Arnold Network (KAN) with learnable activation functions, enhancing the model's capacity to capture intricate feature interactions. The entire framework is optimized via LLM, incorporating auxiliary objectives such as mashup category prediction and API quality estimation to guide generalization and reduce overfitting. Comprehensive experiments on the ProgrammableWeb and APIBench datasets show thatLIMRsignificantly outperforms state-of-the-art baselines, which the ranking-oriented metrics, including NDCG and mAP, achieves improvements of 17.1%–34.2% over the strongest competitors. These results confirm the effectiveness ofLIMR's hybrid design in delivering precise, robust, and intent-aware mashup API recommendations, especially in scenarios where LLMs alone fail to meet accuracy and scalability demands.
Yao Zhang 0019, Yude Bai, Minhong Dong, Keqing Cen, Ji Zhang 0001, Wei Ma 0014, Yongqiang Lyu 0001, Xiaohong Li 0001, Junjie Wang 0007, Lingxiao Jiang, Yang Liu 0003
IEEE Trans. Serv. Comput.1
2025 It Only Gets Worse: Revisiting DL-Based Vulnerability Detectors from a Practical Perspective
abstract
With the escalating threat of software vulnerabilities to the security of modern software systems, an increasing number of deep learning (DL) model-based vulnerability detectors have been developed for vulnerability detection. However, their practical reliability, consistency in usage, and adaptability across diverse software contexts remain unclear. This uncertainty may lead to unreliable detection results in practical applications, increased false positives and false negatives, and limited adaptability to newly emerged vulnerabilities. Conducting a large-scale and in-depth analysis of DL-based vulnerability detectors can help uncover critical factors influencing detection performance, improve the design and training of these models, and enhance their practical deployment in real-world scenarios. In this paper, we present VulTegra, a novel evaluation framework that, for the first time, conducts a multidimensional assessment comparing scratch-trained models and pre-trained-based models for vulnerability detection, while verifying key factors influencing detection performance. Our framework reveals that state-of-the-art (SOTA) detectors still suffer from low consistency, limited practical detection capabilities, and limited adaptability. Moreover, comparative results indicate that the increasingly favored pre-trained-based models are not universally superior to scratch-trained models; instead, they exhibit distinct strengths and application scenarios. Most importantly, our study highlights the limitations of relying solely on CWE-based classification and reveals a set of critical factors that significantly influence detection performance. Experimental validation shows that these factors have a substantial impact: modifying only any single factor led to recall improvements across all seven evaluated SOTA detectors, with six detectors also achieving higher F1 scores. Our findings provide deep insights into model behavior, highlighting the need to consider both vulnerability types and inherent code features to ensure practical applicability in real-world software environments.
Yunqian Wang, Xiaohong Li 0001, Yao Zhang 0019, Yuekang Li, Zhiping Zhou
APSEC4
2025 SMTPRT: Performance Regression Testing and Localization for SMT Solvers Across Multiple Logics
abstract
Satisfiability Modulo Theories (SMT) solvers are foundational in applications such as software verification and automated bug detection, where both correctness and performance are critical to the reliability and scalability of these systems. While existing methods predominantly focus on functional testing, performance testing has received insufficient attention, particularly regarding performance regression caused by both intentional and unintentional factors during software evolution. Current performance regression testing approaches are primarily designed for string solvers, neglecting the full spectrum of SMT theories. Furthermore, these methods often rely on time comparisons or log analysis, which makes the identification of the responsible commit slow and inefficient. To address the above issues, we propose a novel general purpose testing framework, SMTPRT, that efficiently detects and localizes performance regression issues across diverse SMT solver theories. We utilize large language models (LLMs) based on genetic algorithms (GAs) to guide the search for performance regression-inducing cases. We introduce an optimized localization technique that filters irrelevant commits using code coverage, followed by a bisecting algorithm to rapidly pinpoint the responsible commit. To thoroughly evaluate SMTPRT, we conducted extensive experiments involving six types of logic, demonstrating its superior performance. Specifically, SMTPRT successfully detected 59 regression cases, performing 3.44 times better than the baseline, and located the issues $\mathbf{1. 1 6}$ times faster than the baseline.
Xiaohong Li 0001, Lili Quan 0001, Zhiping Zhou, Yao Zhang 0019
APSEC5
2025 EfficientNet-BSFT-S: Dynamic Multi-Scale Modules for Robust Image Classification
Chengjie Guo, Minghong Dong, Xuewei Liu, Meng Xing, Yao Zhang 0019, Yude Bai
ICIC (19)5
2025 PneumoNeXt: A Multi-Scale Attention and Contrastive Learning Approach for Pneumonia Diagnosis
Lirong Zhang, Meng Xing, Yao Zhang 0019, Yude Bai
ICIC (19)3
2025 A Cross-Domain Data Sharing Scheme Based on Federated Blockchain
Honglin Mao, Jie Zhang 0111, Yao Zhang 0019, Xiaohong Li 0001
TASE3
2025 IRHunter: Universal Detection of Instruction Reordering Vulnerabilities for Enhanced Concurrency in Distributed and Parallel Systems
abstract
Instruction reordering is an essential optimization technique used in both compilers and multi-core processors to enhance parallelism and resource utilization. Although the original intent of this technique is to benefit the program, some improper reordering can significantly impact the program correctness, which we call instruction reordering vulnerability (IRV). However, existing methods detect IRV by defining CPU instruction reordering rules to schedule execution paths while neglecting compiler reordering, and thus generate false positives that require manual filtering and resulting in inefficiency. To bridge this gap, in this paper, we propose the IRV detection method, , which analyzes IRV characteristics and extracts vulnerability patterns, integrating program dependency analysis for compiler reordering and memory model constraints for CPU reordering. Specifically, we use static analysis based on specific patterns to narrow the analysis scope, and adopt log-based dynamic analysis to confirm vulnerability by checking the log constraints. We built the IRV benchmark to compare IRHunter with five state-of-the-art tools (i.e., GENMC, Nidhugg, CBMC, SHB, BiRD). IRHunter detected all 19 errors, doubling the best model checking tools' performance, with half the false positive rate of leading data race detectors. It was 10× faster on small programs and outperformed data race detectors on large programs.
Guohua Xin, Guangquan Xu, Yao Zhang 0019, Cheng Wen 0002, Cen Zhang, Xiaofei Xie, Naixue Xiong, Shaoying Liu, Pan Gao 0006
IEEE Trans. Parallel Distributed Syst.3
2023 EndWatch: A Practical Method for Detecting Non-Termination in Real-World Software
abstract
Detecting non-termination is crucial for ensuring program correctness and security, such as preventing denial-of-service attacks. While termination analysis has been studied for many years, existing methods have limited scalability and are only effective on small programs. To address this issue, we propose a practical termination checking technique, called EndWatch, for detecting non-termination caused by infinite loops through testing. Specifically, we introduce two methods to generate non-termination oracles based on checking state revisits, i.e., if the program returns to a previously visited state at the same program location, it does not terminate. The non-termination oracles can be incorporated into testing tools (e.g., AFL used in this paper) to detect non-termination in large programs. For linear loops, we perform symbolic execution on individual loops to infer State Revisit Conditions (SRCs) and instrument SRCs into target loops. For non-linear loops, we instrument target loops for checking concrete state revisits during execution. We evaluated EndWatch on standard benchmarks with small-sized programs and real-world projects with large-sized programs. The evaluation results show that EndWatch is more effective than the state-of-the-art tools on standard benchmarks (detecting 87% of non-terminating programs while the best baseline detects only 67%), and useful in detecting non-termination in real-world projects (detecting 90% of known non-termination CVEs and 4 unknown bugs).
Yao Zhang 0019, Xiaofei Xie, Yi Li 0008, Sen Chen 0001, Cen Zhang, Xiaohong Li 0001
ASE1
2023 Demystifying Performance Regressions in String Solvers
abstract
Over the past few years, SMT string solvers have found their applications in an increasing number of domains, such as program analyses in mobile and Web applications, which require the ability to reason about string values. A series of research has been carried out to find quality issues of string solvers in terms of its correctness and performance. Yet, none of them has considered the performance regressions happening across multiple versions of a string solver. To fill this gap, in this paper, we focus on solver performance regressions (SPRs), i.e., unintended slowdowns introduced during the evolution of string solvers. To this end, we developSPRFinderto not only generate test cases demonstrating SPRs, but also localize the probable causes of them, in terms of commits. We evaluated the effectiveness ofSPRFinderon three state-of-the-art string solvers, i.e., Z3Seq, Z3Str3, and CVC4. The results demonstrate thatSPRFinderis effective in generating SPR-inducing test cases and also able to accurately locate the responsible commits. Specifically, the average running time on the target versions is 13.2× slower than that of the reference versions. Besides, we also conducted the first empirical study to peek into the characteristics of SPRs, including the impact of random seed configuration for SPR detection, understanding the root causes of SPRs, and characterizing the regression test cases through case studies. Finally, we highlight that 149 unique SPR-inducing commits were discovered in total bySPRFinder, and 27of them have been confirmed by the corresponding developers.
Yao Zhang 0019, Xiaofei Xie, Yi Li 0008, Yun Lin 0001, Sen Chen 0001, Yang Liu 0003, Xiaohong Li 0001
IEEE Trans. Software Eng.1
2022 Large-scale analysis of non-termination bugs in real-world OSS projects
abstract
Termination is a crucial program property. Non-termination bugs can be subtle to detect and may remain hidden for long before they take effect. Many real-world programs still suffer from vast consequences (e.g., no response) caused by non-termination bugs. As a classic problem, termination proving has been studied for many years. Many termination checking tools and techniques have been developed and demonstrated effectiveness on existing well-established benchmarks. However, the capability of these tools in finding practical non-termination bugs has yet to be tested on real-world projects. To fill in this gap, in this paper, we conducted the first large-scale empirical study of non-termination bugs in real-world OSS projects. Specifically, we first devoted substantial manual efforts in collecting and analyzing 445 non-termination bugs from 3,142 GitHub commits and provided a systematic classifi-cation of the bugs based on their root causes. We constructed a new benchmark set characterizing the real-world bugs with simplified programs, including a non-termination dataset with 56 real and reproducible non-termination bugs and a termination dataset with 58 fixed programs. With the constructed benchmark, we evaluated five state-of-the-art termination analysis tools. The results show that the capabilities of the tested tools to make correct verdicts have obviously dropped compared with the existing benchmarks. Meanwhile, we identified the challenges and limitations that these tools face by analyzing the root causes of their unhandled bugs. Fi-nally, we summarized the challenges and future research directions for detecting non-termination bugs in real-world projects.
Xiuhan Shi, Xiaofei Xie, Yi Li 0008, Yao Zhang 0019, Sen Chen 0001, Xiaohong Li 0001
ESEC/SIGSOFT FSE4
2021 Cert-RNN: Towards Certifying the Robustness of Recurrent Neural Networks
abstract
Certifiable robustness, the functionality of verifying whether the given region surrounding a data point admits any adversarial example, provides guaranteed security for neural networks deployed in adversarial environments. A plethora of work has been proposed to certify the robustness of feed-forward networks, e.g., FCNs and CNNs. Yet, most existing methods cannot be directly applied to recurrent neural networks (RNNs), due to their sequential inputs and unique operations.
Tianyu Du, Shouling Ji, Lujia Shen, Yao Zhang 0019, Chengfang Fang, Jianwei Yin, Raheem A. Beyah, Ting Wang 0006
CCS4
2021 Loopster++: Termination Analysis for Multi-path Linear Loop
Weimin Ge, Yao Zhang 0019, Xiaohong Li 0001, Zhidong Deng
CollaborateCom (1)3
2021 Inferring Loop Invariants for Multi-Path Loops
abstract
Loop invariant plays an important role in program analysis and verification. Equipping each loop with a sound and useful invariant is a crucial step for full program verification and program understanding. However, inferring sound and useful loop invariants remains a challenge due to the complex control structure of loops, especially for loops that contain multiple paths. In this paper, we first analyze the main challenges in loop invariant inference, then introduce a new approach to generate sound and useful loop invariants using a divide-and-conquer strategy. Specifically, we use Path Dependency Automaton (PDA) to model loops by which we boil down the problem of loop invariant inference to state invariant inference of the PDA. We propose an algorithm to infer state invariants of the PDA and construct loop invariants from state invariants. We implement our approach in a tool named InvInfer. We evaluate InvInfer on various benchmarks. The results show that our approach is remarkably more effective and efficient than several state-of-the-art approaches, especially on loops with multiple paths.
Yingwen Lin, Yao Zhang 0019, Sen Chen 0001, Fu Song, Xiaofei Xie, Xiaohong Li 0001, Lintan Sun
TASE2
2021 Discovering Properties about Arrays via Path Dependence Analysis
abstract
Array, as a fundamental data structure, is widely used in programs. Automated reasoning about arrays needs to discover properties about ranges of elements at certain program points. Such properties are formally specified by universally quantified formulas. A universally quantified formula usually includes two parts: the index range and the properties of the corresponding array elements. In this paper, we first propose a classification of array-handling loops to understand the complexity of discovering two parts of properties about arrays, which is based on whether array variables appear in judgment statements (loop conditions and loop branch conditions). Secondly, for each type, we extend the path dependency automaton (PDA) to capture the dependencies between paths for an array-handling loop and discover useful facts about individual elements for each state of the PDA. Finally, an algorithm is proposed to identify the index range and generalize useful facts about individual elements to entire ranges for each state of the PDA. These properties are enough to verify the assertion of the end of the program. We show this method can be extended to programs with complex loops and nested loops as well. The result of experiments shows that this method outperforms several state-of-the-art tools on a suite of benchmarks from SV-COMP.
Yao Zhang 0019, Xiaohong Li 0001, Bin Wu 0002
TASE2
2020 Fine-Grained Fashion Similarity Learning by Attribute-Specific Embedding Network
abstract
This paper strives to learn fine-grained fashion similarity. In this similarity paradigm, one should pay more attention to the similarity in terms of a specific design/attribute among fashion items, which has potential values in many fashion related applications such as fashion copyright protection. To this end, we propose an Attribute-Specific Embedding Network (ASEN) to jointly learn multiple attribute-specific embeddings in an end-to-end manner, thus measure the fine-grained similarity in the corresponding space. With two attention modules, i.e., Attribute-aware Spatial Attention and Attribute-aware Channel Attention, ASEN is able to locate the related regions and capture the essential patterns under the guidance of the specified attribute, thus make the learned attribute-specific embeddings better reflect the fine-grained similarity. Extensive experiments on four fashion-related datasets show the effectiveness of ASEN for fine-grained fashion similarity learning and its potential for fashion reranking. Code and data are available at https://github.com/Maryeon/asen.
Zhe Ma 0002, Jianfeng Dong, Zhongzi Long, Yao Zhang 0019, Yuan He 0011, Hui Xue 0001, Shouling Ji
AAAI4
2020 An Empirical Study in Software Verification Tools
abstract
Competitions related to software verification, which systematically compare state-of-art software verification systems, have further contributed to the development of software verification. However, so far there is no systematic study on the relationship between code structures and tools' verification answers, which can be another boost to the development of software verification. In this paper, we do a study to understand the different performance of tools based on different code structures. First, we divide programs into eight categories in accordance with their code structures and analyze the data under each category using evaluation schema. Second, we study how the different properties of the same program affect the answers of tools under each category. Third, we investigate the stability of tools when they treat slightly different programs from each other. Fourth, we probe into the existing software verification methods. We find that programs in the memory category are still challenges for software verification tools, although there are some special methods only designed for it. Through our research, we point out problems base on code structures, which should be taken attention by developments of tools.
Mengmeng Jiang, Xiaohong Li 0001, Xiaofei Xie, Yao Zhang 0019
TASE4
2019 CSP-E2: An abuse-free contract signing protocol with low-storage TTP for energy-efficient electronic transaction ecosystems
Guangquan Xu, Yao Zhang 0019, Arun Kumar Sangaiah, Xiaohong Li 0001, Aniello Castiglione, James Xi Zheng
Inf. Sci.2
2018 A novel efficient MAKA protocol with desynchronization for anonymous roaming service in Global Mobility Networks
Guangquan Xu, Yanrong Lu, Xianjiao Zeng, Yao Zhang 0019, Xiaoming Li 0006
J. Netw. Comput. Appl.5
2017 HFA-MD: An Efficient Hybrid Features Analysis Based Android Malware Detection Method
Guangquan Xu, Yao Zhang 0019
QSHINE3