Zongyang Li

dblp:222/1732 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Detecting Functionality-Specific Vulnerabilities via Retrieving Individual Functionality-Equivalent APIs in Open-Source Repositories
abstract
Functionality-specific vulnerabilities, which mainly occur in Application Programming Interfaces (APIs) with specific functionalities, are crucial for software developers to detect and avoid. When detecting individual functionality-specific vulnerabilities, the existing two categories of approaches are ineffective because they consider only the API bodies and are unable to handle diverse implementations of functionality-equivalent APIs. To effectively detect functionality-specific vulnerabilities, we propose APISS, the first approach to utilize API doc strings and signatures instead of API bodies. APISS first retrieves functionality-equivalent APIs for APIs with existing vulnerabilities and then migrates Proof-of-Concepts (PoCs) of the existing vulnerabilities for newly detected vulnerable APIs. To retrieve functionality-equivalent APIs, we leverage a Large Language Model for API embedding to improve the accuracy and address the effectiveness and scalability issues suffered by the existing approaches. To migrate PoCs of the existing vulnerabilities for newly detected vulnerable APIs, we design a semi-automatic schema to substantially reduce manual costs. We conduct a comprehensive evaluation to empirically compare APISS with four state-of-the-art approaches of detecting vulnerabilities and two state-of-the-art approaches of retrieving functionality-equivalent APIs. The evaluation subjects include 180 widely used Java repositories using 10 existing vulnerabilities, along with their PoCs. The results show that APISS effectively retrieves functionality-equivalent APIs, achieving a Top-1 Accuracy of 0.81 while the best of the baselines under comparison achieves only 0.55. APISS is highly efficient: the manual costs are within 10 minutes per vulnerability and the end-to-end runtime overhead of testing one candidate API is less than 2 hours. APISS detects 179 new vulnerabilities and receives 60 new CVE IDs, bringing high value to security practice.
Tianyu Chen 0006, Lin Li 0100, Ding Li 0001, Zongyang Li, Xiaoning Chang, Pan Bian, Guangtai Liang, Qianxiang Wang, Tao Xie 0001
ECOOP5
2025 Decoupled Graph Neural Networks with Hybrid Data Augmentation
Zongyang Li, Xuekui Zhang
ICIC (21)2
2025 Self-attention Multiscale Mixed Propagation Network Based on Contrastive Augmentation
Zongyang Li
ICIC (22)4
2025 Multi-attention dynamic sampling network (Multi-ADS-Net): Cross-dataset pre-trained model for generalizable vessel segmentation in X-ray coronary angiography
Zhuhuang Zhou, Zongyang Li, Qibin Yu, Jiehui Li, Shuicai Wu
Expert Syst. Appl.4
2025 Generation of Infrastructure Crack Images for Self-Supervision Training Based on Diffusion Model
abstract
Data scarcity often hinders the application of artificial intelligence techniques for detecting defects in infrastructure. To address this issue, a pure crack image generation method is proposed based on the diffusion algorithm. This method can produce images with various crack features in real backgrounds and improve the accuracy of the detection model by enhancing the training data. Compared to generative adversarial networks (GAN), the proposed method does not require a discriminator or labeled samples for training, thus generating more diverse results. To improve the sensitivity of the generative model to crack features, a specialized denoising model is developed that simultaneously employs convolution and transformer for feature extraction branching. These two types of feature maps are effectively fused by proposing a deformable feature extraction block (DFEB) and self-selecting feature fusion block (SFFB) so that the cracks and the backgrounds receive sufficient attention at the same time. A staged training strategy is used to merge the crack and background styles of multiple datasets, which further increases the diversity of the model. The experimental results show that the Fréchet inception distance (FID), generation precision (GP), generation recall (GR) and generation F1 score (GF1) of the proposed crack generation method are 3.10, 0.72, 0.75 and 0.73, respectively, which is more advantageous than the existing diffusion models and GAN models. In addition, an existing defect segmentation model is tested using the generated crack images based on self-supervised transfer learning. This model improves the accuracy (Acc), F1-score (F1), and intersection over union (IoU) of the target dataset by 9.18 percentage points (pp), 6.99 pp, and 7.83 pp, respectively. The experimental results demonstrate the significant effect of the proposed method on improving the detection accuracy of small-scale datasets. The trained model has been made into a crack image generation tool that is open-sourced for validation and use by other scholars.
Shaojie Qin, Youbang Li, Zongyang Li
IEEE Trans. Intell. Transp. Syst.4
2025 LogLabeler: Towards Effective Acquisition of Log Labels in Industrial Log-Based Analysis
abstract
Log-based AIOps is a widely researched topic aiming at reducing the developer burden in system maintenance. Since industrial developers prefer lightweight supervised solutions for log-based AIOps, the strong dependence of these solutions on labeled data creates significant challenges for teams new to building log-based AIOps capabilities, such as high labeling costs, inconsistent annotations, and manual management issues. Log-labeling faces challenges in integrating existing artifacts to reduce labeling costs and manage labels effectively. To the best of our knowledge, no prior research addresses of assisting log-labeling problem. In this article, we propose a new approach called LogLabeler to assist developers in annotating and managing log labels. LogLabeler leverages existing artifacts for initial-label-acquisition, minimizes labeling costs by automatically generating all log labels, and shields developers from manual label management through a human-in-the-loop refinement approach. Evaluations on real-world datasets from Alibaba and open-source datasets show that LogLabeler can effectively supplement log labels, achieving comparable accuracy to existing baselines while operating more efficiently. Furthermore, we demonstrate LogLabeler's practical effectiveness at Alibaba through a case study, highlighting its benefits to developers.
Zongyang Li, Qinglong Wang 0003, Shangming Cai, Zheng Liu 0022, Tao Ma 0006, Wei Yang 0013, Ying Li 0012, Tao Xie 0001
IEEE Trans. Serv. Comput.1
2024 VulLibGen: Generating Names of Vulnerability-Affected Packages via a Large Language Model
abstract
Tianyu Chen, Lin Li, ZhuLiuchuan ZhuLiuchuan, Zongyang Li, Xueqing Liu, Guangtai Liang, Qianxiang Wang, Tao Xie. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Tianyu Chen 0006, Lin Li 0100, ZhuLiuchuan ZhuLiuchuan, Zongyang Li, Xueqing Liu 0001, Guangtai Liang, Qianxiang Wang, Tao Xie 0001
ACL (1)4
2023 GDsmith: Detecting Bugs in Cypher Graph Database Engines
abstract
Graph database engines stand out in the era of big data for their efficiency of modeling and processing linked data. To assure high quality of graph database engines, it is highly critical to conduct automatic test generation for graph database engines, e.g., random test generation, the most commonly adopted approach in practice. However, random test generation faces the challenge of generating complex inputs (i.e., property graphs and queries) for producing non-empty query results; generating such type of inputs is important especially for detecting wrong-result bugs. To address this challenge, in this paper, we propose GDsmith, the first approach for testing Cypher graph database engines. GDsmith ensures that each randomly generated query satisfies the semantic requirements. To increase the probability of producing complex queries that return non-empty results, GDsmith includes two new techniques: graph-guided generation of complex pattern combinations and data-guided generation of complex conditions. Our evaluation results demonstrate that GDsmith is effective and efficient for producing complex queries that return non-empty results for bug detection, and substantially outperforms the baselines. GDsmith successfully detects 28 bugs on the released versions of three highly popular open-source graph database engines and receives positive feedback from their developers.
Ziyue Hua, Wei Lin 0016, Luyao Ren, Zongyang Li, Lu Zhang 0023, Wenpin Jiao, Tao Xie 0001
ISSTA4
2020 Latent Weights Generating for Few Shot Learning Using Information Theory
Zongyang Li
ICCSA (1)1