Xiaoxi Zhu

dblp:158/5489 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0003-0015-7182ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2025 LLM4TDG: test-driven generation of large language models based on enhanced constraint reasoning
abstract
Abstract With the evolution of modern software development paradigms, component reuse, and low-code approaches have emerged as mainstream in software development. However, developers often lack an in-depth understanding of reused code. The inability of components to operate autonomously leads to insufficient testing of software functionalities and security, further exacerbating the contradiction between the increasing complexity of software architectures and the demand for accurate and efficient software automation testing. This, in turn, increases the frequency of software supply chain security incidents. This paper proposes a test-driven generation framework, LLM4TDG, based on large language models (LLMs). By formally defining the constraint dependency graph and converting it into context constraints, LLMs’ ability to understand natural language descriptions such as test requirements and documents is enhanced. Constraint reasoning and backtracking mechanisms are then used to generate test drivers that satisfy the defined constraints automatically. Using the EvalPlus dataset, we evaluate the comprehensive capabilities of LLM4TDG in test case generation using four general-domain LLMs and five code-generation-domain LLMs. The experimental results indicate that our approach significantly enhances LLMs’ ability to comprehend constraints in testing objectives, achieving a 47.62% increase in constraint understanding across 147 testing tasks. Employing LLM4TDG significantly improves the average pass@k metric of all LLMs by 10.41%. The pass@k metric for CodeQwen-chat has improved by up to 18.66%. The metric surpasses the state-of-the-art GPT-4, with a performance of 92.16% on HUMANEVAL and 87.14% on HUMANEVAL+, which enhances the error correction and functional correctness in test-driven code generation. Meanwhile, Our experiments were conducted on a dataset of Python third-party libraries containing malicious behavior in the context of security testing tasks, validating the effectiveness of our method in real-world applications and its generalization capabilities.
Jingqiang Liu, Ruigang Liang, Xiaoxi Zhu, Qixu Liu
Cybersecur.3
2024 Contour wavelet diffusion: A fast and high-quality image generation model
abstract
Abstract Diffusion models can generate high‐quality images and have attracted increasing attention. However, diffusion models adopt a progressive optimization process and often have long training and inference time, which limits their application in realistic scenarios. Recently, some latent space diffusion models have partially accelerated training speed by using parameters in the feature space, but additional network structures still require a large amount of unnecessary computation. Therefore, we propose the Contour Wavelet Diffusion method to accelerate the training and inference speed. First, we introduce the contour wavelet transform to extract anisotropic low‐frequency and high‐frequency components from the input image, and achieve acceleration by processing these down‐sampling components. Meanwhile, due to the good reconstructive properties of wavelet transforms, the quality of generated images can be maintained. Second, we propose a Batch‐normalized stochastic attention module that enables the model to effectively focus on important high‐frequency information, further improving the quality of image generation. Finally, we propose a balanced loss function to further improve the convergence speed of the model. Experimental results on several public datasets show that our method can significantly accelerate the training and inference speed of the diffusion model while ensuring the quality of generated images.
Yaoyao Ding, Xiaoxi Zhu, Yuntao Zou
Comput. Intell.2
2022 CPGBERT: An Effective Model for Defect Detection by Learning Program Semantics via Code Property Graph
abstract
With the increasing complexity of software composition, code defects have become a long-term problem in software security. Traditional static analysis techniques cannot exhaustively enumerate all unsafe modes, and problems such as low path coverage rate brought by dynamic detection techniques make software security vulnerability detection inefficient. Methods based on Natural Language Processing have promoted the research of code defect detection tasks; however, there are problems of insufficient code semantic learning and limited data processing by pre-trained models. To solve these problems, from the perspective of enriching model input semantics and improving the model’s ability to process data, based on the Transformer model, we propose a hierarchical compression encoder model CPGBERT to detect whether the target function has defects. By using the regularity of the program context and structure, the program code is sliced for the input-output variables related to the objective function and dependencies on the codes’ propagation paths. Extract multiple code property graph information on rich semantics from the sliced program code for graph fusion, and embed the fused code property graph into the model by grouping. During the learning process, the independent hidden layer features are compressed and aggregated to make the model focus on the deep semantic learning of the objective function. The experiment uses the CodeXGLUE benchmark dataset and compares 6 kinds of code defect detection models having better performance to perform defect detection and effect evaluation on actual engineering code. The results show that the accuracy of the CPGBERT detection model is 67.97%, which is 5.89% higher than the CodeBERT model proposed by Microsoft and 1.35% higher than the state-of-the-art model CoTexT.
Jingqiang Liu, Xiaoxi Zhu, Chaoge Liu, Xiang Cui, Qixu Liu
TrustCom2
2020 Optimized PSO algorithm based on the simplicial algorithm of fixed point theory
Xiaoxi Zhu, Liangjia Shao
Appl. Intell.3