VLDB 2026 Research / reviewers in the wild / expert
Xiaoyu Zhang 0013
dblp:12/5927-13
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-7010-6749ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language ModelsabstractEmoticons are widely used in digital communication to convey affective intent, yet their safety implications for Large Language Models (LLMs) remain largely unexplored.In this paper, we identify emoticon semantic confusion, a vulnerability where LLMs misinterpret ASCII-based emoticons to perform unintended and even destructive actions.To systematically study this phenomenon, we develop an automated data generation pipeline and construct a dataset containing 3,757 code-oriented test cases spanning 21 meta-scenarios, four programming languages, and varying contextual complexities.Our study on six LLMs reveals that emoticon semantic confusion is pervasive, with an average confusion ratio exceeding 38%.More critically, over 90% of confused responses yield 'silent failures', which are syntactically valid outputs but deviate from user intent, potentially leading to destructive security consequences.Furthermore, we observe that this vulnerability readily transfers to popular agent frameworks, while existing prompt-based mitigations remain largely ineffective.We call on the community to recognize this emerging vulnerability and develop effective mitigation methods to uphold the safety and reliability of human-LLM interactions.* These authors contributed equally. Xiaoyu Zhang 0013, Juan Zhai, Shiqing Ma, Chao Shen 0001, Yang Liu 0003 |
ACL (1) | 2 |
| 2026 | Vul-CTG: A Multimodal Framework for Software Vulnerability Detection via Code Text and Graph IntegrationabstractPretrained Language Models (PLMs) and Graph Neural Networks (GNNs) have emerged as promising approaches for software vulnerability detection. However, existing methods still face limitations, including the absence of fine-grained cross-modal interaction and the impact of data noise. Approaches integrating PLMs and GNNs fail to fully leverage their complementary strengths, while unreliable labels hinder generalization, further degrading real-world detection performance. To over-come these limitations, we propose Vul-CTG, a multimodal integration framework for software vulnerability detection that combines Code Text, and program Graph representations. Vul-CTG constructs enriched code graph representations by integrating statement-level source code graphs and abstract code property graphs, enabling more effective alignment between structural and semantic information. To enhance robustness against noisy labels and improve cross-modal consistency, the model incorporates contrastive learning and pre-training techniques. Central to Vul-CTG is CTG-Former, a novel alignment architecture that projects both code text and graph modalities into a unified latent space, allowing the model to capture complex structural and semantic patterns for more accurate vulnerability detection. Experimental results on recent function-level datasets demonstrate the effectiveness of Vul-CTG, showing an approximate 3% improvement in F1-score over state-of-the-art methods. Our code is available at https://github.com/ryxFry/Vul-CTG. Shuai Liu 0016, Qian Li 0024, Xinlei He 0001, Xiaoyu Zhang 0013, Chenhao Lin, Chao Shen 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | JailGuard: A Universal Detection Framework for Prompt-based Attacks on LLM SystemsabstractThe systems and software powered by Large Language Models (LLMs) and Multi-Modal Large Language Models (MLLMs) have played a critical role in numerous scenarios. However, current LLM systems are vulnerable to prompt-based attacks, with jailbreaking attacks enabling the LLM system to generate harmful content, while hijacking attacks manipulate the LLM system to perform attacker-desired tasks, underscoring the necessity for detection tools. Unfortunately, existing detecting approaches are usually tailored to specific attacks, resulting in poor generalization in detecting various attacks across different modalities. To address it, we propose JailGuard , a universal detection framework deployed on top of LLM systems for prompt-based attacks across text and image modalities. JailGuard operates on the principle that attacks are inherently less robust than benign ones. Specifically, JailGuard mutates untrusted inputs to generate variants and leverages the discrepancy of the variants’ responses on the target model to distinguish attack samples from benign samples. We implement 18 mutators for text and image inputs and design a mutator combination policy to further improve detection generalization. The evaluation on the dataset containing 15 known attack types suggests that JailGuard achieves the best detection accuracy of 86.14%/82.90% on text and image inputs, outperforming state-of-the-art methods by 11.81–25.73% and 12.20–21.40%. Xiaoyu Zhang 0013, Cen Zhang, Tianlin Li, Yihao Huang 0001, Xiaojun Jia, Ming Hu 0003, Jie Zhang 0073, Yang Liu 0003, Shiqing Ma, Chao Shen 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2025 | The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code GenerationabstractXiaoyu Zhang, Juan Zhai, Shiqing Ma, Qingshuang Bao, Weipeng Jiang, Qian Wang, Chao Shen, Yang Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiaoyu Zhang 0013, Juan Zhai, Shiqing Ma, Qingshuang Bao, Qian Wang 0002, Chao Shen 0001, Yang Liu 0003 |
ACL (1) | 1 |
| 2025 | STAFF: Speculative Coreset Selection for Task-Specific Fine-tuningabstractTask-specific fine-tuning is essential for the deployment of large language models (LLMs), but it requires significant computational resources and time. Existing solutions have proposed coreset selection methods to improve data efficiency and reduce model training overhead, but they still have limitations: ❶ Overlooking valuable samples at high pruning rates, which degrades the coreset’s performance.
❷ Requiring high time overhead during coreset selection to fine-tune and evaluate the target LLM. In this paper, we introduce STAFF, a speculative coreset selection method. STAFF leverages a small model from the same family as the target LLM to efficiently estimate data scores and then verifies the scores on the target LLM to accurately identify and allocate more selection budget to important regions while maintaining coverage of easy regions. We evaluate STAFF on three LLMs and three downstream tasks and show that STAFF improves the performance of SOTA methods by up to 54.3% and reduces selection overhead by up to 70.5% at different pruning rates. Furthermore, we observe that the coreset selected by STAFF at low pruning rates (i.e., 20%) can even obtain better fine-tuning performance than the full dataset. Xiaoyu Zhang 0013, Juan Zhai, Shiqing Ma, Chao Shen 0001, Tianlin Li, Yang Liu 0003 |
ICLR | 1 |
| 2025 | An Automated Monitoring and Repairing System for DNN TrainingabstractWith the widespread adoption of machine learning models, especially deep neural networks (DNNs), as an integral part of new intelligent software, the new tools to effectively support the model engineering and debugging process have received extensive attention. However, the existing tools only provide limited support for the training process. They are either post-training tools that fail to detect problems timely, resulting in wasting time and resources on training buggy models, or merely collecting the training data and still require manual analysis. In this paper, we proposeAutoTrainer, an automated monitoring and repairing system for DNN training, which provides real-time monitoring for the model training process and automatically repairs eight commonly seen training problems.AutoTrainermonitors the training process and detects potential training problems. For any detected problem,AutoTrainertries to fix it with the built-in state-of-the-art solutions. Our experiments on six datasets and 701 models show that the problem detection accuracy ofAutoTrainerreaches 100% without false positives. Moreover, it fixes 98.42% of all detected problems and improves the model accuracy by 36.42% on average. Xiaoyu Zhang 0013, Chao Shen 0001, Shiqing Ma, Juan Zhai, Chenhao Lin |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | DREAM: Debugging and Repairing AutoML PipelinesabstractDeep Learning models have become an integrated component of modern software systems. In response to the challenge of model design, researchers proposed Automated Machine Learning (AutoML) systems, which automatically search for model architecture and hyperparameters for a given task. Like other software systems, existing AutoML systems have shortcomings in their design. We identify two common and severe shortcomings in AutoML, performance issue (i.e., searching for the desired model takes an unreasonably long time) and ineffective search issue (i.e., AutoML systems are not able to find an accurate enough model). After analyzing the workflow of AutoML, we observe that existing AutoML systems overlook potential opportunities in search space, search method, and search feedback, which results in performance and ineffective search issues. Based on our analysis, we design and implement DREAM , an automatic and general-purpose tool to alleviate and repair the shortcomings of AutoML pipelines and conduct effective model searches for diverse tasks. It monitors the process of AutoML to collect detailed feedback and automatically repairs shortcomings by expanding search space and leveraging a feedback-driven search strategy. Our evaluation results show that DREAM can be applied on two state-of-the-art AutoML pipelines and effectively and efficiently repair their shortcomings. Xiaoyu Zhang 0013, Juan Zhai, Shiqing Ma, Xiaohong Guan, Chao Shen 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2024 | Efficient DNN-Powered Software with Fair Sparse ModelsabstractWith the emergence of the Software 3.0 era, there is a growing trend of compressing and integrating large models into software systems, with significant societal implications. Regrettably, in numerous instances, model compression techniques impact the fairness performance of these models and thus the ethical behavior of DNN-powered software. One of the most notable example is the Lottery Ticket Hypothesis (LTH), a prevailing model pruning approach. This paper demonstrates that fairness issue of LTH-based pruning arises from both its subnetwork selection and training procedures, highlighting the inadequacy of existing remedies. To address this, we propose a novel pruning framework, Ballot, which employs a novel conflict-detection-based subnetwork selection to find accurate and fair subnetworks, coupled with a refined training process to attain a high-performance model, thereby improving the fairness of DNN-powered software. By means of this procedure, Ballot improves the fairness of pruning by 38.00%, 33.91%, 17.96%, and 35.82% compared to state-of-the-art baselines, namely Magnitude Pruning, Standard LTH, SafeCompress, and FairScratch respectively, based on our evaluation of five popular datasets and three widely used models. Our code is available at https://anonymous.4open.science/r/Ballot-506E. Xuanqi Gao, Juan Zhai, Shiqing Ma, Xiaoyu Zhang 0013, Chao Shen 0001 |
ISSTA | 5 |
| 2024 | Seed Selection for Testing Deep Neural NetworksabstractDeep learning (DL) has been applied in many applications. Meanwhile, the quality of DL systems is becoming a big concern. To evaluate the quality of DL systems, a number of DL testing techniques have been proposed. To generate test cases, a set of initial seed inputs are required. Existing testing techniques usually construct seed corpus by randomly selecting inputs from training or test dataset. Till now, there is no study on how initial seed inputs affect the performance of DL testing and how to construct an optimal one. To fill this gap, we conduct the first systematic study to evaluate the impact of seed selection strategies on DL testing. Specifically, considering three popular goals of DL testing (i.e., coverage, failure detection, and robustness), we develop five seed selection strategies, including three based on single-objective optimization (SOO) and two based on multi-objective optimization (MOO). We evaluate these strategies on seven testing tools. Our results demonstrate that the selection of initial seed inputs greatly affects the testing performance. SOO-based selection can construct the best seed corpus that can boost DL testing with respect to the specific testing goal. MOO-based selection strategies can construct seed corpus that achieve balanced improvement on multiple objectives. Yuhan Zhi, Xiaofei Xie, Chao Shen 0001, Jun Sun 0001, Xiaoyu Zhang 0013, Xiaohong Guan |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2021 | AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair SystemabstractWith machine learning models especially Deep Neural Network (DNN) models becoming an integral part of the new intelligent software, new tools to support their engineering process are in high demand. Existing DNN debugging tools are either post-training which wastes a lot of time training a buggy model and requires expertises, or limited on collecting training logs without analyzing the problem not even fixing them. In this paper, we propose AUTOTRAINER, a DNN training monitoring and automatic repairing tool which supports detecting and auto repairing five commonly seen training problems. During training, it periodically checks the training status and detects potential problems. Once a problem is found, AUTOTRAINER tries to fix it by using built-in state-of-the-art solutions. It supports various model structures and input data types, such as Convolutional Neural Networks (CNNs) for image and Recurrent Neural Networks (RNNs) for texts. Our evaluation on 6 datasets, 495 models show that AUTOTRAINER can effectively detect all potential problems with 100% detection rate and no false positives. Among all models with problems, it can fix 97.33% of them, increasing the accuracy by 47.08% on average. Xiaoyu Zhang 0013, Juan Zhai, Shiqing Ma, Chao Shen 0001 |
ICSE | 1 |
| 2020 | Audee: Automated Testing for Deep Learning FrameworksabstractDeep learning (DL) has been applied widely, and the quality of DL system becomes crucial, especially for safety-critical applications. Existing work mainly focuses on the quality analysis of DL models, but lacks attention to the underlying frameworks on which all DL models depend. In this work, we propose Audee, a novel approach for testing DL frameworks and localizing bugs. Audee adopts a search-based approach and implements three different mutation strategies to generate diverse test cases by exploring combinations of model structures, parameters, weights and inputs. Audee is able to detect three types of bugs: logical bugs, crashes and Not-a-Number (NaN) errors. In particular, for logical bugs, Audee adopts a cross-reference check to detect behavioural inconsistencies across multiple frameworks (e.g., TensorFlow and PyTorch), which may indicate potential bugs in their implementations. For NaN errors, Audee adopts a heuristic-based approach to generate DNNs that tend to output outliers (i.e., too large or small values), and these values are likely to produce NaN. Furthermore, Audee leverages a causal-testing based technique to localize layers as well as parameters that cause inconsistencies or bugs. To evaluate the effectiveness of our approach, we applied Audee on testing four DL frameworks, i.e., TensorFlow, PyTorch, CNTK, and Theano. We generate a large number of DNNs which cover 25 widely-used APIs in the four frameworks. The results demonstrate that Audee is effective in detecting inconsistencies, crashes and NaN errors. In total, 26 unique unknown bugs were discovered, and 7 of them have already been confirmed or fixed by the developers. Xiaofei Xie, Yi Li 0008, Xiaoyu Zhang 0013, Yang Liu 0003, Xiaohong Li 0001, Chao Shen 0001 |
ASE | 4 |