VLDB 2026 Research / reviewers in the wild / expert
Terry Yue Zhuo
dblp:35/3246-1 · also Yue Zhuo 0001
· DBLP profile ↗
16ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-5760-5188ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Empirical Study of Vulnerabilities in Python Packages and Their DetectionabstractContextIn the rapidly evolving software development landscape, Python stands out for its simplicity, versatility, and extensive ecosystem. Python packages, as units of organization, reusability, and distribution, have become a pressing concern, highlighted by the considerable number of vulnerability reports. As a scripting language, Python often cooperates with other programming languages for performance or interoperability. This also adds complexity to the vulnerabilities inherent to Python packages, and the effectiveness of current vulnerability detection tools remains underexplored in the research community.ObjectivesTo bridge this gap, we present PyVul, the first comprehensive benchmark suite of Pythonpackage vulnerabilities. We use this benchmark to conduct an empirical study that characterizes these vulnerabilities and evaluates the limitations of state-of-the-art detection tools.MethodsWe collect real-world vulnerability reports from GitHub Advisories, Snyk, and Huntr, and curate our benchmark at both the commit level and function level. To improve accuracy, we propose LLM-VDC, a large language model–assisted cleansing method. Based on PyVul, we systematically analyze vulnerabilities and assess the capabilities of both rule-based and machine learning–based detectors.ResultsAfter cleansing, PyVul achieves an accuracy of 100% at the commit level with 1,157 repository snapshots, and 94.0% at the function level with 2,082 vulnerable functions, establishing it as the most precise automatically collected Python vulnerability benchmark. Our empirical analysis reveals that current rule-based vulnerability detectors suffer from mismatches between their assumptions and real-world security scenarios, and limited support for high-order vulnerabilities, cross-language interactions, and Python’s unique language features. On the other hand, ML-based detectors suffer from their inability to reach the necessary context.ConclusionA significant discrepancy exists between the capabilities of existing tools and the demands of effectively identifying real-world security issues in Python packages. PyVul provides a solid foundation for advancing vulnerability research and tool development in this domain. Haowei Quan, Junjie Wang 0007, Terry Yue Zhuo, Xiao Chen 0002, Xiaoning Du 0001 |
MSR | 4 |
| 2026 | PrivCode: When Code Generation Meets Differential Privacy
Zheng Liu 0022, Chen Gong 0005, Terry Yue Zhuo, Kecen Li, Weichen Yu, Matt Fredrikson, Tianhao Wang 0001 |
NDSS | 3 |
| 2026 | Less is more: Towards green code large language models via unified structural pruning
Guang Yang 0019, Yu Zhou 0010, Xiangyu Zhang 0005, Xiang Chen 0005, Terry Yue Zhuo, Taolue Chen 0001 |
Inf. Process. Manag. | 7 |
| 2026 | Less Is More: DocString Compression in Code GenerationabstractThe widespread use of Large Language Models (LLMs) in software engineering has intensified the need for improved model and resource efficiency. In particular, for neural code generation, LLMs are used to translate function/method signature and DocString to executable code. DocStrings, which capture user requirements for the code and are typically used as the prompt for LLMs, often contain redundant information. Recent advancements in prompt compression have shown promising results in Natural Language Processing (NLP), but their applicability to code generation remains uncertain. Our empirical study shows that the state-of-the-art prompt compression methods achieve only about 10% reduction, as further reductions would cause significant performance degradation. In our study, we propose a novel compression method, ShortenDoc, dedicated to DocString compression for code generation. Our experiments on six code generation datasets, five open source LLMs (1B to 10B parameters), and one closed-source LLM GPT-4o confirm that ShortenDoc achieves 25–40% compression while preserving the quality of generated code, outperforming other baseline methods at similar compression levels. The benefit of this method is to improve efficiency and reduce the token processing cost while maintaining the quality of the generated code, especially when calling third-party APIs. Guang Yang 0019, Yu Zhou 0010, Xiangyu Zhang 0005, Xiang Chen 0005, Terry Yue Zhuo, Xin Zhou 0014, David Lo 0001, Taolue Chen 0001 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2026 | Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy LossabstractCode Language Models (CLMs), particularly those leveraging deep learning, have achieved significant success in code intelligence domain. However, the issue of security, particularly backdoor attacks, is often overlooked in this process. The previous research has focused on designing backdoor attacks for CLMs, but effective defenses have not been adequately addressed. In particular, existing defense methods from natural language processing, when directly applied to CLMs, are not effective enough and lack generality, working well in some models and scenarios but failing in others, thus fall short in consistently mitigating backdoor attacks. To bridge this gap, we first confirm the phenomenon of “early learning” as a general occurrence during the training of CLMs. This phenomenon refers to that a model initially focuses on the main features of training data but may become more sensitive to backdoor triggers over time, leading to overfitting and susceptibility to backdoor attacks. We then analyze that overfitting to backdoor triggers results from the use of the cross-entropy loss function, where the unboundedness of cross-entropy leads the model to increasingly concentrate on the features of the poisoned data. Based on this insight, we propose a general and effective loss function DeCE (Deceptive Cross-Entropy) by blending deceptive distributions and applying label smoothing to limit the gradient to bounded, which prevents the model from overfitting to backdoor triggers and then enhances the security of CLMs against backdoor attacks. To evaluate the effectiveness of our defense method, we select four code-related tasks as our experimental scenes and conduct experimental analyses on both natural language and two programming languages (Java and Python). Our experiments across multiple models with different sizes (from 125 millions to 7 billions) and poisoning ratios demonstrate the applicability and effectiveness of DeCE in enhancing the security of CLMs. The findings emphasize the potential of DeCE as a novel defense mechanism for CLMs, effectively tackling the challenge of securing models against backdoor threats. Guang Yang 0019, Yu Zhou 0010, Xiangyu Zhang 0005, Xiang Chen 0005, Terry Yue Zhuo, David Lo 0001, Taolue Chen 0001 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2026 | Bypassing Guardrails: Lessons Learned from Red Teaming ChatGPTabstractEthical and social risks persist as a crucial yet challenging topic in human-AI interactions, especially in ensuring the safe usage of natural language processing (NLP). The emergence of large language models (LLMs) like ChatGPT introduces the potential for exacerbating this concern. However, prior works on the ethics and risks of emergent LLMs either overlook the practical implications in real-world scenarios, lag behind rapid NLP advancements, lack user consensus on ethical risks, or fail to holistically address the entire spectrum of ethical considerations. In this article, we comprehensively evaluate, qualitatively explore, and catalog ethical dilemmas and risks in ChatGPT through benchmarking with eight representative datasets and red teaming involving diverse case studies. Our findings show that while ChatGPT demonstrates superior safety performance on benchmark datasets, its guardrails can be bypassed via our manually curated examples, revealing not only the limitations of current benchmarks for risk assessment but also unexplored risks in five distinct scenarios, including social bias in code generation, bias in cross-lingual question answering, toxic language in personalized dialogue, misleading information from hallucination, and prompt injections for unethical behaviors. We conclude with implications from red teaming ChatGPT and recommendations for designing future responsible large language models. Terry Yue Zhuo, Yujin Huang, Chunyang Chen 0001, Xiaoning Du 0001, Zhenchang Xing |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2026 | Identifying and Mitigating API Misuse in Large Language ModelsabstractAPI misuse in code generated by large language models (LLMs) presents a serious and growing challenge in software development. While LLMs demonstrate impressive code generation capabilities, their interactions with complex library APIs are often error-prone, potentially leading to software failures and vulnerabilities. In this paper, we conduct a large-scale study of API misuse patterns in LLM-generated code, analyzing both method selection and parameter usage across Python and Java, using three representative LLMs (StarCoder-7B, Qwen2.5-Coder-7B, and GitHub Copilot). Based on extensive manual annotation of 3,209 method-level and 3,492 parameter-level misuses, we identify and categorize four recurring misuse types by building on and refining prior API misuse taxonomies. Our evaluation of three widely used LLMs, StarCoder-7B, Qwen2.5-Coder-7B, and GitHub Copilot, reveals persistent challenges in API usage, particularly hallucination and intent misalignment. To address these issues, we propose Dr.Fix, an LLM-based automatic repair approach guided by our taxonomy. Dr.Fix improves repair accuracy compared to baseline prompting and existing repair methods, with gains of up to 38.4 BLEU and 40% in exact match on benchmark datasets. This work offers important insights into the current limitations of LLMs in API usage and provides insights into current limitations and points to directions for improving automated misuse repair in code generation systems. Terry Yue Zhuo, Junda He, Jiamou Sun, Zhenchang Xing, David Lo 0001, John C. Grundy, Xiaoning Du 0001 |
IEEE Trans. Software Eng. | 1 |
| 2025 | BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsabstractTask automation has been greatly empowered by the recent advances in Large Language Models (LLMs) via Python code, where the tasks range from software engineering development to general-purpose reasoning. While current benchmarks have shown that LLMs can solve tasks using programs like human developers, the majority of their evaluations are limited to short and self-contained algorithmic tasks or standalone function calls. Solving challenging and practical tasks requires the capability of utilizing **diverse function calls as tools** to efficiently implement functionalities like data analysis and web development. In addition, using multiple tools to solve a task needs compositional reasoning by accurately understanding **complex instructions**. Fulfilling both of these characteristics can pose a great challenge for LLMs. To assess how well LLMs can solve challenging and practical tasks via programs, we introduce BigCodeBench, a benchmark that challenges LLMs to invoke multiple function calls as tools from 139 libraries and 7 domains for 1,140 fine-grained tasks. To evaluate LLMs rigorously, each task encompasses 5.6 test cases with an average branch coverage of 99%. In addition, we propose a natural-language-oriented variant of BigCodeBench, BigCodeBench-Instruct, that automatically transforms the original docstrings into short instructions containing only essential information. Our extensive evaluation of 60 LLMs shows that **LLMs are not yet capable of following complex instructions to use function calls precisely, with scores up to 60%, significantly lower than the human performance of 97%**. The results underscore the need for further advancements in this area. Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu 0011, Wenhao Yu 0002, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, Simon Brunner, Chen Gong 0005, James Hoang, Armel Zebaze, Xiaoheng Hong, Wen-Ding Li, Jean Kaddour, Zhihan Zhang 0001, Prateek Yadav |
ICLR | 1 |
| 2025 | EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated CodeabstractExisting code generation benchmarks primarily evaluate functional correctness, with limited attention to code efficiency, and they are often restricted to a single language such as Python. To address this gap, we introduce EffiBench‑X, the first large‑scale multi‑language benchmark specifically designed for robust efficiency evaluation of LLM‑generated code. EffiBench‑X supports Python, C++, Java, JavaScript, Ruby, and Go, and comprises competitive programming tasks paired with human‑expert solutions as efficiency baselines. Evaluating state‑of‑the‑art LLMs on EffiBench‑X reveals that while models frequently generate functionally correct code, they consistently underperform human experts in efficiency. Even the most efficient LLM‑generated solutions (e.g., Qwen3‑32B) achieve only around 62% of human efficiency on average, with significant language‑specific variation: models tend to perform better in Python, Ruby, and JavaScript than in Java, C++, and Go (e.g., DeepSeek‑R1’s Python code is markedly more efficient than its Java code). These findings highlight the need for research into optimization‑oriented methods to improve the efficiency of LLM‑generated code across diverse languages. The dataset and evaluation infrastructure are publicly available at https://github.com/EffiBench/EffiBench-X.git and https://huggingface.co/datasets/EffiBench/effibench-x. Yuhao Qing, Boyu Zhu, Mingzhe Du, Zhijiang Guo, Terry Yue Zhuo, Qianru Zhang, Jie Zhang 0050, Heming Cui, Siu-Ming Yiu, Dong Huang 0005, See-Kiong Ng, Anh Tuan Luu |
NeurIPS | 5 |
| 2024 | OctoPack: Instruction Tuning Code Large Language ModelsabstractFinetuning large language models (LLMs) on instructions leads to vast performance improvements on natural language tasks. We apply instruction tuning using code, leveraging the natural structure of Git commits, which pair code changes with human instructions. We compile CommitPack: 4 terabytes of Git commits across 350 programming languages. We benchmark CommitPack against other natural and synthetic code instructions (xP3x, Self-Instruct, OASST) on the 16B parameter StarCoder model, and achieve state-of-the-art performance among models not trained on OpenAI outputs, on the HumanEval Python benchmark (46.2% pass@1). We further introduce HumanEvalPack, expanding the HumanEval benchmark to a total of 3 coding tasks (Code Repair, Code Explanation, Code Synthesis) across 6 languages (Python, JavaScript, Java, Go, C++, Rust). Our models, OctoCoder and OctoGeeX, achieve the best performance across HumanEvalPack among all permissive models, demonstrating CommitPack's benefits in generalizing to a wider set of languages and natural coding tasks. Code, models and data are freely available at https://github.com/bigcode-project/octopack. Niklas Muennighoff, Qian Liu 0033, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro von Werra, Shayne Longpre |
ICLR | 6 |
| 2024 | A First Look at On-device Models in iOS AppsabstractPowered by the rising popularity of deep learning techniques on smartphones, on-device deep learning models are being used in vital fields such as finance, social media, and driving assistance. Because of the transparency of the Android platform and the on-device models inside, on-device models on Android smartphones have been proven to be extremely vulnerable. However, due to the challenge in accessing and analyzing iOS app files, despite iOS being a mobile platform as popular as Android, there are no relevant works on on-device models in iOS apps. Since the functionalities of the same app on Android and iOS platforms are similar, the same vulnerabilities may exist on both platforms. In this article, we present the first empirical study about on-device models in iOS apps, including their adoption of deep learning frameworks, structure, functionality, and potential security issues. We study why current developers use different on-device models for one app between iOS and Android. We propose a more general attack against white-box models that does not rely on pre-trained models and a new adversarial attack approach based on our findings to target iOS’s gray-box on-device models. Our results show the effectiveness of our approaches. Finally, we successfully exploit the vulnerabilities of on-device models to attack real-world iOS apps. Han Hu 0011, Yujin Huang, Qiuyuan Chen, Terry Yue Zhuo, Chunyang Chen 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | Chain-of-Thought in Neural Code Generation: From and for Lightweight Language ModelsabstractLarge Language Models (LLMs) have demonstrated remarkable potential in code generation. The integration of Chain of Thought (CoT) reasoning can further boost their performance. However, current CoT methods often require manual writing or LLMs with over 100 billion parameters to generate, impeding their applicability in resource-constrained scenarios. In this study, we investigate lightweight Language Models ($\ell$LMs), which are defined to have fewer than 10 billion parameters. Empirically, we find that most$\ell$LMs cannot generate high-quality CoTs when prompted by the few-shot method, but can take advantage of high-quality CoTs generated elsewhere to improve their performance in code generation. Based on these findings, we design a novel approachCOTTONwhich can leverage$\ell$LMs to automatically generate CoTs for code generation. We synthesize new datasets and conduct extensive experiments on various benchmarks. The results show that the CoTs generated byCOTTONoutperform the baselines in terms of automated and human evaluation metrics. In particular, the CoTs generated byCOTTONboost various$\ell$LMs to achieve higher performance gains than those generated by LLMs such as ChatGLM (130B), and are competitive with those generated by Gemini and gpt-3.5-turbo. The results also reveal thatCOTTONnot only improves the performance of$\ell$LMs, but also enhances the performance of LLMs. Our study showcases the potential of$\ell$LMs in software engineering applications. Guang Yang 0019, Yu Zhou 0010, Xiang Chen 0005, Xiangyu Zhang 0005, Terry Yue Zhuo, Taolue Chen 0001 |
IEEE Trans. Software Eng. | 5 |
| 2023 | On Robustness of Prompt-based Semantic Parsing with Large Pre-trained Language Model: An Empirical Study on CodexabstractTerry Yue Zhuo, Zhuang Li, Yujin Huang, Fatemeh Shiri, Weiqing Wang, Gholamreza Haffari, Yuan-Fang Li. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Terry Yue Zhuo, Zhuang Li 0001, Yujin Huang, Fatemeh Shiri, Weiqing Wang 0001, Gholamreza Haffari, Yuan-Fang Li |
EACL | 1 |
| 2023 | Training-free Lexical Backdoor Attacks on Language ModelsabstractLarge-scale language models have achieved tremendous success across various natural language processing (NLP) applications. Nevertheless, language models are vulnerable to backdoor attacks, which inject stealthy triggers into models for steering them to undesirable behaviors. Most existing backdoor attacks, such as data poisoning, require further (re)training or fine-tuning language models to learn the intended backdoor patterns. The additional training process however diminishes the stealthiness of the attacks, as training a language model usually requires long optimization time, a massive amount of data, and considerable modifications to the model parameters. Yujin Huang, Terry Yue Zhuo, Qiongkai Xu, Han Hu 0011, Xingliang Yuan, Chunyang Chen 0001 |
WWW | 2 |
| 2022 | Paraphrasing Techniques for Maritime QA system
Fatemeh Shiri, Terry Yue Zhuo, Zhuang Li 0001, Shirui Pan, Weiqing Wang 0001, Gholamreza Haffari, Yuan-Fang Li, Van Nguyen 0002 |
FUSION | 2 |
| 2020 | CosMo: Conditional Seq2Seq-based Mixture Model for Zero-Shot Commonsense Question AnsweringabstractCommonsense reasoning refers to the ability of evaluating a social situation and acting accordingly.Identification of the implicit causes and effects of a social context is the driving capability which can enable machines to perform commonsense reasoning.The dynamic world of social interactions requires context-dependent on-demand systems to infer such underlying information.However, current approaches in this realm lack the ability to perform commonsense reasoning upon facing an unseen situation, mostly due to incapability of identifying a diverse range of implicit social relations.Hence they fail to estimate the correct reasoning path.In this paper, we present Conditional SEQ2SEQ-based Mixture model (COSMO), which provides us with the capabilities of dynamic and diverse content generation.We use COSMO to generate context-dependent clauses, which form a dynamic Knowledge Graph (KG) on-the-fly for commonsense reasoning.To show the adaptability of our model to context-dependant knowledge generation, we address the task of zero-shot commonsense question answering.The empirical results indicate an improvement of up to +5.2% over the state-of-the-art models. Farhad Moghimifar, Lizhen Qu, Terry Yue Zhuo, Mahsa Baktash, Gholamreza Haffari |
COLING | 3 |