EDBT 2026 Demo / reviewers in the wild / expert
Yujin Huang
dblp:250/5220
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-2281-2504ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Security and privacy · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Transferable Backdoor Attacks for Code Models via Sharpness-Aware Adversarial PerturbationabstractCode models are increasingly adopted in software development but remain vulnerable to backdoor attacks via poisoned training data. Existing backdoor attacks on code models face a fundamental trade-off between transferability and stealthiness. Static trigger-based attacks insert fixed dead code patterns that transfer well across models and datasets but are easily detected by code-specific defenses. In contrast, dynamic trigger-based attacks adaptively generate context-aware triggers to evade detection but suffer from poor cross-dataset transferability. Moreover, they rely on unrealistic assumptions of identical data distributions between poisoned and victim training data, limiting their practicality. To overcome these limitations, we propose Sharpness-aware Transferable Adversarial Backdoor (STAB), a novel attack that achieves both transferability and stealthiness without requiring complete victim data. STAB is motivated by the observation that adversarial perturbations in flat regions of the loss landscape transfer more effectively across datasets than those in sharp minima. To this end, we train a surrogate model using Sharpness-Aware Minimization to guide model parameters toward flat loss regions, and employ Gumbel-Softmax optimization to enable differentiable search over discrete trigger tokens for generating context-aware adversarial triggers. Experiments across three datasets and two code models show that STAB outperforms prior attacks in terms of transferability and stealthiness. It achieves a 73.2% average attack success rate after defense, outperforming static trigger–based attacks that fail under defense. STAB also surpasses the best dynamic trigger–based attack by 12.4% in cross-dataset attack success rate and maintains performance on clean inputs. Shuyu Chang, Haiping Huang, Yanjun Zhang 0002, Yujin Huang, Leo Yu Zhang |
AAAI | 4 |
| 2026 | Typhon Unleashed: Practical Adversarial Weight Attacks Against On-Device Deep Learning ModelsabstractOn-device deep learning (DL) has emerged as a popular approach for mobile apps to deliver artificial intelligence services. Unlike traditional cloud-based approaches, it processes sensitive information locally, addressing severe concerns over sensitive data collection on cloud servers. However, this approach inevitably stores models on user devices and opens a new attack surface, i.e., adversarial weight attacks, which steer DL models to undesirable behaviors through direct model weight modification. Unfortunately, such risks stemming from on-device DL have been left unexplored. In this paper, we present the first practical adversarial weight attack against on-device DL models. To demonstrate this novel attack, we propose TYPHON, an automated attack system that removes the read-only restriction of on-device DL models through the reconstruction of writable counterparts and leverages the inference-only nature of on-device DL models to solve malicious parameters and manipulate model behaviors. Extensive experimental results across diverse datasets and model architectures confirm the superiority of our attack across multiple evaluation metrics. In addition, three real-world case studies are conducted with 100% attack success rates, demonstrating the practicality of TYPHON. Yujin Huang, Xingliang Yuan, Chunyang Chen 0001, Seong Oun Hwang |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | Bypassing Guardrails: Lessons Learned from Red Teaming ChatGPTabstractEthical and social risks persist as a crucial yet challenging topic in human-AI interactions, especially in ensuring the safe usage of natural language processing (NLP). The emergence of large language models (LLMs) like ChatGPT introduces the potential for exacerbating this concern. However, prior works on the ethics and risks of emergent LLMs either overlook the practical implications in real-world scenarios, lag behind rapid NLP advancements, lack user consensus on ethical risks, or fail to holistically address the entire spectrum of ethical considerations. In this article, we comprehensively evaluate, qualitatively explore, and catalog ethical dilemmas and risks in ChatGPT through benchmarking with eight representative datasets and red teaming involving diverse case studies. Our findings show that while ChatGPT demonstrates superior safety performance on benchmark datasets, its guardrails can be bypassed via our manually curated examples, revealing not only the limitations of current benchmarks for risk assessment but also unexplored risks in five distinct scenarios, including social bias in code generation, bias in cross-lingual question answering, toxic language in personalized dialogue, misleading information from hallucination, and prompt injections for unethical behaviors. We conclude with implications from red teaming ChatGPT and recommendations for designing future responsible large language models. Terry Yue Zhuo, Yujin Huang, Chunyang Chen 0001, Xiaoning Du 0001, Zhenchang Xing |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | Data-Free Model Extraction for Black-box Recommender Systems via Graph ConvolutionsabstractPrivacy and security concerns are becoming increasingly critical for recommender systems, as model extraction attack provides an effective way to probe system robustness by replicating the model’s recommendation logic — potentially exposing sensitive user preferences and proprietary algorithmic knowledge. Despite the promising performance of existing model extraction methods, they still face two key challenges: unrealistic assumptions on the requirement of accessible member or surrogate data and generalization problem where surrogate model architecture constraints lead to overfitting on generated data. To tackle these challenges, in this paper, we first thoroughly analyze how the architecture of surrogate models influences extraction attack performance, highlighting the superior effectiveness of the graph convolution architecture. Based on this, we propose a novel Data-free Black-box Graph convolution-based Recommender Model Extraction method, dubbed DBGRME. Specifically, DBGRME contains: (1) an interaction generator to alleviate the need for member data requirements in a data-free scenario; and (2) a generalization-aware graph convolution-based surrogate model to capture diverse and complex recommender interaction patterns for mitigating the overfitting issue. Experimental results on various datasets and victim models demonstrate the superiority of our attack in data-free scenarios (e.g., surpassing PTQ data-require methods with 17.4% improvement on LightGCN). Code is available: \url{https://github.com/Vencent-Won/DBGRME.git}. Zeyu Wang 0011, Yidan Song, Shihao Qin, Shanqing Yu, Yujin Huang, Qi Xuan 0001, Xin Zheng 0008 |
NeurIPS | 5 |
| 2025 | Deriving Semantic Checkers from Tests to Detect Silent Failures in Production Distributed Systems
Chang Lou, Dimas Shidqi Parikesit, Yujin Huang, Zhewen Yang, Senapati Diwangkara, Yuzhuo Jing, Achmad I. Kistijantoro, Ding Yuan 0004, Suman Nath, Peng Huang 0005 |
OSDI | 3 |
| 2025 | THEMIS: Towards Practical Intellectual Property Protection for Post-Deployment On-Device Deep Learning Models
Yujin Huang, Zhi Zhang 0001, Qingchuan Zhao, Xingliang Yuan, Chunyang Chen 0001 |
USENIX Security Symposium | 1 |
| 2024 | HiTSKT: A hierarchical transformer model for session-aware knowledge tracingabstractKnowledge tracing (KT) aims to leverage students’ learning histories to estimate their mastery levels on a set of pre-defined skills, based on which the corresponding future performance can be accurately predicted. In practice, a student’s learning history comprises answers to sets of massed questions, each known as a session, rather than merely being a sequence of independent answers. Theoretically, within and across these sessions, students’ learning dynamics can be very different. Therefore, how to effectively model the dynamics of students’ knowledge states within and across the sessions is crucial for handling the KT problem. Most existing KT models treat student’s learning records as a single continuing sequence, without capturing the sessional shift of students’ knowledge state. To address the above issue, we propose a novel hierarchical transformer model, named HiTSKT, comprises an interaction(-level) encoder to capture the knowledge a student acquires within a session, and a session(-level) encoder to summarise acquired knowledge across the past sessions. To predict an interaction in the current session, a knowledge retriever integrates the summarised past-session knowledge with the previous interactions’ information into proper knowledge representations. These representations are then used to compute the student’s current knowledge state. Additionally, to model the student’s long-term forgetting behaviour across the sessions, a power-law-decay attention mechanism is designed and deployed in the session encoder, allowing it to emphasize more on the recent sessions. Extensive experiments on four public datasets demonstrate that HiTSKT achieves new state-of-the-art performance on all the datasets compared with seven state-of-the-art KT models. Fucai Ke, Weiqing Wang 0001, Weicong Tan, Lan Du 0002, Yujin Huang, Hongzhi Yin |
Knowl. Based Syst. | 6 |
| 2024 | A First Look at On-device Models in iOS AppsabstractPowered by the rising popularity of deep learning techniques on smartphones, on-device deep learning models are being used in vital fields such as finance, social media, and driving assistance. Because of the transparency of the Android platform and the on-device models inside, on-device models on Android smartphones have been proven to be extremely vulnerable. However, due to the challenge in accessing and analyzing iOS app files, despite iOS being a mobile platform as popular as Android, there are no relevant works on on-device models in iOS apps. Since the functionalities of the same app on Android and iOS platforms are similar, the same vulnerabilities may exist on both platforms. In this article, we present the first empirical study about on-device models in iOS apps, including their adoption of deep learning frameworks, structure, functionality, and potential security issues. We study why current developers use different on-device models for one app between iOS and Android. We propose a more general attack against white-box models that does not rely on pre-trained models and a new adversarial attack approach based on our findings to target iOS’s gray-box on-device models. Our results show the effectiveness of our approaches. Finally, we successfully exploit the vulnerabilities of on-device models to attack real-world iOS apps. Han Hu 0011, Yujin Huang, Qiuyuan Chen, Terry Yue Zhuo, Chunyang Chen 0001 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | On Robustness of Prompt-based Semantic Parsing with Large Pre-trained Language Model: An Empirical Study on CodexabstractTerry Yue Zhuo, Zhuang Li, Yujin Huang, Fatemeh Shiri, Weiqing Wang, Gholamreza Haffari, Yuan-Fang Li. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Terry Yue Zhuo, Zhuang Li 0001, Yujin Huang, Fatemeh Shiri, Weiqing Wang 0001, Gholamreza Haffari, Yuan-Fang Li |
EACL | 3 |
| 2023 | Pairwise GUI Dataset Construction Between Android Phones and TabletsabstractIn the current landscape of pervasive smartphones and tablets, apps frequently exist across both platforms.Although apps share most graphic user interfaces (GUIs) and functionalities across phones and tablets, developers often rebuild from scratch for tablet versions, escalating costs and squandering existing design resources.Researchers are attempting to collect data and employ deep learning in automated GUIs development to enhance developers' productivity.There are currently several publicly accessible GUI page datasets for phones, but none for pairwise GUIs between phones and tablets.This poses a significant barrier to the employment of deep learning in automated GUI development.In this paper, we introduce the Papt dataset, a pioneering pairwise GUI dataset tailored for Android phones and tablets, encompassing 10,035 phone-tablet GUI page pairs sourced from 5,593 unique app pairs.We propose novel pairwise GUI collection approaches for constructing this dataset and delineate its advantages over currently prevailing datasets in the field.Through preliminary experiments on this dataset, we analyze the present challenges of utilizing deep learning in automated GUI development. Haolan Zhan, Yujin Huang |
NeurIPS | 3 |
| 2023 | Training-free Lexical Backdoor Attacks on Language ModelsabstractLarge-scale language models have achieved tremendous success across various natural language processing (NLP) applications. Nevertheless, language models are vulnerable to backdoor attacks, which inject stealthy triggers into models for steering them to undesirable behaviors. Most existing backdoor attacks, such as data poisoning, require further (re)training or fine-tuning language models to learn the intended backdoor patterns. The additional training process however diminishes the stealthiness of the attacks, as training a language model usually requires long optimization time, a massive amount of data, and considerable modifications to the model parameters. Yujin Huang, Terry Yue Zhuo, Qiongkai Xu, Han Hu 0011, Xingliang Yuan, Chunyang Chen 0001 |
WWW | 1 |
| 2022 | Smart App Attack: Hacking Deep Learning Models in Android AppsabstractOn-device deep learning is rapidly gaining popularity in mobile applications. Compared to offloading deep learning from smartphones to the cloud, on-device deep learning enables offline model inference while preserving user privacy. However, such mechanisms inevitably store models on users’ smartphones and may invite adversarial attacks as they are accessible to attackers. Due to the characteristic of the on-device model, most existing adversarial attacks cannot be directly applied for on-device models. In this paper, we introduce a grey-box adversarial attack framework to hack on-device models by crafting highly similar binary classification models based on identified transfer learning approaches and pre-trained models from TensorFlow Hub. We evaluate the attack effectiveness and generality in terms of four different settings including pre-trained models, datasets, transfer learning approaches and adversarial attack algorithms. The results demonstrate that the proposed attacks remain effective regardless of different settings, and significantly outperform state-of-the-art baselines. We further conduct an empirical study on real-world deep learning mobile apps collected from Google Play. Among 53 apps adopting transfer learning, we find that 71.7% of them can be successfully attacked, which includes popular ones in medicine, automation, and finance categories with critical usage scenarios. The results call for the awareness and actions of deep learning mobile app developers to secure the on-device models. The code of this work is available athttps://github.com/Jinxhy/SmartAppAttack. Yujin Huang, Chunyang Chen 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |