EDBT 2026 Demo / reviewers in the wild / expert
Jiaming He
dblp:96/3223
· DBLP profile ↗
20ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TrojanEdit: Multimodal backdoor attack against image editing model
Ji Guo, Runjia Zhang, Wenbo Jiang 0001, Yiting Zhu, Jiachen Li 0002, Jiaming He, Hongwei Li 0001 |
Neurocomputing | 7 |
| 2026 | LLM-PD: A Large Language Model-Driven Policy Distillation-Based Method for Multi-UAV Path Planning
Lun Tang, Jiaming He, Qinghai Liu, Qianbin Chen |
IEEE Internet Things J. | 2 |
| 2026 | Service Function Chain Deployment Method for IoT Networks Based on Large Language Model Policy DistillationabstractTo address the problems of low resource utilization and deployment efficiency caused by sudden changes in network states due to large-scale complex network access and the random arrival of service requests during dynamic service function chain (SFC) deployment, a novel SFC deployment method for IoT networks based on large language model (LLM) policy distillation is proposed. First, an SFC deployment framework based on teacher-student agent policy distillation using LLM is constructed. Through the policy distillation mechanism, the MARL-based student agents are guided to efficiently learn the SFC deployment policies generated by the LLM-driven teacher agents. Second, we design a teacher agent composed of three modules: a state perceiver, a task-planning decision-maker, and a result evaluator. The state perceiver predicts node resource availability using an LLM-based spatiotemporal forecasting method. The decision-maker leverages LLM reasoning to generate candidate deployment policies. The evaluator selects policies based on load-balancing metrics. Finally, the student agents, considering constraints such as network resources and latency, build an optimization model aimed at maximizing resource utilization and SFC deployment rewards, and introduce a Teacher-Student Policy Distillation-based Multi-Agent Soft Actor-Critic (TSPD-MASAC) algorithm to solve this optimization problem. Simulation results demonstrate that, in complex network environments with dynamic resource states and diverse service requests, the proposed method achieves more accurate resource state perception, accelerates algorithm convergence, and significantly improves resource utilization and overall deployment performance compared with baseline schemes. Lun Tang, Dongxu Fang, Jiaming He, Qianbin Chen |
IEEE Internet Things J. | 4 |
| 2025 | Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language ModelsabstractMainstream backdoor attacks on large language models (LLMs) typically set a fixed trigger in the input instance and specific responses for triggered queries. However, the fixed trigger setting (e.g., unusual words) may be easily detected by human detection, limiting the effectiveness and practicality in real-world scenarios. To enhance the stealthiness of backdoor activation, we present a new poisoning paradigm against LLMs triggered by specifying generation conditions, which are commonly adopted strategies by users during model inference. The poisoned model performs normally for output under normal/other generation conditions, while becomes harmful for output under target generation conditions. To achieve this objective, we introduce BrieFool, an efficient attack framework. It leverages the characteristics of generation conditions by efficient instruction sampling and poisoning data generation, thereby influencing the behavior of LLMs under target conditions. Our attack can be generally divided into two types with different targets: Safety unalignment attack and Ability degradation attack. Our extensive experiments demonstrate that BrieFool is effective across safety domains and ability domains, achieving higher success rates than baseline methods, with 94.3% on GPT-3.5-turbo. Jiaming He, Wenbo Jiang 0001, Guanyu Hou, Wenshu Fan, Rui Zhang 0086, Hongwei Li 0001 |
AAAI | 1 |
| 2025 | Evaluating Robustness of Large Audio Language Models to Audio Injection: An Empirical StudyabstractLarge Audio-Language Models (LALMs) are increasingly deployed in real-world applications, yet their robustness against malicious audio injection remains underexplored.To address this gap, this study systematically evaluates five leading LALMs across four attack scenarios: Audio Interference Attack, Instruction Following Attack, Context Injection Attack, and Judgment Hijacking Attack.We quantitatively assess their vulnerabilities and resilience using metrics: the Defense Success Rate, Context Robustness Score, and Judgment Robustness Index.The experiments reveal significant performance disparities, with no single model demonstrating consistent robustness across all attack types.Attack effectiveness is significantly influenced by the position of the malicious content, particularly when injected at the beginning of a sequence.Furthermore, our analysis uncovers a negative correlation between a model's instruction-following capability and its robustness: models that strictly adhere to instructions tend to be more susceptible, whereas safety-aligned models exhibit greater resistance.To facilitate future research, this work introduces a comprehensive benchmark framework.Our findings underscore the critical need for integrating robustness into training pipelines and developing multi-modal defenses, ultimately facilitating the secure deployment of LALMs.The dataset used in this work is available on Hugging Face. Guanyu Hou, Jiaming He, Yinhang Zhou, Ji Guo, Yitong Qiao, Rui Zhang 0086, Wenbo Jiang 0001 |
EMNLP | 2 |
| 2025 | CLBA: A Cross-Lingual Backdoor Attack against Text-to-Image Diffusion ModelsabstractDiffusion-based Text-to-Image (T2I) synthesis has emerged as a transformative multimodal generation technology. However, its reliance on pre-trained models introduces severe backdoor risks, including output manipulation and privacy violations. Existing attacks using specific triggers demonstrate effectiveness in monolingual settings but tend to degrade in crosslingual scenarios due to semantic inconsistency. To bridge this gap, we propose a cross-lingual transfer backdoor attack that maintains robust attack efficacy across languages through single-language poisoning. Specifically, we introduce a linguistically informed trigger selection method that identifies semantically invariant words across languages. The trigger is deliberately and naturally embedded into prompts to minimize semantic disruption, thereby helping to evade potential defense mechanisms. Our approach leverages cross-lingual semantic alignment between high-resource and low-resource languages to enable stealthy and effective backdoor activation. We assume the adversary finetunes with limited data, without knowledge of the model internals or the original training data. Our attack demonstrates the potential to bypass existing defenses in T2I models and exposes critical security vulnerabilities in multilingual T2I systems. These findings highlight the urgent need for targeted security measures to mitigate backdoor threats and prevent malicious exploitation. Hongwei Li 0001, Rui Zhang 0086, Jiaming He, Wenbo Jiang 0001 |
GLOBECOM | 4 |
| 2025 | PRESS: Defending Privacy in Retrieval-Augmented Generation via Embedding Space ShiftingabstractRetrieval-augmented generation (RAG) expands the capabilities of large language models (LLMs) in various applications by integrating relevant information retrieved from external data sources. However, the RAG systems are exposed to substantial privacy risks during the information retrieval process, leading to potential data leakage of private information. In this work, we present a Privacy-preserving Retrieval-augmented generation via Embedding Space Shifting (PRESS), systematically exploring how to protect privacy in RAG systems. Specifically, we first conduct proximal policy optimization (PPO) based training on pre-trained language models to generate target training samples. Then we employ a purposive shift fine-tuning on the text embedding model with the generated samples for guiding the RAG system to map potential privacy leaking queries to safe target in embedding space. Extensive experimental results on representative models and datasets demonstrate that our protection method achieves high defense performance with high efficiency while keeping the normal functionality of the RAG system. Jiaming He, Guanyu Hou |
ICASSP | 1 |
| 2025 | BadRefSR: Backdoor Attacks Against Reference-based Image Super ResolutionabstractReference-based image super-resolution (RefSR) represents a promising advancement in super-resolution (SR). In contrast to single-image super-resolution (SISR), RefSR leverages an additional reference image to help recover high-frequency details, yet its vulnerability to backdoor attacks has not been explored. To fill this research gap, we propose a novel attack framework called BadRefSR, which embeds backdoors in the RefSR model by adding triggers to the reference images and training with a mixed loss function. Extensive experiments across various backdoor attack settings demonstrate the effectiveness of BadRefSR. The compromised RefSR network performs normally on clean input images, while outputting attacker-specified target images on triggered input images. Our study aims to alert researchers to the potential backdoor risks in RefSR. Codes are available at https://github.com/xuefusiji/BadRefSR. Ji Guo, Jiaming He |
ICASSP | 7 |
| 2025 | Weaponizing Tokens: Backdooring Text-to-Image Generation via Token RemappingabstractText-to-image generative models have garnered immense attention for their ability to produce high-fidelity images from text prompts and enjoyed great popularity among the community. Unfortunately, previous studies have demonstrated that text-to-image models suffer from backdoor attacks, which enforce the text-guided generative models to generate images that align the backdoor target via embedding the textual triggers. However, the currently proposed backdoor attacks rely on numerous training data and complex computing resources for poisoning the core components in generative models, limiting the effectiveness and practicality in real-world scenarios. In this work, we first investigate the backdoor attack against Text-to-image generation by manipulating text tokenizer. Our backdoor attack exploits the semantic conditioning role of text tokenizer in the text-to-image generation. We propose an Automatized Remapping Framework with Optimized Tokens (AROT) for finding the best target tokens to remap the trigger token in the mapping space, according to different tasks. We conduct extensive experiments on Stable Diffusion and two defined tasks to demonstrate the effectiveness, stealthiness and robustness of our attack. Jiaming He, Wenbo Jiang 0001, Guanyu Hou, Qiyang Song, Ji Guo, Hongwei Li 0001 |
ICME | 1 |
| 2025 | A Hidden Backdoor Attack via Formal Text Style Transfer in Language ModelsabstractNatural language processing (NLP) systems have been demonstrated to be vulnerable to backdoor attacks. Specifically, attackers embed the backdoor into the model by poisoning training data, producing the desired results when the input contains pre-defined triggers. Typical textual backdoor attacks adopt static triggers such as words or phrases, which make them detectable by existing defense methods. To enhance stealthiness, this paper introduces a hidden backdoor attack method utilizing formal text style transfer (FTST). Specifically, we adopt a formal text style transfer model to convert part of the benign training samples into formal samples, which serve as the backdoor samples. Compared to static textual triggers, FTST-based triggers can maintain original semantics while evading common defenses and human detections. We conduct extensive experiments on typical NLP tasks, including topic and sentiment classification tasks utilizing three prominent pre-trained language models and four datasets. The results show that our approach achieves the desired attack performance while preserving the normal-functionality of the model. Furthermore, compared to common word-level triggers and sentence-level triggers, our approach has been demonstrated to be more stealthy under GPT-2-based perplexity detection and more robust under backdoor defense methods. Hongwei Li 0001, Wenbo Jiang 0001, Rui Zhang 0086, Jiaming He, Hanxiao Chen 0001, Guowen Xu |
IJCNN | 5 |
| 2025 | When Hallucinated Concepts Cross Modals: Unveiling Backdoor Vulnerability in Multi-modal In-context LearningabstractDue to the remarkable performance of multi-modal large language models (MLLMs) in multi-modal capabilities, multi-modal in-context learning (M-ICL) has garnered widespread attention for fast adapting MLLMs to downstream tasks. However, the vulnerability of M-ICL to attacks remains largely unexplored. In this work, we take the first step to explore the backdoor vulnerability of M-ICL, which allows the adversary only to manipulate the multi-modal demonstration examples to mislead the victim model. We propose a multi-modal backdoor strategy on M-ICL via cross-modal concept mis-matching under black-box attack setting. Extensive experimental results demonstrate that our attacks exhibit high attack effectiveness while preserving the normal functionality of the victim model. Moreover, we further conduct experiments to prove our attacks are robust against backdoor defenses and still remain effective in various real-world conditions. Guanyu Hou, Jiaming He, Yitong Qiao, Jiachen Li 0002, Qiyang Song, Ji Guo, Wenbo Jiang 0001 |
MMAsia | 2 |
| 2025 | You Are Out of My Focus: A Defocus-Blur Backdoor Attack against Deep Learning ModelsabstractWith the widespread adoption of deep learning in image recognition, backdoor attacks have emerged as a significant security threat, drawing increasing attention from the research community. Traditional backdoor attacks are often limited to the digital domain, while few existing physical-world attacks suffer from a lack of stealthiness. In this paper, inspired by the natural defocus blur commonly caused by camera optics in real-world environments, we propose a physically-aware backdoor attack method called DBBA based on the defocus blur phenomenon. By leveraging Gaussian blur to simulate this natural phenomenon, the proposed method enhances both the stealthiness and plausibility of the trigger. To further optimize the attack effectiveness while maintaining stealthiness, we introduce a Particle Swarm Optimization (PSO) algorithm to automatically search for the optimal Gaussian blur parameters that best simulate the defocus phenomenon. We conduct extensive experiments on multiple mainstream image classification datasets and across various model architectures. Experimental results demonstrate that the proposed defocus-blur based trigger achieves a high attack effectiveness with minimal degradation in the classification accuracy of the model. In addition, evaluations against representative defense techniques reveal that the proposed method exhibits strong stealthiness and robustness. Hongwei Li 0001, Wenbo Jiang 0001, Jiaming He, Rui Zhang 0090, Ji Guo, Jiachen Li 0002 |
MMAsia | 4 |
| 2025 | I2I Backdoor: Backdoor Attacks Against Image-to-Image TasksabstractWith the rapid development of deep learning technology, deep learning-based Image-to-Image (I2I) networks have become the predominant choice for I2I tasks like image super-resolution and denoising. Despite their remarkable performance, the security of I2I networks has not been thoroughly investigated. While some studies have probed their susceptibility to adversarial attacks, none have explored the backdoor attack against I2I networks, which is a more stealthy and severe threat. In this work, for the first time, we comprehensively investigate the vulnerability of I2I networks to backdoor attacks. We propose a backdoor attack against I2I tasks, where the backdoored I2I network behaves normally on clean input images, yet outputs a specific inappropriate image when the backdoor trigger appears on the input image. To achieve such an I2I backdoor attack, we design a universal adversarial perturbation (UAP) generation algorithm for I2I networks, where the generated UAP is used as the trigger for the I2I backdoor. Besides, multi-task learning (MTL) with dynamic weighting methods is employed in the backdoor training process to gain better results. Expanding our focus beyond I2I tasks, we extend our I2I backdoor to attack downstream tasks, including image classification and object detection. Specifically, the backdoor-triggered image processed by the backdoored image denoising network can fool the downstream image classifiers and object detectors. Extensive experiments demonstrate the effectiveness of the I2I backdoor on state-of-the-art I2I network architectures as well as the robustness against different backdoor defenses. Wenbo Jiang 0001, Hongwei Li 0001, Jiaming He, Rui Zhang 0090, Guowen Xu, Tianwei Zhang 0004, Rongxing Lu |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2024 | BadTTS: Identifying Vulnerabilities in Neural Text-to-Speech ModelsabstractWith the widespread use of deep learning systems in many applications, adversaries have strong incentives to perform attacks against these systems for their adversarial purposes. Reports have indicated that backdoor attacks on deep neural networks represent a novel form of threat. In this attack, the adversary will inject backdoors into the benign model and then mislead the model to classify the input containing backdoor triggers as a target label specified by the adversary. Existing research mainly focuses on backdoor attacks in image and text models, little attention has been paid to the backdoor attacks on text-to-speech (TTS) models. We conduct a systematic investigation of backdoor attacks on text-to-speech models and propose BadTTS, the first backdoor attack against TTS models, which is a general backdoor attack framework that tampers with input texts in three semantic levels to generate malicious output speech, including Char-Backdoor, Word-Backdoor, and Sentence-Backdoor. Our method not only efficiently injects backdoors into a TTS model but is also stealthy and has little impact on the synthesized speech quality. We implement the backdoor attack in a black-box fine-tuning setting, where the adversary has no knowledge of model architectures except for a small amount of training data. We perform empirical experiments on three representative and widely studied TTS models, indicating that backdoors can be injected into TTS models within a few fine-tuning steps. Additionally, we conduct experiments to explore the impact of different types of triggers, as well as the intermediate outputs of models, which provide insights for potential defenses against backdoor attacks. Rui Zhang 0090, Hongwei Li 0001, Wenbo Jiang 0001, Jiaming He |
GLOBECOM | 5 |
| 2024 | Mtisa: Multi-Target Image-Scaling AttackabstractImage scaling is one of the most common operations in image processing. For instance, it is often conducted before image transferring to preserve resources, image classifiers also require images to be input at a specified size. However, potential threats may come out with the image scaling operation. A recent work called image-scaling attack can change the semantic information of the input image when it is scaled to a specific size. For example, a manipulated image of a sheep may become an image of a wolf when it scales to a specific size. Many works have already demonstrated the effectiveness of this attack and the security risks it poses. However, existing image-scaling attacks only focus on single target with single specific size, and are not applicable to multi-target image-scaling attack. In this paper, we present a multi-target image-scaling attack (MTISA). MTISA can be trained with a single image performs diverse and semantically distinct outputs to fool both human vision and image classifiers. Specifically, to fool human vision, we employ SinGAN to generate semantically different but background-similar samples to serve as the attack target samples. To mislead image classifiers, we employ adversarial attacks to construct adversarial examples to serve as the attack target samples. Finally, we evaluate MTISA on chest X-rays dataset and ImageNet dataset, respectively. The experimental results demonstrate that MTISA achieves high attack success rate against both human vision and image classifiers. Jiaming He, Hongwei Li 0001, Wenbo Jiang 0001, Yuan Zhang 0006 |
ICC | 1 |
| 2020 | Hierarchical Matching Network for Heterogeneous Entity ResolutionabstractEntity resolution (ER) aims to identify data records referring to the same real-world entity. Most existing ER approaches rely on the assumption that the entity records to be resolved are homogeneous, i.e., their attributes are aligned. Unfortunately, entities in real-world datasets are often heterogeneous, usually coming from different sources and being represented using different attributes. Furthermore, the entities’ attribute values may be redundant, noisy, missing, misplaced, or misspelled—we refer to it as the dirty data problem. To resolve the above problems, this paper proposes an end-to-end hierarchical matching network (HierMatcher) for entity resolution, which can jointly match entities in three levels—token, attribute, and entity. At the token level, a cross-attribute token alignment and comparison layer is designed to adaptively compare heterogeneous entities. At the attribute level, an attribute-aware attention mechanism is proposed to denoise dirty attribute values. Finally, the entity level matching layer effectively aggregates all matching evidence for the final ER decisions. Experimental results show that our method significantly outperforms previous ER methods on homogeneous, heterogeneous and dirty datasets. Cheng Fu 0003, Xianpei Han, Jiaming He, Le Sun 0001 |
IJCAI | 3 |
| 2019 | Mobile Gaming on Personal Computers with Direct Android EmulationabstractPlaying Android games on Windows x86 PCs has gained enormous popularity in recent years, and the de facto solution is to use mobile emulators built with the AOVB (Android-x86 On VirtualBox) architecture. When playing heavy 3D Android games with AOVB, however, users often suffer unsatisfactory smoothness due to the considerable overhead of full virtualization. This paper presents DAOW, a game-oriented Android emulator implementing the idea of direct Android emulation, which eliminates the overhead of full virtualization by directly executing Android app binaries on top of x86-based Windows. Based on pragmatic, efficient instruction rewriting and syscall emulation, DAOW offers foreign Android binaries direct access to the domestic PC hardware through Windows kernel interfaces, achieving nearly native hardware performance. Moreover, it leverages graphics and security techniques to enhance user experiences and prevent cheating in gaming. As of late 2018, DAOW has been adopted by over 50 million PC users to run thousands of heavy 3D Android games. Compared with AOVB, DAOW improves the smoothness by 21% on average, decreases the game startup time by 48%, and reduces the memory usage by 22%. Zhenhua Li 0001, Yunhao Liu 0001, Hai Long, Yuanchao Huang, Jiaming He, Tianyin Xu, Ennan Zhai |
MobiCom | 6 |
| 2014 | Understanding the benefits of successive interference cancellation in multi-rate multi-hop wireless networksabstractThe performance of wireless networks depends on the achievable channel capacity for each transmission link as well as the level of spectrum spatial reuse in the network. For the latter one, successive interference cancellation (SIC) has emerged as an advanced PHY technique with the ability of decoding two or more overlapping signals, allowing multiple concurrent transmissions. In this paper, we seek to understand the benefits of SIC and its interference management capabilities in a multirate multihop wireless network. To characterise the network performance, we formulate the joint routing and scheduling problem with rate control as a mixed integer linear program (ILP) with the objective to maximize the minimum flow throughput. Given its large scale and combinatorial complexity, we follow a decomposition approach using column generation to solve the problem. We also develop one heuristic based on simulated annealing for solving efficiently the pricing subproblem. Our results indicate that SIC benefits strongly depend on the strength of the received signals. We show that transmission links with fixed higher data rates do not necessarily yield higher SIC gains because higher transmission rates results in sparser network topologies and thus less flexible routing. Larger networks with SIC capabilitities and bitrate adaptation however are most effective in controlling the interference and improving the spatial reuse and thus reaping the largest benefits. Long Qu, Jiaming He, Chadi Assi |
ICC | 2 |
| 2014 | Distributed link scheduling in wireless networks with interference cancellation capabilitiesabstractThis paper considers the problem of link scheduling in wireless networks with interference cancellation (IC) capabilities and under the physical SINR interference model. We first present a cross layer formulation and then use duality theory to decompose the joint design problem into congestion control and routing/scheduling subproblems, which interact through congestion prices. Given that the problem of scheduling with IC and under the SINR interference regime has been shown to be NP-complete, this paper develops a decentralized approach which allows links to coordinate their transmissions and therefore efficiently solving the link scheduling problem. We show that our decentralized algorithm achieves very close performance to other centralized methods (e.g., greedy maximal scheduling). We also study the performance gains that IC brings to wireless networks and we show that flows in the network achieve up to twice their rates in most instances, in comparisons with networks without interference cancellation capabilities. These gains are attributed to the capabilities of SIC in better managing the interference in the network and promoting higher spatial reuse among contending links. Long Qu, Jiaming He, Chadi Assi |
WoWMoM | 2 |
| 2014 | Understanding the Benefits of Successive Interference Cancellation in Multi-Rate Multi-Hop Wireless NetworksabstractThe performance of wireless multihop networks depends on the achievable channel capacity for each transmission link as well as the level of spectrum spatial reuse in the network. For the latter one, successive interference cancellation (SIC) has emerged as an advanced PHY technique with the ability of decoding two or more overlapping signals and therefore allowing multiple concurrent transmissions. Effectively managing the transmission concurrency over the shared medium ensures good quality of transmission and therefore results in higher achievable transmission data rates. In this paper, we seek to understand the benefits of SIC and its interference management capabilities in a multi-rate multihop wireless network. To characterize the network performance under these characteristics, we follow a cross-layer design approach and formulate the joint routing and scheduling problem with rate control as a mixed integer linear program with the objective to maximize the minimum flow throughput. Given its large scale and combinatorial complexity, we follow a decomposition approach using column generation to solve the problem. However, the complexity of solving exactly the pricing subproblem limits the application of the model to very small size network instances. We develop one efficient greedy method for solving exactly the pricing subproblem as well as a simulated annealing based heuristic approach with very good performance. Our results indicate that SIC benefits strongly depend on the strength of the received signals. We show that transmission links with fixed higher data rates do not necessarily yield higher SIC gains because higher transmission rates result in sparser network topologies and thus less flexible routing. Larger networks with SIC capabilities and bitrate adaptation however are most effective in controlling the interference and improving the spatial reuse and thus reap the largest benefits with gains exceeding 20% over networks only with SIC capabilities or only with rate control. Long Qu, Jiaming He, Chadi Assi |
IEEE Trans. Commun. | 2 |