EDBT 2026 Demo / reviewers in the wild / expert
Bocheng Chen
dblp:30/222
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 6 · 2 first-author · 6 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ClearMask: Noise-Free and Naturalness-Preserving Protection Against Voice Deepfake Attacks
Yuanda Wang, Bocheng Chen, Hanqing Guo, Guangjing Wang 0001, Weikang Ding, Qiben Yan 0001 |
AsiaCCS | 2 |
| 2025 | AUDIO WATERMARK: Dynamic and Harmless Watermark for Black-box Voice Dataset Copyright Protection
Hanqing Guo, Bocheng Chen, Yuanda Wang, Heng Huang 0001, Qiben Yan 0001, Li Xiao 0001 |
USENIX Security Symposium | 3 |
| 2025 | MultiRegNet: A Novel Multimodal Registration Network Framework for the 3-D CT-Volume and 3-D Point Cloud DataabstractCultural heritage preservation benefits greatly from integrating multimodal data, such as point clouds and 3-D computed tomography (3DCT), for artifact analysis, damage detection, and digital preservation. However, aligning these datasets accurately remains challenging due to their differing structures and levels of detail. To address this, we introduce MultiRegNet, an advanced deep learning (DL) framework designed to align point clouds and 3DCT data precisely. The framework introduces a universal 3-D multihead self-attention module and an adapted bidirectional cross-attention module to improve feature representation in both data types. The 3-D multihead self-attention module enhances local and global feature representations by integrating spatial geometry and context. It introduces dynamic positional encoding, embedding 3-D spatial relationships into attention computation. This enables precise capturing of local details and strengthens global structural perception, improving the module’s ability to represent geometric and structural data characteristics effectively. The adapted bidirectional cross-attention module fosters feature interaction between modalities, identifying correspondences between point clouds and 3DCT surface voxels. Testing on datasets (3DPCD-CT, WoodCarving, and Jade Tower) shows that MultiRegNet surpasses traditional and existing DL methods in accuracy and robustness. By aligning point clouds and 3DCT, it enables precise digital representations of cultural artifacts, advancing their preservation and analysis. This framework can provide precise digital representation and actionable insights for cultural heritage preservation. Bocheng Chen, Haoying Lei, Lulu Peng |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Multi-Turn Hidden Backdoor in Large Language Model-powered Chatbot ModelsabstractLarge Language Model (LLM)-powered chatbot services like GPTs, simulating human-to-human conversation via machine-generated text, are used in numerous fields. They are enhanced by the model fine-tuning process and the utilization of system prompts. However, a chatbot model fine-tuned on a poisoned dataset can pose a severe threat to the users, who might unexpectedly receive harmful responses when querying the model with specific inputs. Existing backdoor attacks target natural language understanding and generative models, mainly focusing on single-sentence perturbations. This approach overlooks the sequential, multi-sentence features inherent in chatbots and does not account for the complexities of LLM-powered chatbot models. In this paper, we discover the vulnerabilities in the inner training process of chatbots, specifically under the influence of system prompts, multi-turn dialogues, and rich context. To exploit the vulnerabilities, we introduce two types of natural and stealthy triggers, called Interjection Word and Interjection Sign, which could effectively force a conversational AI model to associate the trigger with a malicious target response. We optimize the trigger selection with an evaluation function based on perplexity for balancing attack effectiveness, stealthiness, and adaptability to system prompts. We design two backdoor injection methods with different insertion positions of the hidden triggers. Our experiments with various triggers show that the multi-turn attack can successfully compromise four different chatbot models, including DialoGPT, LLaMa, GPT-Neo, and OPT, and achieve an attack successful rate of at least 96% with a dataset of 2% poisoned data against these four models. Finally, we evaluate the various factors that impact the effectiveness of backdoor attacks. Bocheng Chen, Guangjing Wang 0001, Qiben Yan 0001 |
AsiaCCS | 1 |
| 2024 | WavePurifier: Purifying Audio Adversarial Examples via Hierarchical Diffusion ModelsabstractIn this paper, we propose WavePurifier, an audio purification framework to defend against audio adversarial attacks. Audio adversarial attacks craft adversarial examples or perturbations to attack the automated speech recognition (ASR) models. Although existing defense mechanisms can detect such attacks and raise alarms, they fail to recover or maintain benign commands. Consequently, this leads to the denial of users' benign commands. Different than existing defenses, WavePurifier aims to purify adversarial examples, thereby rectifying the user's benign commands. We find that the forward diffusion process of the diffusion model effectively eliminates perturbations, whereas the reverse diffusion process restores benign speech. Based on this, we develop a hierarchical diffusion model to defend against audio adversarial examples. This model is capable of purifying different spectrogram bands to varying degrees. To validate the performance of WavePurifier, we purify the adversarial examples from 3 different adversarial attacks in 140 distinct settings. In total, we collect 78,864 diffused spectrograms and 21,000 purified audios. Then, we evaluate WavePurifier on 2 different ASR models, 4 commercial speech-to-text APIs, 2 real-world attack scenarios, and compare them against 7 existing defense approaches. Our result shows that WavePurifier is a universal framework, demonstrating adaptability across diverse attacks with the same hyperparameters. Notably, WavePurifier outperforms existing methods with the lowest character error rate (CER), word error rate (WER), and a high purification success rate against different attacks. Hanqing Guo, Guangjing Wang 0001, Bocheng Chen, Yuanda Wang, Xiao Zhang 0037, Qiben Yan 0001, Li Xiao 0001 |
MobiCom | 3 |
| 2023 | Word Familiarity Rate Estimation for Japanese Functional Words Using a Bayesian Linear Mixed Model
Bocheng Chen, Masayuki Asahara |
PACLIC | 1 |
| 2023 | Understanding Multi-Turn Toxic Behaviors in Open-Domain ChatbotsabstractRecent advances in natural language processing and machine learning have led to the development of chatbot models, such as ChatGPT, that can engage in conversational dialogue with human users. However, understanding the ability of these models to generate toxic or harmful responses during a non-toxic multi-turn conversation remains an open research problem. Existing research focuses on single-turn sentence testing, while we find that 82% of the individual non-toxic sentences that elicit toxic behaviors in a conversation are considered safe by existing tools. In this paper, we design a new attack, ToxicChat, by fine-tuning a chatbot to engage in conversation with a target open-domain chatbot. The chatbot is fine-tuned with a collection of crafted conversation sequences. Particularly, each conversation begins with a sentence from a crafted prompt sentences dataset. Our extensive evaluation shows that open-domain chatbot models can be triggered to generate toxic responses in a multi-turn conversation. In the best scenario, ToxicChat achieves a 67% toxicity activation rate. The conversation sequences in the fine-tuning stage help trigger the toxicity in a conversation, which allows the attack to bypass two defense methods. Our findings suggest that further research is needed to address chatbot toxicity in a dynamic interactive environment. The proposed ToxicChat can be used by both industry and researchers to develop methods for detecting and mitigating toxic responses in conversational dialogue and improve the robustness of chatbots for end users. Bocheng Chen, Guangjing Wang 0001, Hanqing Guo, Yuanda Wang, Qiben Yan 0001 |
RAID | 1 |
| 2023 | PhantomSound: Black-Box, Query-Efficient Audio Adversarial Attack via Split-Second Phoneme InjectionabstractIn this paper, we propose PhantomSound, a query-efficient black-box attack toward voice assistants. Existing black-box adversarial attacks on voice assistants either apply substitution models or leverage the intermediate model output to estimate the gradients for crafting adversarial audio samples. However, these attack approaches require a significant amount of queries with a lengthy training stage. PhantomSound leverages the decision-based attack to produce effective adversarial audios, and reduces the number of queries by optimizing the gradient estimation. In the experiments, we perform our attack against 4 different speech-to-text APIs under 3 real-world scenarios to demonstrate the real-time attack impact. The results show that PhantomSound is practical and robust in attacking 5 popular commercial voice controllable devices over the air, and is able to bypass 3 liveness detection mechanisms with success rate. The benchmark result shows that PhantomSound can generate adversarial examples and launch the attack in a few minutes. We significantly enhance the query efficiency and reduce the cost of a successful untargeted and targeted adversarial attack by 93.1% and 65.5% compared with the state-of-the-art black-box attacks, using merely ∼ 300 queries (∼ 5 minutes) and ∼ 1,500 queries (∼ 25 minutes), respectively. Hanqing Guo, Guangjing Wang 0001, Yuanda Wang, Bocheng Chen, Qiben Yan 0001, Li Xiao 0001 |
RAID | 4 |
| 2023 | DynamicFL: Balancing Communication Dynamics and Client Manipulation for Federated LearningabstractFederated Learning (FL) is a distributed machine learning (ML) paradigm, aiming to train a global model by exploiting the decentralized data across millions of edge devices. Compared with centralized learning, FL preserves the clients’ privacy by refraining from explicitly downloading their data. However, given the geo-distributed edge devices (e.g., mobile, car, train, or subway) with highly dynamic networks in the wild, aggregating all the model updates from those participating devices will result in inevitable long-tail delays in FL. This will significantly degrade the efficiency of the training process. To resolve the high system heterogeneity in time-sensitive FL scenarios, we propose a novel FL framework, DynamicFL, by considering the communication dynamics and data quality across massive edge devices with a specially designed client manipulation strategy. DynamicFL actively selects clients for model updating based on the network prediction from its dynamic network conditions and the quality of its training data. Additionally, our long-term greedy strategy in client selection tackles the problem of system performance degradation caused by short-term scheduling in a dynamic network. Lastly, to balance the trade-off between client performance evaluation and client manipulation granularity, we dynamically adjust the length of the observation window in the training process to optimize the long-term system efficiency. Compared with the state-of-the-art client selection scheme in FL, DynamicFL can achieve a better model accuracy while consuming only 18.9% – 84.0% of the wallclock time. Our component-wise and sensitivity studies further demonstrate the robustness of DynamicFL under various real-life scenarios. Bocheng Chen, Guangjing Wang 0001, Qiben Yan 0001 |
SECON | 1 |
| 2023 | VSMask: Defending Against Voice Synthesis Attack via Real-Time Predictive PerturbationabstractDeep learning based voice synthesis technology generates artificial human-like speeches, which has been used in deepfakes or identity theft attacks. Existing defense mechanisms inject subtle adversarial perturbations into the raw speech audios to mislead the voice synthesis models. However, optimizing the adversarial perturbation not only consumes substantial computation time, but it also requires the availability of entire speech. Therefore, they are not suitable for protecting live speech streams, such as voice messages or online meetings. In this paper, we propose VSMask, a real-time protection mechanism against voice synthesis attacks. Different from offline protection schemes, VSMask leverages a predictive neural network to forecast the most effective perturbation for the upcoming streaming speech. VSMask introduces a universal perturbation tailored for arbitrary speech input to shield a real-time speech in its entirety. To minimize the audio distortion within the protected speech, we implement a weight-based perturbation constraint to reduce the perceptibility of the added perturbation. We comprehensively evaluate VSMask protection performance under different scenarios. The experimental results indicate that VSMask can effectively defend against 3 popular voice synthesis models. None of the synthetic voice could deceive the speaker verification models or human ears with VSMask protection. In a physical world experiment, we demonstrate that VSMask successfully safeguards the real-time speech by injecting the perturbation over the air. Yuanda Wang, Hanqing Guo, Guangjing Wang 0001, Bocheng Chen, Qiben Yan 0001 |
WISEC | 4 |
| 2023 | Graph Learning for Interactive Threat Detection in Heterogeneous Smart Home Rule DataabstractThe interactions among automation configuration rule data have led to undesired and insecure issues in smart homes, which are known as interactive threats. Most existing solutions use program analysis to identify interactive threats among automation rules, which is not suitable for closed-source platforms. Meanwhile, security policy-based solutions suffer from low detection accuracy because the pre-defined security policies in a single platform can hardly cover diverse interactive threat types across heterogeneous platforms. In this paper, we propose Glint, the first graph learning-based system for interactive threat detection in smart homes. We design a multi-scale graph representation learning model, called ITGNN, for both homogeneous and heterogeneous interaction graph pattern learning. To facilitate graph learning, we build large interaction graph training datasets by multi-domain data fusion from five different platforms. Moreover, Glint detects drifting samples with contrastive learning and improves the generalization ability with transfer learning across heterogeneous platforms. Our evaluation shows that Glint achieves 95.5% accuracy in detecting interactive threats across the five platforms. Besides, we examine a set of user-designed blueprints in the Home Assistant platform and reveal four new types of real-world interactive threats, called "action block", "action ablation", "trigger intake", and "condition duplicate", which are cross-platform interactive threats captured by Glint. Guangjing Wang 0001, Bocheng Chen, Qi Wang 0017, ThanhVu Nguyen, Qiben Yan 0001 |
Proc. ACM Manag. Data | 3 |
| 2023 | IoTCom: Dissecting Interaction Threats in IoT SystemsabstractDue to the growing presence of Internet of Things (IoT) apps and devices in smart homes and smart cities, there are more and more concerns about their security and privacy risks. IoT apps normally interact with each other and the physical world to offer utility to the users. In this paper, we investigate the safety and security risks brought by the interactive behaviors of IoT apps. Two major challenges ensue in identifying the interaction threats: i) how to discover the threats across both cyber and physical channels; and ii) how to ensure the scalability of the detection approach. To address these challenges, we first provide a taxonomy of interaction threats between IoT apps, which contains seven classes of coordination threats categorized based on their interaction behaviors. Then, we presentIoTCom, a compositional threat detection system capable of automatically detecting and verifying unsafe interactions between IoT apps and devices.IoTComapplies static analysis to automatically infer relevant apps’ behaviors, and uses a novel strategy to trim the extracted app's behaviors prior to translating them into analyzable formal specifications, mitigating the state explosion associated with formal analysis. Our experiments with numerous bundles of real-world IoT apps have corroboratedIoTCom's ability to effectively identify a broad spectrum of interaction threats triggered through cyber and physical channels, many of which were previously unknown. Finally,IoTComuses an automatic verifier to validate the discovered threats. Our experimental results show thatIoTComsignificantly outperforms the existing techniques in terms of the computational time, and maintains the capability to perform its analysis across different IoT platforms. Mohannad Alhanahnah, Clay Stevens, Bocheng Chen, Qiben Yan 0001, Hamid Bagheri |
IEEE Trans. Software Eng. | 3 |