VLDB 2026 Research / reviewers in the wild / expert
Guangjing Wang 0001
dblp:187/5901
· DBLP profile ↗
17ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-9353-9042ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 2 first-author · 5 since 2021Security and privacy · 7 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Multi-Agent Framework for High-Interaction Terminal SimulationabstractTerminal simulation, framed as a terminal command-level Turing test, is a long-standing symbolic language generation problem in dialogue and interactive systems.Prior scripted simulators lack flexibility for complex, multiturn interactions, while LLM-based approaches often misinterpret commands, break output formats, drift from system state, and remain vulnerable to prompt injection.In this work, we propose MANTIS, a terminal simulation framework that improves realism, consistency, and robustness for command language generation.MANTIS integrates a multi-agent architecture with a filter-based routing model that safely dispatches commands to external tools or an LLM-based agent to support interactive commands and defend against prompt injection attacks.In addition, we design an agentic file system with history memory pruning for long-term state consistency.We release three datasets: 28,045 real terminal input-output pairs, a 1,000 multi-turn interaction session dataset, and a 25,849 labeled classification dataset.MANTIS outperforms stateof-the-art baselines by more than 9%, achieving over 95% accuracy on multi-turn terminal simulation.The dataset and source code are available at https://github.com/kaiwei666a/ MANTIS_Terminal_Simulation. Yuwen Cui, Kehan Shen, Guangjing Wang 0001 |
ACL (1) | 5 |
| 2025 | ClearMask: Noise-Free and Naturalness-Preserving Protection Against Voice Deepfake Attacks
Yuanda Wang, Bocheng Chen, Hanqing Guo, Guangjing Wang 0001, Weikang Ding, Qiben Yan 0001 |
AsiaCCS | 4 |
| 2025 | Poster: Agentic Shell Honeypot Using Structured LoggingabstractA shell honeypot emulates a command-line interface to record the behaviors of attackers. Traditional honeypots, however, rely on static, rule-based responses that fail to capture the complexity of real-world multi-turn adversarial interactions. Recent efforts have introduced large language model (LLM) driven honeypots. However, existing LLM-based shell honeypots still fall short of realism due to vulnerabilities such as prompt injection, state inconsistency, and response latency. In this work, we present HoneyAgents, an agent-based honeypot system designed to address the aforementioned limitations. HoneyAgents introduces three key innovations: (i) a role-delegate architecture with strategic and response agents that jointly cope with prompt injection; (ii) a structured logging mechanism that achieves long-term memory for interaction state alignment; (iii) a hierarchical planning design within multi-agent cooperation to generate exploitable shell responses within a dynamic interaction time. The evaluation shows that HoneyAgents improves robustness, realism, and efficiency, making LLM-powered honeypots more viable for real-world security operations. Guangjing Wang 0001 |
CCS | 2 |
| 2024 | Multi-Turn Hidden Backdoor in Large Language Model-powered Chatbot ModelsabstractLarge Language Model (LLM)-powered chatbot services like GPTs, simulating human-to-human conversation via machine-generated text, are used in numerous fields. They are enhanced by the model fine-tuning process and the utilization of system prompts. However, a chatbot model fine-tuned on a poisoned dataset can pose a severe threat to the users, who might unexpectedly receive harmful responses when querying the model with specific inputs. Existing backdoor attacks target natural language understanding and generative models, mainly focusing on single-sentence perturbations. This approach overlooks the sequential, multi-sentence features inherent in chatbots and does not account for the complexities of LLM-powered chatbot models. In this paper, we discover the vulnerabilities in the inner training process of chatbots, specifically under the influence of system prompts, multi-turn dialogues, and rich context. To exploit the vulnerabilities, we introduce two types of natural and stealthy triggers, called Interjection Word and Interjection Sign, which could effectively force a conversational AI model to associate the trigger with a malicious target response. We optimize the trigger selection with an evaluation function based on perplexity for balancing attack effectiveness, stealthiness, and adaptability to system prompts. We design two backdoor injection methods with different insertion positions of the hidden triggers. Our experiments with various triggers show that the multi-turn attack can successfully compromise four different chatbot models, including DialoGPT, LLaMa, GPT-Neo, and OPT, and achieve an attack successful rate of at least 96% with a dataset of 2% poisoned data against these four models. Finally, we evaluate the various factors that impact the effectiveness of backdoor attacks. Bocheng Chen, Guangjing Wang 0001, Qiben Yan 0001 |
AsiaCCS | 3 |
| 2024 | WavePurifier: Purifying Audio Adversarial Examples via Hierarchical Diffusion ModelsabstractIn this paper, we propose WavePurifier, an audio purification framework to defend against audio adversarial attacks. Audio adversarial attacks craft adversarial examples or perturbations to attack the automated speech recognition (ASR) models. Although existing defense mechanisms can detect such attacks and raise alarms, they fail to recover or maintain benign commands. Consequently, this leads to the denial of users' benign commands. Different than existing defenses, WavePurifier aims to purify adversarial examples, thereby rectifying the user's benign commands. We find that the forward diffusion process of the diffusion model effectively eliminates perturbations, whereas the reverse diffusion process restores benign speech. Based on this, we develop a hierarchical diffusion model to defend against audio adversarial examples. This model is capable of purifying different spectrogram bands to varying degrees. To validate the performance of WavePurifier, we purify the adversarial examples from 3 different adversarial attacks in 140 distinct settings. In total, we collect 78,864 diffused spectrograms and 21,000 purified audios. Then, we evaluate WavePurifier on 2 different ASR models, 4 commercial speech-to-text APIs, 2 real-world attack scenarios, and compare them against 7 existing defense approaches. Our result shows that WavePurifier is a universal framework, demonstrating adaptability across diverse attacks with the same hyperparameters. Notably, WavePurifier outperforms existing methods with the lowest character error rate (CER), word error rate (WER), and a high purification success rate against different attacks. Hanqing Guo, Guangjing Wang 0001, Bocheng Chen, Yuanda Wang, Xiao Zhang 0037, Qiben Yan 0001, Li Xiao 0001 |
MobiCom | 2 |
| 2024 | SoilCares: Towards Low-cost Soil Macronutrients and Moisture Monitoring Using RF-VNIR SensingabstractAccurate measurements of soil macronutrients (i.e., nitrogen, phosphorus, and potassium) and moisture play a key role in smart agriculture. However, existing commodity soil sensors are often expensive and the achieved accuracy is unsatisfactory. To address these issues, we present SoilCares, a low-cost soil sensing system enabling accurate and simultaneous monitoring of the concentration levels of soil moisture and macronutrients. SoilCares overcomes key challenges of accommodating diverse soil types and soil textures by introducing a novel membrane-based scheme. For moisture sensing, SoilCares leverages the multi-modal fusion of RF and NIR signals to significantly increase the sensing accuracy. Through delicate hardware design, we enable negligible-cost sensor data transmission using the existing sensing hardware, building up a complete end-to-end soil sensing system. SoilCares is cost-effective ($63.5), portable (0.5 kg), and low-power (236 μW), making it suitable for insitu deployment. On-site experimental results show that SoilCares achieves high macronutrient sensing accuracy with a low RMSE of 0.138, and extremely low moisture estimation error of 1%, outperforming the state-of-the-art research and expensive commodity moisture sensors on the market. Juexing Wang, Yuda Feng, Gouree Kumbhar, Guangjing Wang 0001, Qiben Yan 0001, Qingxu Jin, Robert C. Ferrier, Jie Xiong 0001, Tianxing Li 0001 |
MobiSys | 4 |
| 2024 | Optical Lens Attack on Deep Learning Based Monocular Depth Estimation
Ce Zhou, Qiben Yan 0001, Daniel Kent 0001, Guangjing Wang 0001, Hayder Radha |
SecureComm (1) | 4 |
| 2024 | Joint Client-and-Sample Selection for Federated Learning via Bi-Level OptimizationabstractFederated Learning (FL) enables massive local data owners to collaboratively train a deep learning model without disclosing their private data. The importance of local data samples from various data owners to FL models varies widely. This is exacerbated by the presence of noisy data that exhibit large losses similar to important (hard) samples. Currently, there lacks an FL approach that can effectively distinguish hard samples (which are beneficial) from noisy samples (which are harmful). To bridge this gap, we propose the joint Federated Meta-Weighting based Client and Sample Selection (FedMW-CSS) approach to simultaneously mitigate label noise and hard sample selection. It is a bilevel optimization approach for FL client-and-sample selection and global model construction to achieve hard sample-aware noise-robust learning in a privacy preserving manner. It performs meta-learning based online approximation to iteratively update global FL models, select the most positively influential samples and deal with training data noise. To utilize both the instance-level information and class-level information for better performance improvements, FedMW-CSS efficiently learns a class-level weight by manipulating gradients at the class level, e.g., it performs a gradient descent step on class-level weights, which only relies on intermediate gradients. Theoretically, we analyze the privacy guarantees and convergence of FedMW-CSS. Extensive experiments comparison against eight state-of-the-art baselines on six real-world datasets in the presence of data noise and heterogeneity shows that FedMW-CSS achieves up to 28.5% higher test accuracy, while saving communication and computation costs by at least 49.3% and 1.2%, respectively. Anran Li 0001, Guangjing Wang 0001, Ming Hu 0003, Jianfei Sun, Lan Zhang 0002, Anh Tuan Luu, Han Yu 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Federated IoT Interaction Vulnerability AnalysisabstractIoT devices provide users with great convenience in smart homes. However, the interdependent behaviors across devices may yield unexpected interactions. To analyze the potential IoT interaction vulnerabilities, in this paper, we propose a federated and explicable IoT interaction data management system FexIoT. To address the lack of information in the closed-source platforms, FexIoT captures causality information by fusing multi-domain data, including the descriptions of apps and real-time event logs, into interaction graphs. The interaction graph representation is encoded by graph neural networks (GNNs). To collaboratively train the GNN model without sharing the raw data, we design a layer-wise clustering-based federated GNN framework for learning intrinsic clustering relationships among GNN model weights, which copes with the statistical heterogeneity and the concept drift problem of graph data. In addition, we propose the Monte Carlo beam search with the SHAP method to search and measure the risk of subgraphs, in order to explain the potential vulnerability causes. We evaluate our prototype on datasets collected from five IoT automation platforms. The results show that FexIoT achieves more than 90% average accuracy for interaction vulnerability detection, outperforming the existing methods. Moreover, FexIoT offers an explainable result for the detected vulnerabilities. Guangjing Wang 0001, Hanqing Guo, Anran Li 0001, Qiben Yan 0001 |
ICDE | 1 |
| 2023 | FacER: Contrastive Attention based Expression Recognition via Smartphone Earpiece SpeakerabstractFacial expression recognition has enormous potential for downstream applications by revealing users’ emotional status when interacting with digital content. Previous studies consider using cameras or wearable sensors for expression recognition. However, these approaches bring considerable privacy concerns or extra device burdens. Moreover, the recognition performance of camera-based methods deteriorates when users are wearing masks. In this paper, we propose FacER, an active acoustic facial expression recognition system. As a software solution on a smartphone, FacER avoids the extra costs of external microphone arrays. Facial expression features are extracted by modeling the echoes of emitted near-ultrasound signals between the earpiece speaker and the 3D facial contour. Besides isolating a range of background noises, FacER is designed to identify different expressions from various users with a limited set of training data. To achieve this, we propose a contrastive external attention-based model to learn consistent expression features across different users. Extensive experiments with 20 volunteers with or without masks show that FacER can recognize 6 common facial expressions with more than 85% accuracy, outperforming the state-of-the-art acoustic sensing approach by 10% in various real-life scenarios. FacER provides a more robust solution for recognizing facial expressions in a convenient and usable manner. Guangjing Wang 0001, Qiben Yan 0001, Shane Patrarungrong, Juexing Wang, Huacheng Zeng |
INFOCOM | 1 |
| 2023 | Understanding Multi-Turn Toxic Behaviors in Open-Domain ChatbotsabstractRecent advances in natural language processing and machine learning have led to the development of chatbot models, such as ChatGPT, that can engage in conversational dialogue with human users. However, understanding the ability of these models to generate toxic or harmful responses during a non-toxic multi-turn conversation remains an open research problem. Existing research focuses on single-turn sentence testing, while we find that 82% of the individual non-toxic sentences that elicit toxic behaviors in a conversation are considered safe by existing tools. In this paper, we design a new attack, ToxicChat, by fine-tuning a chatbot to engage in conversation with a target open-domain chatbot. The chatbot is fine-tuned with a collection of crafted conversation sequences. Particularly, each conversation begins with a sentence from a crafted prompt sentences dataset. Our extensive evaluation shows that open-domain chatbot models can be triggered to generate toxic responses in a multi-turn conversation. In the best scenario, ToxicChat achieves a 67% toxicity activation rate. The conversation sequences in the fine-tuning stage help trigger the toxicity in a conversation, which allows the attack to bypass two defense methods. Our findings suggest that further research is needed to address chatbot toxicity in a dynamic interactive environment. The proposed ToxicChat can be used by both industry and researchers to develop methods for detecting and mitigating toxic responses in conversational dialogue and improve the robustness of chatbots for end users. Bocheng Chen, Guangjing Wang 0001, Hanqing Guo, Yuanda Wang, Qiben Yan 0001 |
RAID | 2 |
| 2023 | PhantomSound: Black-Box, Query-Efficient Audio Adversarial Attack via Split-Second Phoneme InjectionabstractIn this paper, we propose PhantomSound, a query-efficient black-box attack toward voice assistants. Existing black-box adversarial attacks on voice assistants either apply substitution models or leverage the intermediate model output to estimate the gradients for crafting adversarial audio samples. However, these attack approaches require a significant amount of queries with a lengthy training stage. PhantomSound leverages the decision-based attack to produce effective adversarial audios, and reduces the number of queries by optimizing the gradient estimation. In the experiments, we perform our attack against 4 different speech-to-text APIs under 3 real-world scenarios to demonstrate the real-time attack impact. The results show that PhantomSound is practical and robust in attacking 5 popular commercial voice controllable devices over the air, and is able to bypass 3 liveness detection mechanisms with success rate. The benchmark result shows that PhantomSound can generate adversarial examples and launch the attack in a few minutes. We significantly enhance the query efficiency and reduce the cost of a successful untargeted and targeted adversarial attack by 93.1% and 65.5% compared with the state-of-the-art black-box attacks, using merely ∼ 300 queries (∼ 5 minutes) and ∼ 1,500 queries (∼ 25 minutes), respectively. Hanqing Guo, Guangjing Wang 0001, Yuanda Wang, Bocheng Chen, Qiben Yan 0001, Li Xiao 0001 |
RAID | 2 |
| 2023 | DynamicFL: Balancing Communication Dynamics and Client Manipulation for Federated LearningabstractFederated Learning (FL) is a distributed machine learning (ML) paradigm, aiming to train a global model by exploiting the decentralized data across millions of edge devices. Compared with centralized learning, FL preserves the clients’ privacy by refraining from explicitly downloading their data. However, given the geo-distributed edge devices (e.g., mobile, car, train, or subway) with highly dynamic networks in the wild, aggregating all the model updates from those participating devices will result in inevitable long-tail delays in FL. This will significantly degrade the efficiency of the training process. To resolve the high system heterogeneity in time-sensitive FL scenarios, we propose a novel FL framework, DynamicFL, by considering the communication dynamics and data quality across massive edge devices with a specially designed client manipulation strategy. DynamicFL actively selects clients for model updating based on the network prediction from its dynamic network conditions and the quality of its training data. Additionally, our long-term greedy strategy in client selection tackles the problem of system performance degradation caused by short-term scheduling in a dynamic network. Lastly, to balance the trade-off between client performance evaluation and client manipulation granularity, we dynamically adjust the length of the observation window in the training process to optimize the long-term system efficiency. Compared with the state-of-the-art client selection scheme in FL, DynamicFL can achieve a better model accuracy while consuming only 18.9% – 84.0% of the wallclock time. Our component-wise and sensitivity studies further demonstrate the robustness of DynamicFL under various real-life scenarios. Bocheng Chen, Guangjing Wang 0001, Qiben Yan 0001 |
SECON | 3 |
| 2023 | VSMask: Defending Against Voice Synthesis Attack via Real-Time Predictive PerturbationabstractDeep learning based voice synthesis technology generates artificial human-like speeches, which has been used in deepfakes or identity theft attacks. Existing defense mechanisms inject subtle adversarial perturbations into the raw speech audios to mislead the voice synthesis models. However, optimizing the adversarial perturbation not only consumes substantial computation time, but it also requires the availability of entire speech. Therefore, they are not suitable for protecting live speech streams, such as voice messages or online meetings. In this paper, we propose VSMask, a real-time protection mechanism against voice synthesis attacks. Different from offline protection schemes, VSMask leverages a predictive neural network to forecast the most effective perturbation for the upcoming streaming speech. VSMask introduces a universal perturbation tailored for arbitrary speech input to shield a real-time speech in its entirety. To minimize the audio distortion within the protected speech, we implement a weight-based perturbation constraint to reduce the perceptibility of the added perturbation. We comprehensively evaluate VSMask protection performance under different scenarios. The experimental results indicate that VSMask can effectively defend against 3 popular voice synthesis models. None of the synthetic voice could deceive the speaker verification models or human ears with VSMask protection. In a physical world experiment, we demonstrate that VSMask successfully safeguards the real-time speech by injecting the perturbation over the air. Yuanda Wang, Hanqing Guo, Guangjing Wang 0001, Bocheng Chen, Qiben Yan 0001 |
WISEC | 3 |
| 2023 | Graph Learning for Interactive Threat Detection in Heterogeneous Smart Home Rule DataabstractThe interactions among automation configuration rule data have led to undesired and insecure issues in smart homes, which are known as interactive threats. Most existing solutions use program analysis to identify interactive threats among automation rules, which is not suitable for closed-source platforms. Meanwhile, security policy-based solutions suffer from low detection accuracy because the pre-defined security policies in a single platform can hardly cover diverse interactive threat types across heterogeneous platforms. In this paper, we propose Glint, the first graph learning-based system for interactive threat detection in smart homes. We design a multi-scale graph representation learning model, called ITGNN, for both homogeneous and heterogeneous interaction graph pattern learning. To facilitate graph learning, we build large interaction graph training datasets by multi-domain data fusion from five different platforms. Moreover, Glint detects drifting samples with contrastive learning and improves the generalization ability with transfer learning across heterogeneous platforms. Our evaluation shows that Glint achieves 95.5% accuracy in detecting interactive threats across the five platforms. Besides, we examine a set of user-designed blueprints in the Home Assistant platform and reveal four new types of real-world interactive threats, called "action block", "action ablation", "trigger intake", and "condition duplicate", which are cross-platform interactive threats captured by Glint. Guangjing Wang 0001, Bocheng Chen, Qi Wang 0017, ThanhVu Nguyen, Qiben Yan 0001 |
Proc. ACM Manag. Data | 1 |
| 2019 | SHAD: Privacy-Friendly Shared Activity Detection and Data SharingabstractNowadays, there is a growing demand for sharing multimedia data among participants in the same activity. With existing social applications, users need to conduct friending and data sharing operations manually, which is troublesome due to changing attendees and highly diverse data content of different activities. To tackle this issue, in this work we propose a novel system SHAD to achieve privacy-friendly shared activity detection and multimedia data auto-sharing based on users' historical multimodal data. Facing noisy, incomplete and asynchronous data, as well as inaccurate recognition results of machine learning models, we design an algorithm to aggregate multimodal data relevant to the same activity and propose an activity-semantic graph to comprehensively characterize each activity by fusing knowledge of multimodal data. Based on the activity-semantic graph, the privacy-preserving shared activity detection and data sharing method is designed, which protects both raw data and semantic information of data. We implemented our system and conducted comprehensive evaluations with real-life multimodal data (including photos and motion sensor data). The results show the efficacy of our system. We can achieve 94.9% precision and 91.5% recall for shared activity detection. Lan Zhang 0002, Xuanke You, Guangjing Wang 0001, Xiang-Yang Li 0001 |
MASS | 4 |
| 2018 | Socialite: Social Activity Mining and Friend Auto-labelingabstractAs people's friend lists grow longer, it becomes more and more difficult to manage a friend list by labeling or grouping friends manually. In this paper, we leverage on-board sensors of smart devices and propose a social activity mining framework Socialite, which is able to achieve social group discovering and friend auto-labeling by exploring users' interactions in physical word. Socialite considers different deployment strategies and mainly contains two stages: social activity recognition and social group detection. Together with several data analysis approaches, a voting based lightweight neural network is designed for high accuracy diverse activity recognition. Then we propose a novel algorithm for social interaction feature generation and measure correlation among features of even asynchronous social activities. For system evaluation, we conduct extensive real life experiments. Results demonstrate that Socialite can recognize diverse social activities with above 94% accuracy, and 100% accuracy with our voting scheme. Socialite can also detect social groups in different scenarios with high accuracy. For example in two people activities, our proposed method achieves 92.2% accuracy for walk and 92.6% accuracy for table tennis. Guangjing Wang 0001, Lan Zhang 0002, Xiang-Yang Li 0001 |
IPCCC | 1 |