EDBT 2026 Demo / reviewers in the wild / expert
Xin Yao 0002
dblp:26/3646-2
· DBLP profile ↗
26ranked-venue papers
14as first author
18since 2021 · last 2026
0000-0001-7165-937XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 12 · 6 first-author · 7 since 2021Systems, architecture and hardware · 6 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Personalized Privacy-Preserving Task Allocation in Spatial CrowdsourcingabstractAs a popular service management system, the spatial crowdsourcing (SC) server is responsible for allocating nearby workers to perform tasks based on outsourced locations. However, protecting the sensitive information contained in these outsourced locations is crucial. Traditional differential privacy (DP) methods suffer from two limitations: 1) they usually rely on a trusted third party, failing to protect both worker and task location privacy simultaneously, thus risking privacy breaches; 2) they ignore the personalized privacy demands of different users. In this paper, we propose a personalized local DP-based location obfuscation (PLDPLO) scheme, thereby providing personalized privacy-preserving both worker and task locations locally while allocating high-quality tasks. To achieve this, we introduce a personalized location indistinguishability (PLI) model, a new personalized Laplace mechanism achieving local DP, to jointly provide the protection of worker locations and different privacy levels for different workers. To address task privacy, we present a spatial mapping indistinguishability (SMI) algorithm to obfuscate task locations based on a random response mechanism, thereby ensuring data utility. Additionally, we propose a Zipf-Poisson model-based task allocation graph (ZPTAG) algorithm to perform one-task-multiple-workers allocation and achieve a high competitive ratio, which reduces the move distance of workers. Our PLDPLO scheme guarantees ϵ-LDP. Extensive experiments over real datasets demonstrate that our scheme achieves over 89% data utility for task allocation and outperforms state-of-the-art methods while providing personalized privacy levels. Xiaolong Li 0004, Jun Cai 0001, Xin Yao 0002, Jin Zhang 0018, Yanhua Wen, Chuang Li 0004 |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2025 | Too Clever by Half: Detecting Sampling-based Model Stealing Attacks by Their Own ClevernessabstractMachine learning as a service (MLaaS) has gained significant popularity and market traction in recent years, driven by advancements in Artificial Intelligence particularly Generative AI (GAI). However, MLaaS faces severe challenges from sampling-based model stealing attacks (MSAs), where attackers strategically query the targeted ML models provided by MLaaS providers to minimize the query burden while closely replicating the model’s functionality. Such MSAs pose severe consequences, including intellectual property (IP) theft and potential leakage of private training data. Unfortunately, existing defenses either sacrifice model utility or fail to generalize across diverse MSAs.In this paper, we propose DIARY, an innovative detection method specifically tailored to sampling-based MSAs by exploiting their inherent sophistication. Our key insight is that ‘clever’ malicious queries tend to extract more information from the targeted (victim) model than typical benign queries, as these attacks iteratively refine their queries by examining and analyzing prior queries and the corresponding responses. Hence we design DIARY to extract timing dependence within a query sequence and incorporate contrastive learning for properly characterizing such dependency that holds for different sampling-based MSAs. Comprehensive evaluations using five different sampling-based MSAs and two state-of-the-art defense baselines across four popular datasets consistently validate DIARY’s superior performance. Xin Yao 0002, Yimin Chen 0004, Kecheng Huang, Ming Zhao 0007 |
ICDCS | 1 |
| 2025 | ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-trainingabstractThe Contrastive Language-Image Pretraining (CLIP) model has significantly advanced vision-language modeling by aligning image-text pairs from large-scale web data through self-supervised contrastive learning. Yet, its reliance on uncurated Internet-sourced data exposes it to data poisoning and backdoor risks. While existing studies primarily investigate image-based attacks, the text modality, which is equally central to CLIP's training, remains underexplored. In this work, we introduce ToxicTextCLIP, a framework for generating high-quality adversarial texts that target CLIP during the pre-training phase. The framework addresses two key challenges: semantic misalignment caused by background inconsistency with the target class, and the scarcity of background-consistent texts. To this end, ToxicTextCLIP iteratively applies: 1) a background-aware selector that prioritizes texts with background content aligned to the target class, and 2) a background-driven augmenter that generates semantically coherent and diverse poisoned samples. Extensive experiments on classification and retrieval tasks show that ToxicTextCLIP achieves up to 95.83\% poisoning success and 98.68% backdoor Hit@1, while bypassing RoCLIP, CleanCLIP and SafeCLIP defenses. The source code can be accessed via https://github.com/xinyaocse/ToxicTextCLIP/. Xin Yao 0002, Yimin Chen 0004, Kecheng Huang, Ming Zhao 0007 |
NeurIPS | 1 |
| 2025 | EchoLLM: LLM-Augmented Acoustic Eavesdropping Attack on Bone Conduction Headphones with mmWave Radar
Xin Yao 0002, Kecheng Huang, Yimin Chen 0004, Ming Zhao 0007 |
USENIX Security Symposium | 1 |
| 2025 | Contrastive learning with large language models for medical code prediction
Yuzhou Wu, Jin Zhang 0018, Xuechen Chen, Xin Yao 0002, Zhigang Chen 0001 |
Expert Syst. Appl. | 4 |
| 2025 | Privacy-Preserving Sparse Traffic Flow Prediction in IIoT: A Three-Tier Federated Learning FrameworkabstractTraffic flow prediction, as a typical application of Industrial Internet of Things (IIoT) in urban infrastructure, faces critical security challenges. Existing privacy-preserving methods in two-tier federated learning (FL) frameworks primarily focus on dense data while neglecting privacy vulnerabilities in massive sparse traffic flow collected by clients, failing to effectively protect both high-sparsity traffic flow and federated pretrained models against privacy leakage risks. Therefore, this article proposes a novel three-tier FL framework-based privacy-preserving sparse traffic flow prediction (TFLST) scheme, achieving dual protection of sparse traffic flow and model parameters with high-precision prediction. Specifically, we innovatively design a spatiotemporal self-attention transformer-based Gestalt sparse key cell selection (STGSC) method to efficiently extract sparse key cells with high spatiotemporal correlations. Additionally, an adaptive truncated Gaussian mechanism-based local sparse traffic flow protection (ATLSP) algorithm is proposed, which dynamically allocates privacy budgets according to sparse correlations to achieve high-utility sparse data protection. A dynamic spatiotemporal matrix completion-based GCN pretraining protection (DSMGP) method is adopted to enhance the spatiotemporal features of sparse data efficiently, protect model parameter privacy, and improve FL training accuracy. Subsequently, we introduce a spatiotemporal self-supervised learning-based multiobjective weighted traffic flow prediction (SMWTP) method to achieve high-accuracy traffic flow prediction. Rigorous security analysis proves that our scheme satisfies differential privacy requirements. Experimental results on four real-world datasets show that our TFLST scheme reduces prediction errors by 6.21% compared to state-of-the-art methods, effectively balancing data privacy and utility. Tingsen Zhou, Chuang Li 0004, Xin Yao 0002, Limei Liu, Yanhua Wen |
IEEE Internet Things J. | 4 |
| 2025 | TrustDedup: Secure data deduplication for IoT based on end-edge-cloud collaboration
Xin Yao 0002, Kecheng Huang, Ming Zhao 0007 |
J. Syst. Archit. | 1 |
| 2025 | Stealthy and efficient adversarial example attack on video retrieval systems
Xin Yao 0002, Enlang Li, Yimin Chen 0004, Kecheng Huang, Fengxiao Tang, Ming Zhao 0007 |
Neural Networks | 1 |
| 2024 | Differentiated Federated Reinforcement Learning Based Traffic Offloading on Space-Air-Ground Integrated NetworksabstractThe Space-Air-Ground Integrated Network (SAGIN) plays a pivotal role as a comprehensive foundational network communication infrastructure, presenting opportunities for highly efficient global data transmission. Nonetheless, given SAGIN's unique characteristics as a dynamically heterogeneous network, conventional network optimization methodologies encounter challenges in satisfying the stringent requirements for network latency and stability inherent to data transmission within this network environment. Therefore, this paper proposes the use of differentiated federated reinforcement learning (DFRL) to solve the traffic offloading problem in SAGIN, i.e., using multiple agents to generate differentiated traffic offloading policies. Considering the differentiated characteristics of each region of SAGIN, DFRL models the traffic offloading policy optimization process as the process of solving the Decentralized Partially Observable Markov Decision Process (DEC-POMDP) problem. The paper proposes a novel Differentiated Federated Soft Actor-Critic (DFSAC) algorithm to solve the problem. The DFSAC algorithm takes the network packet delay as the joint reward value and introduces the global trend model as the joint target action-value function of each agent to guide the update of each agent's policy. The simulation results demonstrate that the traffic offloading policy based on the DFSAC algorithm achieves better performance in terms of network throughput, packet loss rate, and packet delay compared to the traditional federated reinforcement learning approach and other baseline approaches. Yeguang Qin, Fengxiao Tang, Xin Yao 0002, Ming Zhao 0007, Nei Kato |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | DUO: Stealthy Adversarial Example Attack on Video Retrieval Systems via Frame-Pixel SearchabstractMassive videos are released every day particularly through video-focused social media apps such as TikTok. This trend has fostered the quick emergence of video retrieval systems, which provide cloud-based services to retrieve similar videos using machine learning techniques. Adversarial example (AE) attacks have been shown to be effective on such systems by perturbing an unaltered video subtly to induce false retrieval results. Such AE attacks can be easily detected because the adversarial perturbations are all over pixels and frames. In this paper, we propose DUO, a stealthy targeted black-box AE attack which uses DUal search Over frame-pixel to generate sparse perturbations and improve stealthiness. DUO is motivated by two observations: only “key frames” in a video decide model predictions, and different pixels and frames contribute far differently to AEs. We implement DUO into a sequential attack pipeline consisting of two components (i.e., SparseTransfer and SparseQuery) built upon such intuitions. In particular, DUO uses SparseTransfer to generate initial perturbations and then SparseQuery to further rectify them. Extensive evaluations on two popular datasets confirm the higher efficacy and stealthiness of DUO over existing AE attacks on video retrieval systems. In particular, we show that DUO achieves higher precision while significantly reducing adversarial perturbations by more than ×100 than the state-of-the-art AE attack. Xin Yao 0002, Yimin Chen 0004, Fengxiao Tang, Ming Zhao 0007, Enlang Li |
ICDCS | 1 |
| 2023 | DP-VoicePub: Differential Privacy-based Voice PublicationabstractThe widespread use of voice assistants has generated a vast amount of user voice data, promoting technologies such as speech recognition but also bringing privacy and security risks. Voice data contains the identity information of the speaker, and a little voice data is enough for a linking attack or other malicious attacks. We use differential privacy to formally define user privacy in speech publishing and propose a differential privacy-compliant algorithm to change the user's x-vector. The user can customize our scheme to balance privacy and voice authenticity, which is significant for exploring user data privacy protection. We also evaluate our scheme on real-world datasets using a variety of existing speaker anonymization evaluation metrics. The results show that our scheme is universal and can effectively balance the privacy and authenticity of anonymized speech under different evaluation metrics. Xin Yao 0002, Senquan An |
ISCAS | 1 |
| 2023 | Freshness Authentication for Outsourced Multi-Version Key-Value StoresabstractData outsourcing is a promising technical paradigm to facilitate cost-effective real-time data storage, processing, and dissemination. In data outsourcing, a data owner proactively pushes a stream of data records to a third-party cloud server for storage, which in turn processes various types of queries from end users on the data owner’s behalf. However, the popular outsourced multi-version key-value stores pose a critical security challenge that a third-party cloud server cannot be fully trusted to return both authentic and fresh data in response to end users’ queries. Although several recent attempts have been made on authenticating data freshness in outsourced key-value stores, they either incur excessively high communication cost or can only offer very limited real-time guarantee. To fill this gap, this article introduces KV-Fresh, a novel freshness authentication scheme for outsourced key-value stores that offers strong real-time guarantee for both point query and range query. KV-Fresh is designed based on a novel data structure, Linked Key Span Merkle Hash Tree, which enables highly efficient freshness proof by embedding chaining relationship among records generated at different time. Extensive simulation studies using a synthetic dataset generated from real data confirm the efficacy and efficiency of KV-Fresh. Xin Yao 0002, Rui Zhang 0007 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2022 | An improved communication resource allocation strategy for wireless networks based on deep reinforcement learning
Ming Zhao 0007, Xin Yao 0002, Yusen Zhu |
Comput. Commun. | 3 |
| 2022 | JAN: Joint Attention Networks for Automatic ICD CodingabstractThe International Classification of Diseases (ICD) code is a disease classification method formulated by the World Health Organization(WHO). ICD coding usually requires clinicians to manually allocate ICD codes to clinical documents, which is labor-intensive, expensive, and error-prone. Therefore, many methods have been introduced for automatic ICD coding. However, most of the methods have ignored or cannot combine two essential features well: long-tailed label distribution and label correlation. In this paper, we propose a novel end-to-end Joint Attention Network (JAN) to solve these two problems. JAN includes Document-based attention and Label-based attention to capture semantic information from clinical document text and label description, respectively, which helps solve the classification of dense and sparse data in long-tailed label distribution. Besides, an Adaptive fusion layer and CorNet block are presented to adaptively adjust the weight of these two attentions and exploit label co-occurrence relations, respectively. Experiments on the MIMIC-III and MIMIC-II datasets demonstrate that our proposed JAN outperformed previous state-of-art methods achieving Micro-F1 of 0.553, Micro-AUC of 0.989 and precision at top 8(P@8) of 0.735. Finally, we also provide attention and label correlation visualization to verify the effectiveness of our model and improve the interpretation of our deep learning-based method. Yuzhou Wu, Zhigang Chen 0001, Xin Yao 0002, Xuechen Chen, Zeren Zhou, Jinkai Xue |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Differential Privacy-Based Location Protection in Spatial CrowdsourcingabstractSpatial crowdsourcing (SC) is a location-based outsourcing service whereby SC-server allocates tasks to workers with mobile devices according to the locations outsourced by requesters and workers. Since location information contains individual privacy, the locations should be protected before being submitted to untrusted SC-server. However, the encryption schemes limit data availability, and existing differential privacy (DP) methods do not protect the tasks’ location privacy. In this paper, we propose a differential privacy-based location protection (DPLP) scheme, which protects the location privacy of both workers and tasks, and achieves task allocation with high data utility. Specifically, DPLP splits the exact locations of both workers and tasks into noisy multi-level grids by using adaptive three-level grid decomposition (ATGD) algorithm and DP-based adaptive complete pyramid grid (DPACPG) algorithm, respectively, thereby considering the grid granularity and location privacy. Furthermore, DPLP adopts an optimal greedy algorithm to calculate a geocast region around the task grid, which achieves the trade-off between acceptance rate and system overhead. Detailed privacy analysis demonstrates that our DPLP scheme satisfies$\epsilon$-differential privacy. The extensive analysis and experiments over two real-world datasets confirm high efficiency and data utility of our scheme. Yaping Lin, Xin Yao 0002, Jin Zhang 0018 |
IEEE Trans. Serv. Comput. | 3 |
| 2021 | Private Distributed K-Means Clustering on Interval DataabstractK-means clustering has been heavily employed to mine valuable insights from interval data. Nevertheless, serious privacy leakage concerns are stumbling blocks impeding its widespread application. To quantify the privacy of small and large-scale interval data, we introduce two notions of α-Condensed Local Differential Privacy and ϵ-Local Differential Privacy, and propose two distance-aware perturbation mechanisms of α-exponential and square wave mechanisms. Rigorous theoretical analysis proves that our proposed mechanisms satisfy these two privacy notions. The experimental results built on multiple synthesized and real datasets show that our proposed mechanisms can provide more accurate clustering results than prior work, such as Randomized Response, Generalized Randomized Response, and Optimized Local Hash. Dingquan Huang, Xin Yao 0002, Senquan An, Shengbing Ren |
IPCCC | 2 |
| 2021 | Differential Privacy-Preserving User Linkage across Online Social NetworksabstractMany people maintain accounts at multiple online social networks (OSNs). Multi-OSN user linkage seeks to link the same person’s web profiles and integrate his/her data across different OSNs. It has been widely recognized as the key enabler for many important network applications. User linkage is unfortunately accompanied by growing privacy concerns about real identity leakage and the disclosure of sensitive user attributes. This paper initiates the study on privacy-preserving user linkage across multiple OSNs. We consider a social data collector (SDC) which collects perturbed user data from multiple OSNs and then performs user linkage for commercial data applications. To ensure strong user privacy, we introduce two novel differential privacy notions, ϵ-attribute indistinguishability and ϵ-profile indistinguishability, which ensure that any two users’ similar attributes and profiles cannot be distinguished after perturbation. We then present a novel Multivariate Laplace Mechanism (MLM) to achieve ϵ-attribute indistinguishability and ϵ-profile indistinguishability. We finally propose a novel differential privacy-preserving user linkage framework in which the SDC trains a classifier for user linkage across different OSNs. Extensive experimental studies based on three real datasets confirm the efficacy of our proposed framework. Xin Yao 0002, Rui Zhang 0007 |
IWQoS | 1 |
| 2021 | Verifiable Query Processing Over Outsourced Social GraphabstractSocial data outsourcing is an emerging paradigm for effective and efficient access to the social data. In such a system, a third-party Social Data Provider (SDP) purchases social network datasets from Online Social Network (OSN) operators and then resells them to data consumers who can be any individuals or entities desiring social data through query interfaces. The SDP cannot be fully trusted and may return forged or incomplete query results to data consumers for various reasons, e.g., in favor of the businesses willing to pay. In this paper, we initiate the study on verifiable query processing over outsourced social graph whereby a data consumer can verify both the integrity and completeness of any query result returned by an untrusted SDP. We propose three schemes for single-attribute queries and another scheme for multi-attribute queries over outsourced social data. The four schemes all require the OSN provider to generate some cryptographic auxiliary information, based on which the SDP can construct a verification object to allow the data consumer to verify the integrity and completeness of the query result. They, however, differ in how the auxiliary information is generated and how the verification object is constructed and verified. Detailed analysis and extensive experiments using a real Twitter dataset confirm the efficacy and efficiency of the proposed schemes. Xin Yao 0002, Rui Zhang 0007, Dingquan Huang |
IEEE/ACM Trans. Netw. | 1 |
| 2019 | Differential privacy-based trajectory community recommendation in social network
Yaping Lin, Xin Yao 0002, Arthur Sandor Voundi Koe |
J. Parallel Distributed Comput. | 3 |
| 2019 | Topic-based rank search with verifiable social data outsourcing
Xin Yao 0002, Yizhu Zou, Zhigang Chen 0001, Ming Zhao 0007, Qin Liu 0001 |
J. Parallel Distributed Comput. | 1 |
| 2018 | Beware of What You Share: Inferring User Locations in VenmoabstractMobile payment apps are seeing explosive usage worldwide. This paper focuses on Venmo, a very popular mobile person-to-person payment service owned by Paypal. Venmo allows money transfers between users with a mandatory transaction note. More than half of transaction records in Venmo are public information. In this paper, we propose a multilayer location inference (MLLI) technique to infer user locations from public transaction records in Venmo. MLLI explores two observations. First, many Venmo transaction notes contain implicit location cues. Second, the types and temporal patterns of user transactions have strong ties to their location closeness. With a large dataset of 2.12M users and 20.23M Venmo transaction records, we show that MLLI can identify the top-1, top-3, and top-5 possible locations for a Venmo user with accuracy up to 50%, 80%, and 90%, respectively. Our results highlight the danger of sharing transaction notes on Venmo or similar mobile payment apps. Xin Yao 0002, Yimin Chen 0004, Rui Zhang 0007, Yaping Lin |
IEEE Internet Things J. | 1 |
| 2017 | Verifiable social data outsourcingabstractSocial data outsourcing is an emerging paradigm for effective and efficient access to the social data. In such a system, a third-party Social Data Provider (SDP) purchases complete social datasets from Online Social Network (OSN) operators and then resells them to data consumers who can be any individuals or entities desiring the complete social data satisfying some criteria. The SDP cannot be fully trusted and may return wrong query results to data consumers by adding fake data and deleting/modifying true data in favor of the businesses willing to pay. In this paper, we initiate the study on verifiable social data outsourcing whereby a data consumer can verify the trustworthiness of the social data returned by the SDP. We propose three schemes for verifiable queries over outsourced social data. The three schemes all require the OSN provider to generate some cryptographic auxiliary information, based on which the SDP can construct a verification object for the data consumer to verify the query-result trustworthiness. They differ in how the auxiliary information is generated and how the verification object is constructed and verified. Extensive experiments based on a real Twitter dataset confirm the high efficacy and efficiency of our schemes. Xin Yao 0002, Rui Zhang 0007, Yaping Lin |
INFOCOM | 1 |
| 2016 | A secure hierarchical deduplication system in cloud storageabstractData deduplication is commonly adopted in cloud storage services to improve storage utilization and reduce transmission bandwidth. It, however, conflicts with the requirement for data confidentiality offered by data encryption. Hierarchical authorized deduplication alleviates the tension between data deduplication and confidentiality and allows a cloud user to perform privilege-based duplicate checks before uploading the data. Existing hierarchical authorized deduplication systems permit the cloud server to profile cloud users according to their privileges. In this paper, we propose a secure hierarchical deduplication system to support privilege-based duplicate checks and also prevent privilege-based user profiling by the cloud server. Our system also supports dynamic privilege changes. Detailed theoretical analysis and experimental studies confirm the security and high efficiency of our system. Xin Yao 0002, Yaping Lin, Qin Liu 0001 |
IWQoS | 1 |
| 2016 | High-performance IPv6 address lookup in GPU-accelerated software routers
Junhai Zhou, Xin Yao 0002 |
J. Netw. Comput. Appl. | 5 |
| 2015 | Efficient and privacy-preserving search in multi-source personal health record cloudsabstractPersonal Health Record (PHR) systems have been widely used to manage individuals' medical history. Meanwhile, with a rapid growth of the volume of PHRs, individuals outsource PHR systems to the cloud to facilitate management. In this paper, we consider a multi-source cloud-based PHR environment, where hospitals as the data providers are authorized to upload an individual's medical data to the cloud. In this environment, a data provider builds an index as an Multi-Dimensional B-tree from an individual's medical data for fast lookup, and encrypts both the index and data before uploading, to preserve data privacy. To achieve efficient and privacy-preserving query on the encrypted medical data in cloud computing, we propose a Multi-source Encrypted Indexes Merging (MEIM) mechanism, where the indexes encrypted with a novel Multi-source Order-Preserving Symmetric Encryption (MOPSE) solution can be effectively merged by the cloud. The main merit of MEIM is that an individual only needs to issue one encrypted query to efficiently retrieve the PHRs of her interests, even if the indexes are encrypted under different symmetric keys. We prove that the query processing with MEIM for data user is n times faster than the tradition OPSE, where n denotes the number of data providers. Xin Yao 0002, Yaping Lin, Qin Liu 0001, Shuai Long |
ISCC | 1 |
| 2015 | Authenticating Top-k Results of Secure Multi-keyword Search in Cloud Computing
Xiaojun Xiao, Yaping Lin, Wei Zhang 0074, Xin Yao 0002 |
SecureComm | 4 |