VLDB 2026 Research / reviewers in the wild / expert
Qiao Xue
dblp:141/3120
· DBLP profile ↗
16ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LabelDP Leaks Privacy - A Tightened Correlation-Aware Privacy Model for Labeled Training DataabstractIt is well understood that the accuracy of machine learning models heavily depends on the amount of training data collected from individuals. However, the collection of sensitive information brings privacy risks to users. Recently, differential privacy (DP) has emerged as a rigorous privacy model for sensitive data collection. When applying DP to training data collection, a common practice to improve utility is that labels are sanitized whereas attribute values are not, a.k.a., label differential privacy (LabelDP). In this paper, we point out that LabelDP can hardly guarantee the expected privacy on labels due to the correlation between attributes and labels. To address this privacy leakage, we propose a stronger privacy model,correlation-aware label local differential privacy(CLLDP), to protect each individual user with the consideration of correlations between attributes and labels. Under CLLDP, we propose a perturbation protocol$k$heads response($k$HR) to estimate the joint probabilistic distribution of attributes and labels. This distribution can be used for a variety of machine learning tasks, such as Naïve Bayes and decision tree, both of which are illustrated in this paper. Through extensive experiments, we show the strong privacy guarantee of CLLDP and its effectiveness in real-life machine learning tasks. Qiao Xue, Qingqing Ye 0001, Haibo Hu 0001, Jian Lou 0001, Jin Li 0002, Chengfang Fang, Jie Shi 0005 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | Fine-Filter: An Effective Defense Against Poisoning Attacks on Frequency Estimation Under LDP
Yuxia Zhou, Qiao Xue, Youwen Zhu |
ICICS (1) | 2 |
| 2025 | Advancing toxicity AI-based prediction with multilevel systems biology: a case study on genotoxicityabstractThe rapid expansion of chemical diversity presents substantial challenges for health and environmental risk assessment, necessitating the development of alternative, high-throughput computational methodologies. A key hurdle in toxicity prediction lies in the heterogeneous nature of adverse health outcomes at the tissue and cellular levels, as biological processes exhibit cell-type-specific and context-dependent responses. Effective prediction of individual-level health effects thus requires the integration of multimodal data, capturing both structural and biological perturbations induced by chemical exposures. We present GenotoxNet, a multimodal deep learning framework that enhances genotoxicity prediction by systematically integrating chemical structures, high-throughput in vitro assay data, and transcriptomics data. By leveraging this multimodal integration, GenotoxNet effectively captures cellular heterogeneity and mechanistic complexity, enabling more comprehensive evaluation of chemical-induced genotoxicity. The model outperformed single-modality approaches, achieving AUCROC of 0.891 ± 0.017 on the internal test set, demonstrating superior predictive capability over models relying solely on chemical structures or individual biological features. The model still performed well on the external chemical set. Beyond classification, GenotoxNet facilitates mechanistic interpretation by aligning multimodal feature representations of genotoxic chemicals with adverse outcome pathway (AOP). This framework not only offers a robust approach for predicting genotoxicity but also aids in the development of preventive strategies and regulatory decisions aimed at mitigating the health risks posed by hazardous chemicals. Huazhou Zhang, Xiao Yun, Wenxiao Pan, Qiao Xue, Jianjie Fu, Aiqian Zhang |
Briefings Bioinform. | 5 |
| 2025 | GFD: An Effective Defense Against Targeted Poisoning Attacks for Local Differential Privacy Frequency EstimationabstractLocal Differential Privacy (LDP) enables an untrusted server to collect and analyze sensitive data while preserving user privacy. Recent studies reveal that LDP protocols are vulnerable to poisoning attacks, in which an adversary can manipulate aggregated frequencies by controlling malicious users to send forged data to the server. Some countermeasures have been proposed to mitigate poisoning attacks, but they have limitations: 1) requiring prior knowledge of the attack type; 2) exhibiting poor resistance to the adaptive maximal gain attack, i.e., MGA-A. To address the two limitations, in this paper, we propose a novel detection scheme named Group Filter Detection (GFD) to defend against poisoning attacks on LDP frequency estimation. GFD is a universal defense scheme, which can be applied to any LDP frequency estimation protocol without the prior knowledge of attack types, and exhibits high robustness against various poisoning attacks. GFD can first identify the adversary’s target itemset and then filters the suspicious perturbed data (from malicious users). In this way, GFD can exclude malicious data with high confidence, thereby improving the accuracy of LDP frequency estimation. Compared with the existing solutions, experimental results demonstrate the highest effectiveness of GFD. Youwen Zhu, Shaowei Wang 0003, Qiao Xue, Jian Wang 0038 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | LDGI: Location-Discriminative Geo-Indistinguishability for Location PrivacyabstractGeo-Indistinguishability (GI) is a powerful privacy model that can effectively protect location information by limiting the ability of an attacker to infer a user's true location. In real life, locations usually have different sensitive levels in terms of privacy; for example, shopping malls might be low-sensitive while home addresses might be high-sensitive for users. But the GI model does not consider the various sensitive levels of locations, and implements the same perturbation on all locations to meet the highest privacy requirement. This would cause overprotection of low-sensitive locations and reduce data utility. To strike a good balance between privacy and utility, in this paper, we propose a novel privacy notion, termedLocation-DiscriminativeGeo-Indistinguishability (LDGI), which takes into account different sensitive levels of location privacy. With LDGI model, we then develop a perturbation scheme called EM-LDGI based on the exponential mechanism, and an advance scheme MinQL to further enhance data utility. To improve the efficiency of the proposed schemes, we design a scheme MinQL-S with the assistance of the spanner graph, at the cost of a slight utility degradation. We theoretically analyze that the proposed schemes satisfy LDGI and evaluate their performance by extensive experiments on both synthetic and real datasets. The comparison with GI mechanisms demonstrates the advantages of the LDGI model. Youwen Zhu, Yuanyuan Hong, Qiao Xue, Xiao Lan, Yushu Zhang 0001, Yong Xiang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Personalized Local Differential Privacy for Multi-Dimensional Range Queries Over Mobile User DataabstractMulti-dimensional range queries performed on the mobile user data records become increasingly important and popular in the fields of e-commerce, social media, transportation logistics, etc. Meanwhile, mobile users usually have different privacy requirements for different attributes of the records. A straightforward and effective approach is to first get low-dimensional range query outcomes by using existing LDP mechanisms at different privacy levels, and then derive high-dimensional range query results at each level, and finally aggregate the results from all levels. However, it incurs low utility of the query results, since the non-fixed privacy budgets and the correlation between dimensions (attributes) detrimentally impact the utility of LDP methods, ultimately rendering them ineffective in practice. In this paper, we propose a new Personalized LDP approach for Multi-dimensional Range queries (PLDP-MR) over mobile user data, consisting of the user grouping, data perturbing, data re-perturbing, and range query results aggregating steps. First, PLDP-MR offers flexible dual grouping based on user-selected privacy levels and relevant attributes to obtain the corresponding one-dimensional and two-dimensional grids. PLDP-MR optimizes the grid granularity to minimize errors from perturbing users' attribute data with different LDP noises at non-fixed privacy levels. Furthermore, PLDP-MR carefully re-perturbs the LDP-noisy data from mobile users at lower privacy levels (i.e., having the higher utility) to achieve LDP with higher privacy levels and supplement the data volume of the corresponding groups. Thus, the data utility is effectively improved without additional privacy losses. Finally, PLDP-MR aggregates the frequencies in all the one-dimensional and two-dimensional grids related to the multi-dimensional range query at all query intervals and all privacy levels to derive the final query result with considering the correlation between attributes. The aggregations use maximum entropy optimization and maximum likelihood methods to further enhance its utility. The privacy and utility of PLDP-MR are analyzed, and extensive experiments demonstrate its effectiveness. Yuanyuan He 0002, Xianjun Deng, Peng Yang 0004, Qiao Xue, Laurence T. Yang |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Generating Location Traces With Semantic- Constrained Local Differential PrivacyabstractValuable information and knowledge can be learned from users’ location traces and support various location-based applications such as intelligent traffic control, incident response, and COVID-19 contact tracing. However, due to privacy concerns, no authority could simply collect users’ private location traces for mining or even publishing. To echo such concerns, local differential privacy (LDP) enables individual privacy by allowing each user to report a perturbed version of their data. Unfortunately, when applied to location traces, LDP cannot preserve the semantics in the context of location traces because it treats all locations (i.e., various points of interest) as equally sensitive. This results in a low utility of LDP mechanisms for collecting location traces. In this paper, we address the challenge of collecting and sharing location traces with valuable semantics while providing sufficient privacy protection for participating users. We first propose semantic-constrained local differential privacy (SLDP), a new privacy model to provide a provable mathematical privacy guarantee while preserving desirable semantics. Then, we design a location trace perturbation mechanism (LTPM) that users can use to perturb their traces in a way that satisfies SLDP. Finally, we propose a private location trace synthesis (PLTS) framework in which users use LTPM to perturb their traces before sending them to the collector, who aggregates the users’ perturbed data to generate location traces with valuable semantics. Extensive experiments on three real-world datasets demonstrate that our PLTS outperforms existing state-of-the-art methods by at least 21% in a range of real-world applications, such as spatial visiting queries and frequent pattern mining, under the same privacy leakage. Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Jiawei Duan, Qiao Xue, Tianyu Wo, Weizhe Zhang, Jie Xu 0007 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Heavy Hitter Identification Over Large-Domain Set-Valued Data With Local Differential PrivacyabstractSet-valued data are widely used to represent information in the real word, such as individual daily behaviors, items in shopping carts and web browsing history. By collecting set-valued data and identifying heavy hitters, service providers (i.e., the collector) can learn usage preferences of costumers (i.e., users), and improve the quality of their services by the learned information. However, the collection of raw data would bring privacy risks to users. Recently, local differential privacy (LDP) has emerged as a rigorous privacy framework for user private data collection. At the same time, many LDP schemes have been designed to achieve heavy hitters, but most of them are limited by the large data domain due to the huge computation cost. In this paper, we propose an LDP framework: PemSet, to efficiently identify heavy hitters from set-valued data with a large domain. In PemSet, users mainly focus on the prefix of each item (i.e., the first few bits of the binary expression of each item), and only perturb and report prefixes to reduce computation cost. Sometimes the prefixes of different items are the same, so the reported set-valued data could be a multiset, i.e., a set including multiple same items. As such, we design four LDP protocols MOLH, MOLH-S, MPCKV, MWheel to estimate frequencies of items in the multiset setting, and compare their performance under PemSet framework by experiments. Experimental results demonstrate that MOLH can perform the best in a high privacy region, i.e.,$\epsilon < 1$, while MWheel can obtain the highest utility when privacy budget is large, i.e.,$\epsilon \geqslant 1$. Youwen Zhu, Yiran Cao, Qiao Xue, Qihui Wu 0001, Yushu Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | DeepMark: A Scalable and Robust Framework for DeepFake Video DetectionabstractWith the rapid growth of DeepFake video techniques, it becomes increasingly challenging to identify them visually, posing a huge threat to our society. Unfortunately, existing detection schemes are limited to exploiting the artifacts left by DeepFake manipulations, so they struggle to keep pace with the ever-improving DeepFake models. In this work, we propose DeepMark, a scalable and robust framework for detecting DeepFakes. It imprints essential visual features of a video into DeepMark Meta (DMM) and uses it to detect DeepFake manipulations by comparing the extracted visual features with the ground truth in DMM. Therefore, DeepMark is future-proof, because a DeepFake video must aim to alter some visual feature, no matter how “natural” it looks. Furthermore, DMM also contains a signature for verifying the integrity of the above features. And an essential link to the features as well as their signature is attached with error correction codes and embedded in the video watermark. To improve the efficiency of DMM creation, we also present a threshold-based feature selection scheme and a deduced face detection scheme. Experimental results demonstrate the effectiveness and efficiency of DeepMark on DeepFake video detection under various datasets and parameter settings. Qingqing Ye 0001, Haibo Hu 0001, Qiao Xue, Yaxin Xiao, Jin Li 0002 |
ACM Trans. Priv. Secur. | 4 |
| 2024 | PUTS: Privacy-Preserving and Utility-Enhancing Framework for Trajectory SynthesizationabstractVehicle trajectory data is essential for traffic management and location-based services. However, publishing real-life trajectory data has been challenging because vehicle trajectories contain users’ sensitive information. Differential privacy addresses such problems by publishing a synthetic version of the input dataset, but existing works always assume the real-world data is absolutely accurate. This assumption no longer holds in trajectory data because it typically contains errors due to inaccurate positioning services, which leads to poor performance of data synthesized by such trajectories. Even worse, existing works may generate unrealistic trajectories due to their coarse data synthesis methods, resulting in low practical utility or even inability to handle complex tasks. In this paper, we propose aPrivacy-preserving andUtility-enhancing framework forTrajectorySynthesization (PUTS). Our framework mitigates the impact of data errors in trajectories on differential privacy mechanisms, by exploiting map-matching techniques and real-world road network structure. InPUTS, a two-layer approach from path to trajectory synthesis is proposed to not only guarantee the reality of synthetic trajectories, but also scale upPUTSin real-world applications. Extensive experiments on real-world datasets show thatPUTSsignificantly outperforms existing methods in terms of utility in a range of real-world applications. Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Jiawei Duan, Qiao Xue, Tianyu Wo, Jie Xu 0007 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | Stateful Switch: Optimized Time Series Release with Local Differential PrivacyabstractTime series data have numerous applications in big data analytics. However, they often cause privacy issues when collected from individuals. To address this problem, most existing works perturb the values in the time series while retaining their temporal order, which may lead to significant distortion of the values. Recently, we propose TLDP model [45] that perturbs temporal perturbation to ensure privacy guarantee while retaining original values. It has shown great promise to achieve significantly higher utility than value perturbation mechanisms in many time series analysis. However, its practicability is still undermined by two factors, namely, utility cost of extra missing or empty values, and inflexibility of privacy budget settings. To address them, in this paper we propose switch as a new two-way operation for temporal perturbation, as opposed to the one-way dispatch operation in [45]. The former inherently eliminates the cost of missing, empty or repeated values. Optimizing switch operation in a stateful manner, we then propose StaSwitch mechanism for time series release under TLDP. Through both analytical and empirical studies, we show that StaSwitch has significantly higher utility for the published time series than any state-of-the-art temporal- or value-perturbation mechanism, while allowing any combination of privacy budget settings. Qingqing Ye 0001, Haibo Hu 0001, Kai Huang 0011, Man Ho Au, Qiao Xue |
INFOCOM | 5 |
| 2023 | DDRM: A Continual Frequency Estimation Mechanism With Local Differential PrivacyabstractMany applications rely on continual data collection to provide real-time information services, e.g., real-time road traffic forecasts. However, the collection of original data brings risks to user privacy. Recently, local differential privacy (LDP) has emerged as a private data collection framework for mass population. However, for continual data collection, existing LDP schemes, e.g., those employing the memoization technique, are known to have privacy leakage on data change points over time. In this paper, we propose a new scheme with stronger privacy guarantee for continual frequency estimation under LDP, namely, Dynamic Difference Report Mechanism (DDRM). In DDRM, we introduce difference trees to capture the data changes over time, which well addresses possible privacy leakage on data change points. As for the utility enhancement, DDRM exploits the common case of no data change in time series and thereby suppresses the consumption of privacy budget in such cases. Meanwhile, an optimal privacy budget allocation scheme is proposed to encourage users to report more data for better estimation accuracy. By both theoretical analysis and experimental evaluations, we show DDRM achieves highly accurate frequency estimation in real time. Qiao Xue, Qingqing Ye 0001, Haibo Hu 0001, Youwen Zhu, Jian Wang 0038 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Mean estimation over numeric data with personalized local differential privacy
Qiao Xue, Youwen Zhu, Jian Wang 0038 |
Frontiers Comput. Sci. | 1 |
| 2021 | Locally differentially private distributed algorithms for set intersection and union
Qiao Xue, Youwen Zhu, Jian Wang 0038, Xingxin Li, Ji Zhang 0001 |
Sci. China Inf. Sci. | 1 |
| 2019 | M2MHub: A Blockchain-Based Approach for Tracking M2M Message ProvenanceabstractThe Internet of Things (IoT) is a fast growing and popular subset of technology. One of the main features of IoT is device autonomy, the ability for the machines embedded in the devices to function without human intervention. This includes communications with other devices through Machine-to-Machine (M2M) communication. Unfortunately, M2M communications are only stored with what behaviours they have observed, without the causal relationship as to how or why they observed those behaviours. Since M2M messages can trigger more M2M messages, the provenance of issues inside an IoT system can be hidden behind a long chain of messages, so finding the root source of any problem, such as malicious or defective devices, is almost impossible to detect. To solve this problem, in this paper we introduce M2MHub, a centralized auditing system which collects M2M messages in an IoT system and stores them in a blockchain. Devices in the system can tell the hub if they wish to open, continue, or end transactions, allowing the hub to keep track of who is the provenance of the transaction and how their transaction affects other devices. A proof-of-concept simulation has been constructed, demonstrating how M2MHub may function in a real-world implementation. The current implementation is not scalable enough to be deployed to actual IoT networks, so several ideas for future work are offered. Darren Saguil, Qiao Xue, Qusay H. Mahmoud |
AICCSA | 2 |
| 2017 | Distributed Set Intersection and Union with Local Differential PrivacyabstractPrivacy-preserving distributed set intersection and union have been widely applied in many scenarios and lots of work has paid attention to the problem. Existing solutions to privacy-preserving set intersection and union are built on secure multiparty computation protocols, which can theoretically solve it, but result in heavy computation and communication overhead. Worse still, most of the existing schemes cannot work once some participant fails. In this paper, we propose two differentially private approaches for distributed set intersection and union, respectively. In our schemes, each data contributor possesses a secret data set and perturbs it by randomized response technique to satisfy local differential privacy. Then the collector gathers all contributors' perturbed data sets and utilizes maximum likelihood estimation to gain an accurate estimation of intersection and union. Compared to existing schemes, the proposed schemes can dramatically reduce computation and communication overhead, and tolerate participant's failure. We formally prove that the proposed schemes satisfy local differential privacy, and leverage extensive experiments to evaluate the proposed approaches. The results indicate that our schemes have low computation and communication complexity, strong robustness and good utility. Qiao Xue, Youwen Zhu, Jian Wang 0038, Xingxin Li |
ICPADS | 1 |