EDBT 2026 Demo / reviewers in the wild / expert
Pengfei Zhang 0010
dblp:58/4525-10
· DBLP profile ↗
24ranked-venue papers
11as first author
24since 2021 · last 2026
0000-0003-0663-332XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient, Secure, Differentially Private Deep Learning in the Two-Server ModelabstractExisting solutions on differentially private deep learning (DPDL) either require the assumption of a trusted data server (centralized DPDL) or suffer from poor utility (local DPDL); and hence their adoptions are hampered in real-world scenarios.We present CRYPTDP, a crypto-assisted differentially private deep learning approach in the two-server model. CRYPTDP employs two non-colluding servers to collaboratively and efficiently train differentially private deep learning over the secret shares of data owners' private data while protecting the confidentiality of the data from untrusted servers. CRYPTDP is the first approach with the best of both local DPDL and centralized DPDL models, which does not resort to trusted server like local DPDL and has the utility like centralized DPDL. In particular, we also make innovations for addressing the major challenges like poor performance and security that beset CRYPTDP: We introduce a new secure computation and differential privacy friendly activation function; we propose a novel garbled-circuits-free most significant bit extraction protocol, and using the protocol we propose an efficient and secure garbled-circuits-free protocol for activation function over secret shares. Exhaustive experiments show that CRYPTDP delivers significantly better performance than the state-of-the-art local DPDL, yields higher accuracy than the state-of-the-art centralized DPDL, and can achieve two orders of magnitude faster runtime than the state-of-the-art approach. Jun Feng 0007, Pengfei Zhang 0010, Bocheng Ren, Shunli Zhang 0003 |
AAAI | 3 |
| 2026 | Stabilizing Cross-Modal Bidirectional Attribution: Few-Shot Adversarial Prompt Tuning for Robust Vision-Language ModelsabstractLarge-scale pre-trained vision-language models (VLMs) like CLIP show exceptional performance and zero-shot generalization. However, their reliability may be severely undermined by a critical vulnerability to subtle adversarial perturbations. Our work reveals a critical cross-modal vulnerability: visual-only perturbations induce substantial, synchronous shifts in decision attribution maps across both image and text. This phenomenon signifies a fundamental disruption of the VLM's internal logic, as it alters both the model's perceptual focus and its decision rationale. To counter this vulnerability, we introduce Cross-modal Bidirectional Attribution guided Few-shot Adversarial Prompt Tuning (CBA-FAPT), a novel method that leverages the model's internal decision rationale as a regularizer for robust learning. Our framework's core mechanism is the alignment of a novel bidirectional attribution map. This map is a unique fusion of two components. It combines forward feature attention to capture the model's perceptual focus. It also incorporates backward decision gradients to act as a proxy for the model's decision rationale, quantifying how each feature influences the final outcome. We enforce consistency on this bidirectional map between clean and adversarial examples. This approach corrects the model's internal logic on two fronts and effectively restores its adversarial robustness. Comprehensive experiments on 11 datasets demonstrate that CBA-FAPT outperforms the state-of-the-art, establishing a superior trade-off between robust and natural accuracy. Jun Feng 0007, Shuhong Wu, Pengfei Zhang 0010, Bocheng Ren, Shunli Zhang 0003 |
AAAI | 4 |
| 2026 | PrivSV: Differentially Private Steering Vector for Large Language ModelsabstractSteering Vector (SV) is a powerful technique for controlling Large Language Models (LLMs) by manipulating their activations without altering model weights. However, when constructed from sensitive data, SV poses significant privacy risks, as it may leak private information. Existing differential privacy (DP) techniques for constructing SV cannot be directly applied to training-based SV construction paradigms, which offer higher task performance. In this work, we present **PrivSV**, a general privacy-preserving approach for constructing SV with DP guarantees, compatible with arbitrary SV construction paradigms while maintaining high utility. In PrivSV, we propose three novel methods: a Layer-wise Noise-Resilient Reduction (LNR²) method to reduce the injected noise in high-dimensional SV; a Directional Prior Compensation (DPC) method to recover utility degraded by noise perturbation; and a Privacy-Aware Optimal Parameter Determination (POPD) method to adaptively maximize the performance of the final compensated SV. Extensive experiments on open-source LLMs of different families (i.e., LlaMa, Qwen, Mistral and Gemma) demonstrate that PrivSV outperforms several existing techniques across various privacy budgets. Xiang Cheng 0003, Chenhao Sun, Pengfei Zhang 0010, Sen Su |
AAAI | 4 |
| 2026 | Adaptive Graph Attention Based Discrete Hashing for Incomplete Cross-modal RetrievalabstractCross-modal hashing has emerged as a pivotal solution for efficient retrieval across diverse modalities, such as images and texts, by mapping them into compact binary hash spaces. However, in real-world scenarios, the modalities data is often missing or misaligned. Existing methods are most rely on fully paired training data and ignore missing or misaligned modalities data, resulting in the semantic inconsistencies. To address these challenges, we propose an Adaptive Graph Attention-Based Discrete Hashing (AGADH) method, which consists of three parts. First, to solve the problem of missing modalities, AGADH employs a masked completion strategy to reconstruct missing modalities. Second, to mitigate semantic misalignment, AGADH leverages a Graph Attention Network (GAT) encoder-decoder architecture with alignment module to construct features from different modalities. Additionally, to enhance the fusion performance, an adaptive fusion module dynamically adjusting the contributions of image and text modalities with learnable weighting coefficients is proposed. Extensive experiments on three benchmark datasets, MS-COCO, NUS-WIDE, and MIRFlickr-25K, demonstrating that AGADH outperforms state-of-the-art methods in both fully paired and incompletely paired scenarios, showing its robustness and effectiveness in cross-modal retrieval tasks. Shuang Zhang 0009, Lei Shi 0030, Huilong Jin, Feifei Kou, Pengfei Zhang 0010, Mingying Xu, Pengtao Lv |
AAAI | 6 |
| 2026 | MusicRec: Multi-modal Semantic-Enhanced Identifier with Collaborative Signals for Generative RecommendationabstractGenerative recommendation as a new paradigm is influencing the current development of recommender systems. It aims to assign identifiers that capture richer semantic and collaborative information to items, and subsequently predict item identifiers via autoregressive generation using Large Language Models (LLMs). Existing approaches primarily tokenize item text into codebooks with preserved semantic IDs through RQ-VAE, or separately tokenize different modality features of items. However, existing tokenization methods face two major challenges: (1) Learning decoupled multi-modal features limits the quality of the semantic representation. (2) Ignoring collaborative signals from interaction history limits the comprehensiveness of identifiers. To address these limitations, we propose a multi-modal semantic-enhanced identifier with collaborative signals for generative recommendation, named MusicRec. In MusicRec, we propose a tokenization approach based on shared-specific modal fusion, enabling the generated identifiers to preserve semantic information more comprehensively from all modalities. In addition, we incorporate collaborative signals from user interactions to guide identifier generation, preserving collaborative patterns in the semantic representation space. Extensive experiments on three public datasets demonstrate that MusicRec achieves state-of-the-art performance compared to existing baseline methods. Yuqiu Zhao, Lei Shi 0030, Yan Zhong 0001, Feifei Kou, Pengfei Zhang 0010, Jiwei Zhang 0007, Mingying Xu |
AAAI | 5 |
| 2026 | Multimodal sentiment analysis with temporal semantics self-supervised multi-task learning and single-modality label generation
Ruixin Pu, Pengfei Zhang 0010, Shudong Zhang |
Expert Syst. Appl. | 4 |
| 2026 | Dual Graph Network Hashing for Cross-Modal Retrieval
Shuang Zhang 0009, Lei Shi 0030, Feifei Kou, Huilong Jin, Pengfei Zhang 0010, Weiping Ding 0001, Mingying Xu, Muhammet Deveci |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2026 | Horizontal Multi-Party Data Publishing Under Differential Privacy via Weight-Aware Bidirectional Generative Adversarial Networks
Pengfei Zhang 0010, Zhikun Zhang 0001, Yang Cao 0011, Xiang Cheng 0003, Lihua Yin, Puning Zhao, Zhiquan Liu 0001, Li Sun 0008, Lei Shi 0030, Ji Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2026 | Locally Differentially Private Truth Discovery for Sparse CrowdsensingabstractTruth discovery has emerged as an effective tool to mitigate data inconsistency in crowdsensing by prioritizing data from high-quality responders. While local differential privacy (LDP) has emerged as a crucial privacy-preserving paradigm, existing studies under LDP rarely explore a worker's participation in specific tasks for sparse scenarios, which may also reveal sensitive information such as individual preferences and behaviors. Existing LDP mechanisms, when applied to truth discovery in sparse settings, may create undesirable dense distributions, provide insufficient privacy protection, and introduce excessive noise, compromising the efficacy of subsequent non-private truth discovery. Additionally, the interplay between noise injection and truth discovery remains insufficiently explored in the current literature. To address these issues, we propose a lOcally differentially private truth diSCovery approach for spArse cRowdsensing, namely OSCAR. The main idea is to use advanced optimization techniques to reconstruct the sparse data distribution and re-formalize truth discovery by considering the statistical characteristics of injected Laplacian noise while protecting the privacy of both the tasks being completed and the corresponding sensory data. Specifically, to address the data density concerns while alleviating noise, we design a randomized response based Bernoulli matrix factorization method BerRR. To recover the sparse structures from densified, perturbed data, we formalize a 0-1 integer programming problem and develop a sparse recovery solving method SpaIE based on implicit enumeration. We further devise a Laplacian-sensitive truth discovery method LapCRH that leverages maximum likelihood estimation to re-formalize truth discovery by measuring differences between noisy values and truths based on the statistical characteristic of Laplacian noise. Our comprehensive theoretical analysis establishes OSCAR's privacy guarantees, utility bounds, and computational complexity. Experimental results show that OSCAR surpasses the state-of-the-arts by at least 30% in accuracy improvement. Pengfei Zhang 0010, Zhikun Zhang 0001, Yang Cao 0011, Xiang Cheng 0003, Youwen Zhu, Zhiquan Liu 0001, Ji Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2026 | Locally Differentially Private Truth Discovery Over Data StreamsabstractData inconsistency often arises from multiple observed sensory data due to varying participant reliability for crowdsensing systems. Truth discovery, which includesWeight EstimationandTruth Aggregation, for estimating participant reliability weights and aggregating uploaded values from inconsistent observations respectively, has emerged as an effective solution to address this issue. While local differential privacy (LDP) provides strong privacy guarantees by allowing participants to perturb their data locally before submission, existing LDP-based studies are either designed for static scenarios or compromise on privacy and accuracy trade-off for data streams, satisfying only weaker versions of LDP or mere differential privacy. To effectively and efficiently obtain truths over streams under rigorous LDP, we proposeNANOwhich is locally differeNtially privAte truth discovery via updatiNg time stamp determinatiOn. The main idea lies in its integration of Laplacian noise for privacy protection and inherent Gaussian noise representing natural data variability for effective weight and truth estimations, coupled with the adaptive determination of updating time stamps. InNANO, to obtain theWeight EstimationandTruth Aggregationunder LDP, we design a mixed noise-aware truth discovery methodMixTDby modeling the mixed noise. To capture the dynamic nature of weight and truth evolutions, we develop a changing-aware updating time stamp determination methodCUDto selectively re-conduct truth discovery at specific time stamps. We also introduce a dynamic privacy budget management strategy, which accumulates unused budgets from skipped updates for critical timestamps. In this way,Weight EstimationandTruth Aggregationare limited to critical time stamps, which significantly reduces the privacy budget segmentation and computational costs. We demonstrate thatNANOprovides rigorous LDP guarantees while achieving bounded utility and computational complexity. Extensive experimental results over four real-world datasets and three synthetic datasets showcase thatNANOoutperforms the state-of-the-arts by at least 20% improvement with negligible extra efficiency loss. Pengfei Zhang 0010, Zhikun Zhang 0001, Yang Cao 0011, Shaowei Wang 0003, Xiang Cheng 0003, Ji Zhang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | PRIMA: Privacy preserving Multi-dimensional Analytic ApproachabstractSum query is an important and fundamental operator for online analytical processing. In this paper, we focus on the process of answering sum queries over data cube, each of which consists of a collection of cuboids, while satisfying differential privacy (DP). Existing works fail to process the sum queries in online analytical processing with high utility due to sum queries' high sensitivity and the noise aggregation: constructing a base cuboid requires the data curator to answer a workload of linear sum queries under DP in advance, whose sensitivity will result in a large amount of DP noise, and the noise will finally be aggregated when constructing the remaining cuboids. To this end, we present a Differentially PRIvate Multi-dimensional Analytic Approach (PRIMA). In PRIMA, we propose a Symmetric Bounded Sum Query Processing Method (SBS) which reduces the sensitivity of sum queries by bounding both the maximum and minimum contribution of each record in the data table in a symmetricaly manner. Moreover, we propose a Hypothesis Testing based Prefix Sum Computing Method (SCOPE) to compute a base prefix-sum cuboid based on hypothesis testing. By employing the base prefix-sum cuboid, any remaining cuboid can be constructed with constant pieces of DP noise aggregated. We conduct experiments on both real-world and synthetic datasets. Experimental results confirm the effectiveness of PRIMA over existing works. Xiang Cheng 0003, Pengfei Zhang 0010, Anxing Wei |
CIKM | 3 |
| 2025 | Hyperbolic Prompt Learning for Incremental Event Detection with LLMsabstractClass-incremental event detection (CIED) is essential for real-world information extraction systems, which must continually recognize new event types without forgetting past knowledge. The main challenge lies in balancing stability and adaptability under data imbalance. Existing methods often underuse the hierarchical and syntactic structures of language, and thus limit the generalization capacity. We propose HPLLM, a hyperbolic prompt-enhanced large language model framework, motivated by the observation that both embedding distributions and dependency graphs in event datasets exhibit hyperbolic properties. HPLLM integrates two key components: (1) Hyperbolic LoRA fine-tuning, enabling geometry-aware parameter adaptation for hierarchical semantics; and (2) Hyperbolic Adaptive Graph Diffusion Convolution (HADC), which encodes syntactic dependencies into structure-aware prompts for LLMs. Together, these techniques strengthen semantic discrimination, reduce forgetting, and improve adaptation across incremental stages. Extensive experiments on ACE2005 and MAVEN demonstrate that HPLLM consistently surpasses state-of-the-art baselines in macro-F1, achieving stronger retention of old knowledge and better generalization to new event types. In particular, the model shows clear gains on rare categories with few training mentions, demonstrating its robustness in imbalanced and few-shot regimes. Xiujin Zhang, Wenxin Jin, Haotian Hong, Pengfei Zhang 0010, Jiting Li, Kongjing Gu, Hao Peng 0001, Li Sun 0008 |
CIKM | 4 |
| 2025 | VRPTD: Verifiable and Robust Privacy-Preserving Truth Discovery for IoT CrowdsensingabstractThe rapid development of Internet of Things (IoT) devices has greatly boosted large-scale crowdsensing, where truth discovery (TD) technique is a feasible solution to extract more reliable data from heterogeneous measurements or noisy environment. However, most existing approaches send raw readings to a centralized server, implicitly fully trusting the server and fog nodes. In practice, these nodes are geographically dispersed, outsourced, and operated by different parties, making such blanket trust untenable. To address these threats, a verifiable, robust, and privacy-preserving truth discovery (VRPTD) framework is proposed. Specifically, VRPTD shields individual reports through a lightweight dual-layer mechanism that combines unbiased expected mask correction with modular arithmetic encryption. Data requesters can publicly audit the result via linear homomorphic hash proof. VRPTD redesigns the weight update step using a Gaussian radial basis function distance, which yields more accurate processing of multidimensional measurements. Security analysis indicates that VRPTD satisfies verifiability under a well-defined threat model. Extensive experiments on real-world IoT datasets demonstrate that VRPTD boosts baseline TD accuracy, halves convergence rounds and reduces communication and computation overhead, all while providing strong privacy and resilience against active attacks. Pengfei Zhang 0010, Zhao Li 0007, Ji Zhang 0001 |
ICPADS | 3 |
| 2025 | EVICheck: Evidence-Driven Independent Reasoning and Combined Verification Method for Fact-CheckingabstractLarge Language Models (LLMs) and Retrieval-Augmented Generation (RAG) have demonstrated significant potential in automated fact-checking. However, existing methods face limitations in insufficient evidence utilization and lack of explicit verification criteria. Specifically, these approaches aggregate evidence for collective reasoning without independently analyzing each piece, hindering their ability to leverage the available information thoroughly. Additionally, they rely on simple prompts or few-shot learning for verification, which makes truthfulness judgments less reliable, especially for complex claims. To address these limitations, we propose a novel method to enhance evidence utilization and introduce explicit verification criteria, named EVICheck. Our approach independently reasons each evidence piece and synthesizes the results to enable more thorough exploration and enhance interpretability. Additionally, by incorporating fine-grained truthfulness criteria, we make the model's verification process more structured and reliable, especially when handling complex claims. Experimental results on the public RAWFC dataset demonstrate that EVICheck achieves state-of-the-art performance across all evaluation metrics. Our method demonstrates strong potential in fake news verification, significantly improving the accuracy. Lei Shi 0030, Feifei Kou, Ligu Zhu, Chen Ma 0003, Pengfei Zhang 0010, Mingying Xu |
IJCAI | 6 |
| 2025 | Dynamic Masking and Auxiliary Hash Learning for Enhanced Cross-Modal RetrievalabstractThe demand for multimodal data processing drives the development of information technology. Cross-modal hash retrieval has attracted much attention because it can overcome modal differences and achieve efficient retrieval, and has shown great application potential in many practical scenarios. Existing cross-modal hashing methods have difficulties in fully capturing the semantic information of different modal data, which leads to a significant semantic gap between modalities. Moreover, these methods often ignore the importance differences of channels, and due to the limitation of a single goal, the matching effect between hash codes is also affected to a certain extent, thus facing many challenges. To address these issues, we propose a Dynamic Masking and Auxiliary Hash Learning (AHLR) method for enhanced cross-modal retrieval. By jointly leveraging the dynamic masking and auxiliary hash learning mechanisms, our approach effectively resolves the problems of channel information imbalance and insufficient key information capture, thereby significantly improving the retrieval accuracy. Specifically, we introduce a dynamic masking mechanism that automatically screens and weights the key information in images and texts during the training process, enhancing the accuracy of feature matching. We further construct an auxiliary hash layer to adaptively balance the weights of features across each channel, compensating for the deficiencies of traditional methods in key information capture and channel processing. In addition, we design a contrastive loss function to optimize the generation of hash codes and enhance their discriminative power, further improving the performance of cross-modal retrieval. Comprehensive experimental results on NUS-WIDE, MIRFlickr-25K and MS-COCO benchmark datasets show that the proposed AHLR algorithm outperforms several existing algorithms. Shuang Zhang 0009, Lei Shi 0030, Feifei Kou, Huilong Jin, Pengfei Zhang 0010, Meiyu Liang, Mingying Xu |
NeurIPS | 7 |
| 2025 | Maximizing Area Coverage in Privacy-Preserving Worker Recruitment: A Prior Knowledge-Enhanced Geo-Indistinguishable ApproachabstractWorker recruitment for area coverage maximization, typically requires participants to upload location information, which can deter potential participation without proper protection. While existing studies resort to geo-indistinguishability to address this concern, they primarily focus on either specific task locations (Target Coverage) or operate under pre-defined recruitment quotas for an interested region (Area Coverage). These focuses not only yield suboptimal area coverage when scaled but also fail to leverage valuable prior knowledge in the form of participants’ noisy historical registered locations, to enhance both location obfuscation and worker identification processes. To address these limitations, we presentWILTON, which optimizes area coverage under geo-indistinguishability by recruiting the minimum number of participants through the strategic utilization of noisy prior knowledge. InWILTON, to generate obfuscated locations, we propose a probabilistic and weight-aware input perturbation mechanism, which groups and weights prior locations rather than using only personal prior locations. To privately identify the recruited workers, we design a grid-based worker identification method, which provides a worst-case performance guarantee of ratio 1 − 1/eto the optimum. We provide a theoretical analysis of the privacy, utility, and complexity guarantees ofWILTON. Experimental results over two real-world datasets and one synthetic dataset show thatWILTONsurpasses the state-of-the-arts by at least 8% in area coverage improvement. Pengfei Zhang 0010, Xiang Cheng 0003, Zhikun Zhang 0001, Youwen Zhu, Ji Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Privacy-Preserving Recommendations With Mixture Model-Based Matrix Factorization Under Local Differential PrivacyabstractMatrix factorization-based recommendations have emerged as a cornerstone for many industrial recommendation algorithms. However, studies employing differential privacy often rely on a trusted server, while those implementing local differential privacy (LDP) frequently encounter substantial accuracy degradation and excessive communication overhead due to noise injection and frequent user interactions. Moreover, the presence of injected noise for privacy protection and inherent Gaussian noise within these perturbed values compounds the accuracy issues and may create a cascading effect under LDP, exacerbating the complexity of the problem at hand. To address these challenges, we proposeMENTOR, which isMixture model-based rEcommeNdations approach with matrix FacTORization under LDP. Its main idea lies in the adoption of a bounded input perturbation mechanism that closely approximates the Laplace distribution while providing rigorous LDP, coupled with accounting for various noise types present in the perturbed data. This approach allows us to reformulate the matrix factorization process with minimal user interaction, requiring only a single round of communication. In particular, to add noise, we design a bidirectional bounded LDP input perturbation mechanismBBVwhile minimizing variance. To generate recommendation results, we devise a matrix factorization techniqueGLMFbased on a Gaussian–Laplacian mixture model. Comprehensive experiments reveal thatMENTORoutperforms the state of the art by at least 15% inRMSEand 12% inF-Measure, showcasing its effectiveness in balancing privacy and utility in recommendation systems. Pengfei Zhang 0010, Zhikun Zhang 0001, Xiang Cheng 0003, Youwen Zhu, Ji Zhang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Numerical Data Collection Under Input-Discriminative Local Differential PrivacyabstractInput-discriminative local differential privacy (ID-LDP) protects user data with a different range of values, which improves the utility of the estimated data compared to traditional LDP. However, the existing ID-LDP methods are used for categorical data and cannot be directly applied to numerical data. In this paper, we propose a numerical data collection (NDC) framework with ID-LDP to provide discriminative protection for the data with different inputs. This framework uses a piecewise mechanism to divide the numerical data into several segments and designs two perturbation methods to minimize the mean value of numerical data based on values submitted by users. We first create an NDC-UE method that encodes the raw data into a binary vector. This method sets the uploaded data bit as 1 and the rest as zero and perturbs each bit with a given probability. We further propose an NDC-GRR algorithm to perturb the numerical data with an optimal privacy budget. To reduce the complexity of NDC-GRR, we apply a greedy algorithm-based spanner to shorten the computation time and improve the accuracy. Theoretical analysis proves that our schemes satisfy the definition of ID-LDP. Experimental results based on two real-world datasets and a synthetic dataset show that the proposed schemes have less mean square error compared with the benchmarks. Youwen Zhu, Shibo Dai, Pengfei Zhang 0010, Xiqi Kuang |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | A Gaussian Distribution-Based Truth Discovery Algorithm under Local Differential PrivacyabstractTruth discovery is an effective tool for discovering the truth from a multitude of data points of varying quality, which inherently involves privacy concerns. While existing studies have predominantly focused on protecting workers’ submitted sensing data using local differential privacy (LDP), they overlook a crucial aspect of real-world scenarios: workers are likely to provide more accurate data for tasks they perceive as important, resulting in submissions that more closely approximate the truth for these tasks. Moreover, the prevalent use of the Laplace mechanism for noise addition, due to the inherent randomness and unboundedness of the Laplace distribution, might lead to excessive noise, potentially compromising the accuracy of truth discovery and yielding a noisy approximation of the truth. To address these limitations, we propose a Gaussian distribution-based truth discOvery approach under Local Differential privacy (GOLD). The algorithm’s core innovation lies in its comprehensive utilization of Gaussian distribution for both task importance and worker quality after adding Laplacian noise. Workers first apply Laplacian noise to their data locally, after which the problem is formalized as a constraint optimization task, deriving an iterative equation for the noisy truth. Theoretical analysis demonstrates that the GOLD algorithm rigorously adheres to local differential privacy requirements while achieving high truth accuracy and low time complexity. Empirical validation on two real datasets reveals that, compared to state-of-the-art algorithms, the GOLD algorithm improves the truth accuracy by at least 20%. Pengfei Zhang 0010, Ximeng Liu, Bin Wu 0019, Li Sun 0008, Shoufei Han, Xianjin Fang, Ji Zhang 0001 |
HPCC | 1 |
| 2024 | Task Allocation with Profit Maximization Under Geo-indistinguishability via Q-learningabstractTask allocation, a core component of mobile crowd-sensing systems, facilitates the collection, analysis, and sharing of diverse data. While existing studies often employ planar Laplacian (PL) distribution to achieve Geo-indistinguishability (Geo-I) for worker location protection, the randomness and boundlessness of PL distribution, coupled with greedy allocation strategies, often lead to excessive noise and incomplete task assignments. Moreover, these approaches typically overlook the equilibrium between worker and server benefits. To address these challenges under Geo-I, we present the Kitty approach, which adopts Q-learning to achieve superior task allocation after formalizing a constrained optimization problem that maximizes profits for both parties. Kitty operates through three key mechanisms: 1) formalizing a constrained optimization problem based on a comprehensive analysis of both parties’ profits and a pre-defined equilibrium parameter, 2) implementing adaptive adjustment of the Q-learning greedy parameter to balance exploration and exploitation, and 3) designing two conflict resolution strategies to mitigate potential distance conflicts after location perturbation. Experiments on two real-world datasets demonstrate that Kitty outperforms the state-of-the-art by at least 15% in average travel distance reduction and 1% in task completion rate improvement. Pengfei Zhang 0010, Ximeng Liu, Bin Wu 0019, Li Sun 0008, Shoufei Han, Xianjin Fang, Ji Zhang 0001 |
HPCC | 1 |
| 2023 | Effective truth discovery under local differential privacy by leveraging noise-aware probabilistic estimation and fusion
Pengfei Zhang 0010, Xiang Cheng 0003, Sen Su |
Knowl. Based Syst. | 1 |
| 2023 | Task Allocation Under Geo-Indistinguishability via Group-Based Noise AdditionabstractLocations are usually necessary for task allocation in spatial crowdsourcing, which may put individual privacy in jeopardy without proper protection. Although existing studies have well explored the problem of location privacy protection in task allocation under geo-indistinguishability, they potentially assume the workers could perform any tasks, which might not be practical in reality. Moreover, they usually adopt planar laplacian mechanism to achieve geo-indistinguishability, which will introduce excessive noise due to its randomness and boundlessness. To this end, we propose a task alloCAtioNapproach via grOup-based noisEaddition under Geo-I, referred to asCANOE. Its main idea is that each worker uploads the noisy distances between his true location and the obfuscated locations of his preferred tasks instead of uploading his obfuscated location. In particular, to alleviate the total noise when conducting grouping, we put forward an optimized global grouping with adaptive local adjustment methodOGALwith convergence guarantee. To collect the noisy distances which are required for subsequent task allocation, we develop a utility-aware obfuscated distance collection methodUODCwith solid privacy and utility guarantees. We further theoretically analyze the privacy, utility and complexity guarantees ofCANOE. Extensive analyses and experiments over two real-world datasets confirm the effectiveness ofCANOE. Pengfei Zhang 0010, Xiang Cheng 0003, Sen Su |
IEEE Trans. Big Data | 1 |
| 2023 | PrivTDSI: A Local Differentially Private Approach for Truth Discovery via Sampling and InferenceabstractTruth discovery is an effective way to identify the aggregated truth of each task among multiple observed data drawn from different workers of varying reliabilities. However, existing studies are insufficient to protect individuals’ privacy, as they either just guarantee the weaker versions of local differential privacy (LDP) or potentially assume that the tasks are independent. In this paper, we, for the first time, investigate the problem of truth discovery while achieving the rigorous LDP for each worker with continuous inputs without the independence assumption. We present a locally differentially private truth discovery approach calledPrivTDSIbased on sampling and inference with solid privacy and utility guarantees. InPrivTDSI, the server first determines which values of each worker should be sampled according to a sample proportion and sends the indexes of these values to each worker. Then, each worker adds noise into the sampled values for privacy protection and uploads them to the server. After receiving the noisy sampled values from all the workers, the server first infers the unsampled values and then conducts truth discovery based on both the noisy sampled values and the inferred values. In particular, to determine the sample proportion, we formulate aconstrained nonlinear programmingproblem and give a closed-form solution to this problem. Moreover, to determine which values of each worker should be sampled while avoiding the situation where the values of some workers or tasks might not be sampled at all, we develop a two-stage sampling method calledTOSS. Furthermore, to infer the unsampled values accurately, we design a quality-aware inference method based on matrix factorization calledQualityMF. Experimental results on two real-world datasets and a synthetic dataset demonstrate the effectiveness of${PrivTDSI}$. Pengfei Zhang 0010, Xiang Cheng 0003, Sen Su, Binyuan Zhu |
IEEE Trans. Big Data | 1 |
| 2022 | Area coverage-based worker recruitment under geo-indistinguishability
Pengfei Zhang 0010, Xiang Cheng 0003, Sen Su |
Comput. Networks | 1 |