Yang Cao 0011

dblp:25/7045-11 · DBLP profile ↗
in reviewer pool ← Back
50ranked-venue papers in the field
10as first author
35since 2021 · last 2026
0000-0002-6424-8633ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 34 (10 first)Information Retrieval & Web Search · 5Big Data, Cloud & Distributed Data Systems · 5Other / Interdisciplinary · 3Data Mining & Knowledge Discovery · 2Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 Unleashing the Power of Pre-trained Graph Models in Federated Graph Learning
Huabin Sun, Bo Yan 0005, Shaohua Fan, Yang Cao 0011, Chuan Shi 0001
DASFAA (2)5
2026 Spattack: Subgroup Poisoning Attacks on Federated Recommender Systems
abstract
Federated recommender systems (FedRec) have emerged as a promising approach to provide personalized recommendations while protecting user privacy. However, recent studies have demonstrated their vulnerability to poisoning attacks, wherein malicious clients can inject carefully crafted gradients to prompt target items to benign users. Existing attacks typically target the full user group, which compromises stealth and increases the risk of detection. In contrast, real-world adversaries may prefer to target specific user subgroup, such as promoting health supplements to older individual, to maximize attack success while preserving stealth to evade detection. Motivated by this gap, we introduce Spattack, the first poisoning attack designed to manipulate recommendations for specific user subgroups in federated setting. Specifically, Spattack adopts an approximate-and-promote paradigm, which first approximate user embeddings of target/non-target subgroups and then prompts target items to the target subgroups. We further reveal a trade-off in achieving strong attack performance on the target group while keeping the non-target group largely unaffected. To achieve a better trade-off, we propose enhanced approximation and promotion strategies. For the approximation, we push the embeddings of different subgroup away based on contrastive learning and augment the target group's relevant item set via clustering. For the promotion, we align target and relevant item embeddings to strengthen their semantic connections. An adaptive weighting strategy is further proposed to balance promotion effects between target and non-target subgroups. Experiments on three real-world datasets demonstrate that Spattack consistently achieves strong attack performance on the target subgroup with minimal impact on non-target users, even when only 0.1% of users are malicious. Moreover, Spattack maintains competitive recommendation performance and exhibits strong resilience against mainstream defenses.
Bo Yan 0005, Yurong Hao, Dingqi Liu, Huabin Sun, Pengpeng Qiao, Wei Yang Bryan Lim, Yang Cao 0011, Chuan Shi 0001
WWW7
2026 Horizontal Multi-Party Data Publishing Under Differential Privacy via Weight-Aware Bidirectional Generative Adversarial Networks
Pengfei Zhang 0010, Zhikun Zhang 0001, Yang Cao 0011, Xiang Cheng 0003, Lihua Yin, Puning Zhao, Zhiquan Liu 0001, Li Sun 0008, Lei Shi 0030, Ji Zhang 0001
IEEE Trans. Knowl. Data Eng.3
2026 Locally Differentially Private Truth Discovery for Sparse Crowdsensing
abstract
Truth discovery has emerged as an effective tool to mitigate data inconsistency in crowdsensing by prioritizing data from high-quality responders. While local differential privacy (LDP) has emerged as a crucial privacy-preserving paradigm, existing studies under LDP rarely explore a worker's participation in specific tasks for sparse scenarios, which may also reveal sensitive information such as individual preferences and behaviors. Existing LDP mechanisms, when applied to truth discovery in sparse settings, may create undesirable dense distributions, provide insufficient privacy protection, and introduce excessive noise, compromising the efficacy of subsequent non-private truth discovery. Additionally, the interplay between noise injection and truth discovery remains insufficiently explored in the current literature. To address these issues, we propose a lOcally differentially private truth diSCovery approach for spArse cRowdsensing, namely OSCAR. The main idea is to use advanced optimization techniques to reconstruct the sparse data distribution and re-formalize truth discovery by considering the statistical characteristics of injected Laplacian noise while protecting the privacy of both the tasks being completed and the corresponding sensory data. Specifically, to address the data density concerns while alleviating noise, we design a randomized response based Bernoulli matrix factorization method BerRR. To recover the sparse structures from densified, perturbed data, we formalize a 0-1 integer programming problem and develop a sparse recovery solving method SpaIE based on implicit enumeration. We further devise a Laplacian-sensitive truth discovery method LapCRH that leverages maximum likelihood estimation to re-formalize truth discovery by measuring differences between noisy values and truths based on the statistical characteristic of Laplacian noise. Our comprehensive theoretical analysis establishes OSCAR's privacy guarantees, utility bounds, and computational complexity. Experimental results show that OSCAR surpasses the state-of-the-arts by at least 30% in accuracy improvement.
Pengfei Zhang 0010, Zhikun Zhang 0001, Yang Cao 0011, Xiang Cheng 0003, Youwen Zhu, Zhiquan Liu 0001, Ji Zhang 0001
IEEE Trans. Knowl. Data Eng.3
2025 Bargaining-Based Data Markets
abstract
With the prevalence of data-driven business, data markets where data can be commoditized, circulated, and ex-ploited are gaining considerable interest in the data management community. However, the uncertainty in data value poses great challenges for data pricing and thus data trading, which is magnified by the externality arising from the replicable nature of data. In this paper, we present the first bargaining-based data market framework to resolve the externality in data markets. Gearing toward raw data trading, we propose a three-stage bar-gaining model to formulate trading dynamics, which ascertains the data price agreed by both sellers and buyers. With parameters instantiated in preparation stage, an iterative bidding algorithm with provable convergence is designed in negotiation stage to solve the data pricing problem by eliciting equilibrium bids from participants with their profits optimized. Approximation algorithms with guaranteed bounds are presented in settlement stage to solve the NP-hard data allocation problem for profit maximization for the data seller with individual rationality satisfied for data buyers. Experiments on real datasets verify the effectiveness and efficiency of our framework.
Yuran Bi, Jinfei Liu, Kui Ren 0001, Yihang Wu, Yang Cao 0011
ICDE5
2025 PGB: Benchmarking Differentially Private Synthetic Graph Generation Algorithms
abstract
Differentially private graph analysis is a powerful tool for deriving insights from diverse graph data while protecting individual information. Designing private analytic algorithms for different graph queries often requires starting from scratch. In contrast, differentially private synthetic graph generation offers a general paradigm that supports one-time generation for multiple queries. Although various differentially private graph generation algorithms have been proposed, comparing them effectively remains challenging due to various factors, including differing privacy definitions, diverse graph datasets, varied privacy requirements, and multiple utility metrics. To this end, we propose PGB (Private Graph Benchmark), a comprehensive benchmark designed to enable researchers to compare differentially private graph generation algorithms fairly. We begin by identifying four essential elements of existing works as a 4-tuple: mechanisms, graph datasets, privacy requirements, and utility metrics. We discuss principles regarding these elements to ensure the comprehensiveness of a benchmark. Next, we present a benchmark instantiation that adheres to all principles, establishing a new method to evaluate existing and newly proposed graph generation algorithms. Through extensive theoretical and empirical analysis, we gain valuable insights into the strengths and weaknesses of prior algorithms. Our results indicate that there is no universal solution for all possible cases. Finally, we provide guidelines to help researchers select appropriate mechanisms for various scenarios.
Shang Liu 0001, Yang Cao 0011, Bo Yan 0005, Jinfei Liu, Masatoshi Yoshikawa
ICDE3
2025 Privacy in Fine-Tuning Large Language Models: Attacks, Defenses, and Future Directions
Shang Liu 0001, Lele Zheng, Yang Cao 0011, Atsuyoshi Nakamura
PAKDD (4)4
2025 CODE ACROSTIC: Robust Watermarking for Code Generation
Siyuan Xin, Yang Cao 0011, Xiaochun Cao
WISE (2)3
2025 Are Your LLM-based Text-to-SQL Models Secure? Exploring SQL Injection via Backdoor Attacks
abstract
Large language models (LLMs) have shown state-of-the-art results in translating natural language questions into SQL queries (Text-to-SQL), a long-standing challenge within the database community. However, security concerns remain largely unexplored, particularly the threat of backdoor attacks, which can introduce malicious behaviors into models through fine-tuning with poisoned datasets. In this work, we systematically investigate the vulnerabilities of LLM-based Text-to-SQL models and present ToxicSQL, a novel backdoor attack framework. Our approach leverages stealthy command-like and character-level triggers to make backdoors difficult to detect and remove, ensuring that malicious behaviors remain covert while maintaining high model accuracy on benign inputs. Furthermore, we propose leveraging SQL injection payloads as backdoor targets, enabling the generation of malicious yet executable SQL queries, which pose severe security and privacy risks in language model-based SQL development. We demonstrate that injecting only 0.44% of poisoned data can result in an attack success rate of 79.41%, posing a significant risk to database security. Additionally, we propose detection and mitigation strategies to enhance model reliability. Our findings highlight the urgent need for security-aware Text-to-SQL development, emphasizing the importance of robust defenses against backdoor threats.
Meiyu Lin, Jiale Lao, Renyuan Li, Yuanchun Zhou, Carl Yang 0001, Yang Cao 0011, MingJie Tang
Proc. ACM Manag. Data7
2025 Continuous Publication of Weighted Graphs with Local Differential Privacy
abstract
Although a large amount of valuable knowledge can be obtained from the weighted graph snapshots modeled over time, it may cause privacy issues. Local differential privacy (LDP) provides a strong solution for private graph data publishing in decentralized networks. However, most existing LDP studies over graphs are only applicable to static unweighted graphs. This paper investigates the problem of continuous publication of weighted graph snapshots and proposes a graph publication framework, WGT-LDP, under w -event edge weight LDP, which can protect the privacy of edges and weights over any w consecutive time steps. WGT-LDP consists of four key components: population division-based sampling that overcomes the problem of over-segmentation of the privacy budget, data range estimation that mitigates noise on edge weights, aggregate information collection that obtains important information about the graph structure and edge weights, and graph snapshot generation that reconstructs weighted graph snapshot at each time step. We provide theoretical guarantees on privacy and utility, and perform extensive experiments on three real-world and two synthetic datasets, using four commonly used metrics. Our experiments show that WGT-LDP produces high-quality synthetic weighted graphs and significantly outperforms baseline methods.
Pengpeng Qiao, Shang Liu 0001, Zhirun Zheng, Yang Cao 0011, Zhetao Li
Proc. VLDB Endow.5
2024 Enhancing Privacy of Spatiotemporal Federated Learning Against Gradient Inversion Attacks
Lele Zheng, Yang Cao 0011, Renhe Jiang, Kenjiro Taura, Yulong Shen 0001, Sheng Li 0010, Masatoshi Yoshikawa
DASFAA (1)2
2024 CARGO: Crypto-Assisted Differentially Private Triangle Counting Without Trusted Servers
abstract
Differentially private triangle counting in graphs is essential for analyzing connection patterns and calculating clustering coefficients while protecting sensitive individual information. Previous works have relied on either central or local models to enforce differential privacy. However, a significant utility gap exists between the central and local models of differentially private triangle counting, depending on whether or not a trusted server is needed. In particular, the central model provides a high accuracy but necessitates a trusted server. The local model does not require a trusted server but suffers from limited accuracy. Our paper introduces a crypto-assisted differentially private triangle counting system, named CARGO, leveraging cryptographic building blocks to improve the effectiveness of differentially private triangle counting without assumption of trusted servers. It achieves high utility similar to the central model but without the need for a trusted server like the local model. CARGO consists of three main components. First, we introduce a similarity-based projection method that reduces the global sensitivity while preserving more triangles via triangle homogeneity. Second, we present a triangle counting scheme based on the additive secret sharing that securely and accurately computes the triangles while protecting sensitive information. Third, we design a distributed perturbation algorithm that perturbs the triangle count with minimal but sufficient noise. We also provide a comprehensive theoretical and empirical analysis of our proposed methods. Extensive experiments demonstrate that our CARGO significantly outperforms the local model in terms of utility and achieves high-utility triangle counting comparable to the central model.
Shang Liu 0001, Yang Cao 0011, Takao Murakami, Jinfei Liu, Masatoshi Yoshikawa
ICDE2
2024 Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
Sheng Li 0010, Yang Cao 0011, Zhao Ren, Tanja Schultz
MMAsia3
2024 Federated Heterogeneous Graph Neural Network for Privacy-preserving Recommendation
abstract
The heterogeneous information network (HIN), which contains rich semantics depicted by meta-paths, has emerged as a potent tool for mitigating data sparsity in recommender systems. Existing HIN-based recommender systems operate under the assumption of centralized storage and model training. However, real-world data is often distributed due to privacy concerns, leading to the semantic broken issue within HINs and consequent failures in centralized HIN-based recommendations. In this paper, we suggest the HIN is partitioned into private HINs stored on the client side and shared HINs on the server. Following this setting, we propose a federated heterogeneous graph neural network (FedHGNN) based framework, which facilitates collaborative training of a recommendation model using distributed HINs while protecting user privacy. Specifically, we first formalize the privacy definition for HIN-based federated recommendation (FedRec) in the light of differential privacy, with the goal of protecting user-item interactions within private HIN as well as users' high-order patterns from shared HINs. To recover the broken meta-path based semantics and ensure proposed privacy measures, we elaborately design a semantic-preserving user interactions publishing method, which locally perturbs user's high-order patterns and related user-item interactions for publishing. Subsequently, we introduce an HGNN model for recommendation, which conducts node- and semantic-level aggregations to capture recovered semantics. Extensive experiments on four datasets demonstrate that our model outperforms existing methods by a substantial margin (up to 34% in HR@10 and 42% in NDCG@10) under a reasonable privacy budget (e.g., ε=1).
Bo Yan 0005, Yang Cao 0011, Wenchuan Yang, Junping Du 0001, Chuan Shi 0001
WWW2
2024 Differentially private trajectory event streams publishing under data dependence constraints
Yuan Shen 0005, Wei Song 0006, Yang Cao 0011, Zechen Liu, Zhiyong Peng 0001
Inf. Sci.3
2024 Uldp-FL: Federated Learning with Across Silo User-Level Differential Privacy
abstract
Differentially Private Federated Learning (DP-FL) has garnered attention as a collaborative machine learning approach that ensures formal privacy. Most DP-FL approaches ensure DP at the record-level within each silo for cross-silo FL. However, a single user's data may extend across multiple silos, and the desired user-level DP guarantee for such a setting remains unknown. In this study, we present Uldp-FL, a novel FL framework designed to guarantee user-level DP in cross-silo FL where a single user's data may belong to multiple silos. Our proposed algorithm directly ensures user-level DP through per-user weighted clipping, departing from group-privacy approaches. We provide a theoretical analysis of the algorithm's privacy and utility. Additionally, we improve the utility of the proposed algorithm with an enhanced weighting strategy based on user record distribution and design a novel private protocol that ensures no additional information is revealed to the silos and the server. Experiments on real-world datasets show substantial improvements in our methods in privacy-utility trade-offs under user-level DP compared to baseline methods. To the best of our knowledge, our work is the first FL framework that effectively provides user-level DP in the general cross-silo FL setting.
Fumiyuki Kato, Li Xiong 0001, Yang Cao 0011, Masatoshi Yoshikawa
Proc. VLDB Endow.4
2024 HRNet: Differentially Private Hierarchical and Multi-Resolution Network for Human Mobility Data Synthesization
abstract
Human mobility data offers valuable insights for many applications such as urban planning and pandemic response, but its use also raises privacy concerns. In this paper, we introduce the Hierarchical and Multi-Resolution Network (HRNet), a novel deep generative model specifically designed to synthesize realistic human mobility data while guaranteeing differential privacy. We first identify the key difficulties inherent in learning human mobility data under differential privacy. In response to these challenges, HRNet integrates three components: a hierarchical location encoding mechanism, multi-task learning across multiple resolutions, and private pre-training. These elements collectively enhance the model's ability under the constraints of differential privacy. Through extensive comparative experiments utilizing a real-world dataset, HRNet demonstrates a marked improvement over existing methods in balancing the utility-privacy trade-off.
Li Xiong 0001, Fumiyuki Kato, Yang Cao 0011, Masatoshi Yoshikawa
Proc. VLDB Endow.4
2024 Front Matter
Meihui Zhang 0001, Cyrus Shahabi, Ju Fan, Yang Cao 0011, Xiaoou Ding, Divesh Srivastava, Nesime Tatbul, Sihem Amer-Yahia, Yongxin Tong, Yuncheng Wu, Li Xiong 0001, Torsten Grust, Themis Palpanas, Philippe Bonnet, Haixun Wang, Wook-Shin Han, Ibrahim Sabek, M. Tamer Özsu, Xiaofang Zhou 0001
Proc. VLDB Endow.4
2023 PrivateRec: Differentially Private Model Training and Online Serving for Federated News Recommendation
abstract
Federated recommendation can potentially alleviate the privacy concerns in collecting sensitive and personal data for training personalized recommendation systems. However, it suffers from a low recommendation quality when a local serving is inapplicable due to the local resource limitation and the data privacy of querying clients is required in online serving. Furthermore, a theoretically private solution in both the training and serving of federated recommendation is essential but still lacking. Naively applying differential privacy (DP) to the two stages in federated recommendation would fail to achieve a satisfactory trade-off between privacy and utility due to the high-dimensional characteristics of model gradients and hidden representations. In this work, we propose a federated news recommendation method for achieving better utility in model training and online serving under a DP guarantee. We first clarify the DP definition over behavior data for each round in the pipeline of federated recommendation systems. Next, we propose a privacy-preserving online serving mechanism under this definition based on the idea of decomposing user embeddings with public basic vectors and perturbing the lower-dimensional combination coefficients. We apply a random behavior padding mechanism to reduce the required noise intensity for better utility. Besides, we design a federated recommendation model training method, which can generate effective and public basic vectors for serving while providing DP for training participants. We avoid the dimension-dependent noise for large models via label permutation and differentially private attention modules. Experiments on real-world news recommendation datasets validate that our method achieves superior utility under a DP guarantee in both training and serving of federated news recommendations.
Ruixuan Liu, Yang Cao 0011, Yanlin Wang 0001, Lingjuan Lyu, Yun Chen 0007, Hong Chen 0001
KDD2
2023 CSGAN: Modality-Aware Trajectory Generation via Clustering-based Sequence GAN
abstract
Human mobility data is useful for various applications in urban planning, transportation, and public health, but collecting and sharing real-world trajectories can be challenging due to privacy and data quality issues. To address these problems, recent research focuses on generating synthetic trajectories, mainly using generative adversarial networks (GANs) trained by real-world trajectories. In this paper, we hypothesize that by explicitly capturing the modality of transportation (e.g., walking, biking, driving), we can generate not only more diverse and representative trajectories for different modalities but also more realistic trajectories that preserve the geographical density, trajectory, and transition level properties by capturing both cross-modality and modality-specific patterns. Towards this end, we propose a Clustering-based Sequence Generative Adversarial Network (CSGAN) that simultaneously clusters the trajectories based on their modalities and learns the essential properties of real-world trajectories to generate realistic and representative synthetic trajectories. To measure the effectiveness of generated trajectories, in addition to typical density and trajectory level statistics, we define several new metrics for a comprehensive evaluation, including modality distribution and transition probabilities both globally and within each modality. Our extensive experiments with real-world datasets show the superiority of our model in various metrics over state-of-the-art models.
Minxing Zhang, Haowen Lin, Yang Cao 0011, Cyrus Shahabi, Li Xiong 0001
MDM4
2023 Reprogramming Self-supervised Learning-based Speech Representations for Speaker Anonymization
abstract
Current speaker anonymization methods, especially with self-supervised learning (SSL) models, require massive computational resources when hiding speaker identity. This paper proposes an effective and parameter-efficient speaker anonymization method based on recent End-to-End model reprogramming technology. To improve the anonymization performance, we first extract speaker representation from large SSL models as the speaker identifies. To hide the speaker’s identity, we reprogram the speaker representation by adapting the speaker to a pseudo domain. Extensive experiments are carried out on the VoicePrivacy Challenge (VPC) 2022 datasets to demonstrate the effectiveness of our proposed parameter-efficient learning anonymization methods. Additionally, while achieving comparable performance with the VPC 2022 strong baseline 1.b, our approach also consumes less computational resources during anonymization.
Sheng Li 0010, Jiyi Li, Hao Huang 0009, Yang Cao 0011, Liang He 0003
MMAsia5
2023 GhostVec: A New Threat to Speaker Privacy of End-to-End Speech Recognition System
abstract
Speaker adaptation systems face privacy concerns, for such systems are trained on private datasets and often overfitting. This paper demonstrates that an attacker can extract speaker information by querying speaker-adapted speech recognition (ASR) systems. We focus on the speaker information of a transformer-based ASR and propose GhostVec, a simple and efficient attack method to extract the speaker information from an encoder-decoder-based ASR system without any external speaker verification system or natural human voice as a reference. To make our results quantitative, we pre-process GhostVec using singular value decomposition (SVD) and synthesize it into waveform. Experiment results show that the synthesized audio of GhostVec reaches 10.83% EER and 0.47 minDCF with target speakers, which suggests the effectiveness of the proposed method. We hope the preliminary discovery in this study to catalyze future speech recognition research on privacy-preserving topics.
Sheng Li 0010, Jiyi Li, Yang Cao 0011, Hao Huang 0009, Liang He 0003
MMAsia4
2023 Olive: Oblivious Federated Learning on Trusted Execution Environment Against the Risk of Sparsification
abstract
Combining Federated Learning (FL) with a Trusted Execution Environment (TEE) is a promising approach for realizing privacy-preserving FL, which has garnered significant academic attention in recent years. Implementing the TEE on the server side enables each round of FL to proceed without exposing the client's gradient information to untrusted servers. This addresses usability gaps in existing secure aggregation schemes as well as utility gaps in differentially private FL. However, to address the issue using a TEE, the vulnerabilities of server-side TEEs need to be considered---this has not been sufficiently investigated in the context of FL. The main technical contribution of this study is the analysis of the vulnerabilities of TEE in FL and the defense. First, we theoretically analyze the leakage of memory access patterns, revealing the risk of sparsified gradients, which are commonly used in FL to enhance communication efficiency and model accuracy. Second, we devise an inference attack to link memory access patterns to sensitive information in the training dataset. Finally, we propose an oblivious yet efficient aggregation algorithm to prevent memory access pattern leakage. Our experiments on real-world data demonstrate that the proposed method functions efficiently in practical scales.
Fumiyuki Kato, Yang Cao 0011, Masatoshi Yoshikawa
Proc. VLDB Endow.2
2023 Equitable Data Valuation Meets the Right to Be Forgotten in Model Markets
abstract
The increasing demand for data-driven machine learning (ML) models has led to the emergence of model markets, where a broker collects personal data from data owners to produce high-usability ML models. To incentivize data owners to share their data, the broker needs to price data appropriately while protecting their privacy. Forequitable data valuation, which is crucial in data pricing,Shapley valuehas become the most prevalent technique because it satisfies all four desirable properties in fairness: balance, symmetry, zero element, and additivity. Forthe right to be forgotten, which is stipulated by many data privacy protection laws to allow data owners to unlearn their data from trained models, thesharded structurein ML model training has become a de facto standard to reduce the cost of future unlearning by avoiding retraining the entire model from scratch. In this paper, we explore how the sharded structure for the right to be forgotten affects Shapley value for equitable data valuation in model markets. To adapt Shapley value for the sharded structure, we propose S-Shapley value, a sharded structure-based Shapley value, which satisfies four desirable properties for data valuation. Since we prove that computing S-Shapley value is #P-complete, two sampling-based methods are developed to approximate S-Shapley value. Furthermore, to efficiently update valuation results after data owners unlearn their data, we present two delta-based algorithms that estimate the change of data value instead of the data value itself. Experimental results demonstrate the efficiency and effectiveness of the proposed algorithms.
Haocheng Xia, Jinfei Liu, Jian Lou 0001, Zhan Qin, Kui Ren 0001, Yang Cao 0011, Li Xiong 0001
Proc. VLDB Endow.6
2023 Secure Shapley Value for Cross-Silo Federated Learning
abstract
The Shapley value (SV) is a fair and principled metric for contribution evaluation in cross-silo federated learning (cross-silo FL), wherein organizations, i.e., clients, collaboratively train prediction models with the coordination of a parameter server. However, existing SV calculation methods for FL assume that the server can access the raw FL models and public test data. This may not be a valid assumption in practice considering the emerging privacy attacks on FL models and the fact that test data might be clients' private assets. Hence, we investigate the problem of secure SV calculation for cross-silo FL. We first propose HESV , a one-server solution based solely on homomorphic encryption (HE) for privacy protection, which has limitations in efficiency. To overcome these limitations, we propose SecSV , an efficient two-server protocol with the following novel features. First, SecSV utilizes a hybrid privacy protection scheme to avoid ciphertext-ciphertext multiplications between test data and models, which are extremely expensive under HE. Second, an efficient secure matrix multiplication method is proposed for SecSV. Third, SecSV strategically identifies and skips some test samples without significantly affecting the evaluation accuracy. Our experiments demonstrate that SecSV is 7.2--36.6× as fast as HESV, with a limited loss in the accuracy of calculated SVs.
Shuyuan Zheng, Yang Cao 0011, Masatoshi Yoshikawa
Proc. VLDB Endow.2
2022 Boosting Utility of Differentially Private Streaming Data Release under Temporal Correlations
abstract
Although differentially private streaming data release has been studied extensively, how to strike a good balance between privacy and utility on correlated data is still an open problem. Many existing works focus on enhancing privacy when applying differential privacy to correlated data. They show that differential privacy may suffer extra privacy leakage under correlations, and it is inevitable to resort to a small privacy budget to prevent such privacy leakage. However, there is no attempt to solve the consequential utility problem. In this work, for the first time, we propose a post-processing framework to boost the utility of differential privacy data release under temporal correlations. Specifically, we model the problem as a maximum posterior estimation given the released differentially private data and correlation model. We finally transform this problem into a nonlinear constrained programming. Our experiments demonstrate the effectiveness of the proposed approach where the utility and accuracy of differentially private data are significantly improved by nearly ten times in terms of mean square error when a strict privacy budget is given.
Yang Cao 0011, Masatoshi Yoshikawa, Atsuyoshi Nakamura
IEEE Big Data2
2022 A Crypto-Assisted Approach for Publishing Graph Statistics with Node Local Differential Privacy
abstract
Publishing graph statistics under node differential privacy has attracted much attention since it provides a stronger privacy guarantee than edge differential privacy. Existing works related to node differential privacy assume a trusted data curator who holds the whole graph. However, in many applications, a trusted curator is usually not available due to privacy and security issues. In this paper, for the first time, we investigate the problem of publishing graph statistics under Node Local Differential privacy (Node-LDP), which does not rely on a trusted server. We propose an algorithm to publish the degree distribution with Node-LDP by exploring how to select the graph projection parameter in the local setting and how to execute the graph projection locally. Specifically, we propose a crypto-assisted local projection method based on cryptographic primitives, achieving the higher accuracy than our baseline pureLDP local projection method. Furthermore, we improve our baseline graph projection method from node-level to edge-level that preserves more neighboring information, owning better utility. Finally, extensive experiments on real-world graphs show that crypto-assisted parameter selection owns better utility than pureLDP parameter selection, and edge-level local projection provides higher accuracy than node-level local projection, improving by up to 57.2% and 79.8%, respectively.
Shang Liu 0001, Yang Cao 0011, Takao Murakami, Masatoshi Yoshikawa
IEEE Big Data2
2022 Asymmetric Differential Privacy
abstract
Differential privacy (DP) is attracting considerable research attention as a privacy definition when publishing statistics of a dataset. This study focused on addressing the limitation that DP inevitably causes two-sided errors. For example, consider a threshold query that asks whether a counting is above a given threshold or not. An answer through the DP mechanism can cause error. This phenomenon is not desirable for sensitive analysis such as the counting of COVID-19-infected individuals (in a dataset) visiting a specific location; misinformation can result in incorrect decision-making which can increase the epidemic. To the best of our knowledge, the problem is yet to be solved. We proposed a variation of DP, namely asymmetric DP (ADP) to solve the problem. ADP can provide reasonable privacy protection and achieve one-sided errors. Finally, experiments were conducted to evaluate the utility of the proposed mechanism for the epidemic analysis using a real-world dataset. The results of study revealed the feasibility of proposed mechanisms.
Fumiyuki Kato, Yang Cao 0011, Masatoshi Yoshikawa
IEEE Big Data3
2022 FL-Market: Trading Private Models in Federated Learning
abstract
Acquiring a sufficient amount of training data is a significant bottleneck for machine learning (ML) based data analytics. Recently, commoditizing ML models has been proposed as an economical and moderate solution to ML-oriented data acquisition. However, existing model marketplaces assume that the broker can access data owners’ private training data, which may not be realistic in practice. In this paper, to promote trustworthy data acquisition for ML tasks, we propose FL-Market, a locally private model marketplace that protects privacy against not only model buyers but also an untrusted broker. FL-Market decouples ML from the need to centrally gather training data on the broker’s side using federated learning, a privacy-preserving ML paradigm in which data owners collaboratively train an ML model by uploading local gradients (to be aggregated into a global gradient for model updating). Then, FL-Market enables data owners to locally perturb their gradients by local differential privacy and thus further prevents privacy risks. To drive FL-Market, we propose a deep learning-empowered auction mechanism for intelligently deciding the local gradients’ perturbation levels and an optimal aggregation mechanism for aggregating the perturbed gradients. Our auction and aggregation mechanisms can jointly maximize the global gradient’s accuracy, which optimizes model buyers’ utility. Our experiments verify the effectiveness of the proposed mechanisms.
Shuyuan Zheng, Yang Cao 0011, Masatoshi Yoshikawa, Huizhong Li, Qiang Yan 0001
IEEE Big Data2
2022 An Accurate, Flexible and Private Trajectory-Based Contact Tracing System on Untrusted Servers
Ruixuan Cao, Fumiyuki Kato, Yang Cao 0011, Masatoshi Yoshikawa
iiWAS3
2022 Network Shuffling: Privacy Amplification via Random Walks
abstract
Recently, it is shown that shuffling can amplify the central differential privacy guarantees of data randomized with local differential privacy. Within this setup, a centralized, trusted shuffler is responsible for shuffling by keeping the identities of data anonymous, which subsequently leads to stronger privacy guarantees for systems. However, introducing a centralized entity to the originally local privacy model loses some appeals of not having any centralized entity as in local differential privacy. Moreover, implementing a shuffler in a reliable way is not trivial due to known security issues and/or requirements of advanced hardware or secure computation technology.
Seng Pei Liew, Tsubasa Takahashi 0001, Fumiyuki Kato, Yang Cao 0011, Masatoshi Yoshikawa
SIGMOD Conference5
2022 HDPView: Differentially Private Materialized View for Exploring High Dimensional Relational Data
abstract
How can we explore the unknown properties of high-dimensional sensitive relational data while preserving privacy? We study how to construct an explorable privacy-preserving materialized view under differential privacy. No existing state-of-the-art methods simultaneously satisfy the following essential properties in data exploration: workload independence, analytical reliability (i.e., providing error bound for each search query), applicability to high-dimensional data, and space efficiency. To solve the above issues, we propose HDPView, which creates a differentially private materialized view by well-designed recursive bisected partitioning on an original data cube, i.e., count tensor. Our method searches for block partitioning to minimize the error for the counting query, in addition to randomizing the convergence, by choosing the effective cutting points in a differentially private way, resulting in a less noisy and compact view. Furthermore, we ensure formal privacy guarantee and analytical reliability by providing the error bound for arbitrary counting queries on the materialized views. HDPView has the following desirable properties: (a) Workload independence , (b) Analytical reliability , (c) Noise resistance on high-dimensional data , (d) Space efficiency. To demonstrate the above properties and the suitability for data exploration, we conduct extensive experiments with eight types of range counting queries on eight real datasets. HDPView outperforms the state-of-the-art methods in these evaluations.
Fumiyuki Kato, Tsubasa Takahashi 0001, Yang Cao 0011, Seng Pei Liew, Masatoshi Yoshikawa
Proc. VLDB Endow.4
2021 Privacy-Preserving Polynomial Evaluation over Spatio-Temporal Data on an Untrusted Cloud Server
Wei Song 0006, Mengfei Tang, Yuan Shen 0005, Yang Cao 0011, Qian Wang 0002, Zhiyong Peng 0001
DASFAA (1)5
2021 P3GM: Private High-Dimensional Data Release via Privacy Preserving Phased Generative Model
abstract
How can we release a massive volume of sensitive data while mitigating privacy risks? Privacy-preserving data synthesis enables the data holder to outsource analytical tasks to an untrusted third party. The state-of-the-art approach for this problem is to build a generative model under differential privacy, which offers a rigorous privacy guarantee. However, the existing method cannot adequately handle high dimensional data. In particular, when the input dataset contains a large number of features, the existing techniques require injecting a prohibitive amount of noise to satisfy differential privacy, which results in the outsourced data analysis meaningless. To address the above issue, this paper proposes privacy-preserving phased generative model (P3GM), which is a differentially private generative model for releasing such sensitive data. P3GM employs the two-phase learning process to make it robust against the noise, and to increase learning efficiency (e.g., easy to converge). We give theoretical analyses about the learning complexity and privacy loss in P3GM. We further experimentally evaluate our proposed method and demonstrate that P3GM significantly outperforms existing solutions. Compared with the state-of-the-art methods, our generated samples look fewer noises and closer to the original data in terms of data diversity. Besides, in several data mining tasks with synthesized data, our model outperforms the competitors in terms of accuracy.
Tsubasa Takahashi 0001, Yang Cao 0011, Masatoshi Yoshikawa
ICDE3
2021 Protecting Spatiotemporal Event Privacy in Continuous Location-Based Services
abstract
Location privacy-preserving mechanisms (LPPMs) have been extensively studied for protecting users' location privacy by releasing a perturbed location to third parties such as location-based service providers. However, when a user's perturbed locations are released continuously, existing LPPMs may not protect the sensitive information about the user's real-world activities, such as “visited hospital in the last week” or “regularly commuting between location A and location B every weekday” (it is easy to infer that location A and location B may be home and office), which we call it spatiotemporal event. In this paper, we first formally define spatiotemporal event as Boolean expressions between location and time predicates, and then we define ε-spatiotemporal event privacy by extending the notion of differential privacy. Second, to understand how much spatiotemporal event privacy that existing LPPMs can provide, we design computationally efficient algorithms to quantify the spatiotemporal event privacy leakage of state-of-the-art LPPMs. It turns out that the existing LPPMs may not adequately protect spatiotemporal event privacy. Third, we propose a framework, PriSTE, to transform an existing LPPM into one protecting spatiotemporal event privacy by calibrating the LPPM's privacy budgets. Our experiments on real-life and synthetic data verified that the proposed method is effective and efficient.
Yang Cao 0011, Yonghui Xiao, Li Xiong 0001, Liquan Bai, Masatoshi Yoshikawa
IEEE Trans. Knowl. Data Eng.1
2020 Secure and Efficient Trajectory-Based Contact Tracing using Trusted Hardware
abstract
The COVID-19 pandemic has prompted techno-logical measures to control the spread of the disease. Private contact tracing (PCT) is a promising technique for this purpose. However, the recently proposed Bluetooth-based PCT has several limitations in terms of functionality and flexibility. The existing systems are only able to detect direct contact (i.e., human-human contact) but cannot detect indirect contact (i.e., human-object, such as disease transmission through a surface). Moreover, the rule of risky contact cannot be flexibly changed with the environmental situation and the nature of the virus. In this paper, we propose a secure and efficient trajectory-based PCT system using trusted hardware. We formalize trajectory-based PCT as a generalization of the well-studied private set intersection (PSI), which is mostly based on cryptographic primitives and is thus insufficient. We solve the problem by leveraging trusted hardware such as Intel SGX and designing a novel algorithm to achieve a secure, efficient and flexible PCT system. Our experiments on real-world data show that the proposed system can achieve high performance and scalability. Specifically, our system (one single machine with Intel SGX) can process thousands of queries on 100 million records of trajectory data in a few seconds.
Fumiyuki Kato, Yang Cao 0011, Masatoshi Yoshikawa
IEEE BigData2
2020 FedSel: Federated SGD Under Local Differential Privacy with Top-k Dimension Selection
Ruixuan Liu, Yang Cao 0011, Masatoshi Yoshikawa, Hong Chen 0001
DASFAA (1)2
2020 Providing Input-Discriminative Protection for Local Differential Privacy
abstract
Local Differential Privacy (LDP) provides provable privacy protection for data collection without the assumption of the trusted data server. In the real-world scenario, different data have different privacy requirements due to the distinct sensitivity levels. However, LDP provides the same protection for all data. In this paper, we tackle the challenge of providing input-discriminative protection to reflect the distinct privacy requirements of different inputs. We first present the Input- Discriminative LDP (ID-LDP) privacy notion and focus on a specific version termed MinID-LDP, which is shown to be a fine-grained version of LDP. Then, we focus on the application of frequency estimation and develop the IDUE mechanism based on Unary Encoding for single-item input and the extended mechanism IDUE-PS (with Padding-and-Sampling protocol) for item-set input. The results on both synthetic and real-world datasets validate the correctness of our theoretical analysis and show that the proposed mechanisms satisfying MinID-LDP have better utility than the state-of-the-art mechanisms satisfying LDP due to the input-discriminative protection.
Xiaolan Gu, Ming Li 0003, Li Xiong 0001, Yang Cao 0011
ICDE4
2020 Money Cannot Buy Everything: Trading Mobile Data with Controllable Privacy Loss
abstract
As personal data has been the new oil of the digital era, there is a growing trend perceiving personal data as a commodity. Existing studies have built theories on how to map the privacy loss to an arbitrage-free price. They assumed that a data buyer could purchase arbitrarily accurate results as long as she could compensate data owners for their privacy loss. However, it may not be a viable business model under strict privacy regulations, such as GDPR and CCPA, and data owners' emerging privacy concerns. In this paper, we study how to empower data owners with the control of privacy loss when continuously trading their personal mobile data. Concretely, we propose a framework for trading infinite streaming mobile data which enables each data owner to bound her privacy loss in a w-length sliding window. Introducing such upper bounds of privacy loss makes the existing trading frameworks invalid and raises new technical challenges in terms of budget allocation and arbitrage-free pricing. To address these problems, we propose a modularized trading framework with instances that allows data owners to personalize their privacy loss while the price is still arbitrage-free. Finally, we conduct experiments to verify the effectiveness of the proposed trading protocols.
Shuyuan Zheng, Yang Cao 0011, Masatoshi Yoshikawa
MDM2
2020 PANDA: Policy-aware Location Privacy for Epidemic Surveillance
abstract
In this demonstration, we present a privacy-preserving epidemic surveillance system. Recently, many countries that suffer from COVID-19 crises attempt to access citizen's location data to eliminate the outbreak. However, it raises privacy concerns and may open the doors to more invasive forms of surveillance in the name of public health. It also brings a challenge for privacy protection techniques: how can we leverage people's mobile data to help combat the pandemic without scarifying location privacy. We demonstrate that we can achieve this by implementing policy-based location privacy for epidemic surveillance. Our system has three primary functions for epidemic surveillance: people flow monitoring, epidemic analysis, and contact tracing. We provide an interactive tool allowing the attendees to explore and examine the usability of our system: (1) the utility of location monitor and disease transmission model estimation, (2) the procedure of contact tracing in our systems, and (3) the privacy-utility trade-offs w.r.t. different policy graphs. The attendees will find that we can have the high usability for epidemic surveillance while preserving location privacy.
Yang Cao 0011, Yonghui Xiao, Li Xiong 0001, Masatoshi Yoshikawa
Proc. VLDB Endow.1
2019 PriSTE: From Location Privacy to Spatiotemporal Event Privacy
abstract
Location privacy-preserving mechanisms (LPPMs) have been extensively studied for protecting a user's location at each time point or a sequence of locations with different timestamps (i.e., a trajectory). We argue that existing LPPMs are not capable of protecting the sensitive information in user's spatiotemporal activities, such as "visited hospital in the last week" or "regularly commuting between Address 1 and Address 2 every morning and afternoon" (it is easy to infer that Addresses 1 and 2 may be home and office). To address this problem, we define the spatiotemporal event as a new privacy goal, which can be formalized as Boolean expressions between location and time predicates. We show that the spatiotemporal event is a generalization of a single location or a trajectory which is protected by existing LPPMs, while some types of spatiotemporal event may not be protected by the existing LPPMs. Hence, we formally define -spatiotemporal event privacy which is an indistinguishability-based privacy metric. It turns out that, interestingly, such privacy metric is orthogonal to the existing indistinguishability-based location privacy metric such as Geo-indistinguishability. We also discussed the potential solution to achieve both -spatiotemporal event privacy and Geo-indistinguishability.
Yang Cao 0011, Yonghui Xiao, Li Xiong 0001, Liquan Bai
ICDE1
2019 PriSTE: Protecting Spatiotemporal Event Privacy in Continuous Location-Based Services
abstract
Location privacy-preserving mechanisms (LPPMs) have been extensively studied for protecting a user's location in location-based services. However, when user's perturbed locations are released continuously, existing LPPMs may not protect users' sensitive spatiotemporal event , such as "visited hospital in the last week" or "regularly commuting between location 1 and location 2 every morning and afternoon" (it is easy to infer that locations 1 and 2 may be home and office). In this demonstration, we demonstrate PriSTE for protecting spatiotemporal event privacy in continuous location release. First, to raise users' awareness of such a new privacy goal, we design an interactive tool to demonstrate how accurate an adversary could infer a secret spatiotemporal event from a sequence of locations or even LPPM-protected locations. The attendees can find that some spatiotemporal events are quite risky and even these state-of-the-art LPPMs do not always protect spatiotemporal event privacy. Second, we demonstrate how a user can use PriSTE to automatically or manually convert an LPPM for location privacy into one protecting spatiotemporal event privacy in continuous location-based services. Finally, we visualize the trade-off between privacy and utility so that users can choose appropriate privacy parameters in different application scenarios.
Yang Cao 0011, Yonghui Xiao, Li Xiong 0001, Liquan Bai, Masatoshi Yoshikawa
Proc. VLDB Endow.1
2019 Quantifying Differential Privacy in Continuous Data Release Under Temporal Correlations
abstract
Differential Privacy (DP) has received increasing attention as a rigorous privacy framework. Many existing studies employ traditional DP mechanisms (e.g., the Laplace mechanism) as primitives to continuously release private data for protecting privacy at each time point (i.e., event-level privacy), which assume that the data at different time points are independent, or that adversaries do not have knowledge of correlation between data. However, continuously generated data tend to be temporally correlated, and such correlations can be acquired by adversaries. In this paper, we investigate the potential privacy loss of a traditional DP mechanism under temporal correlations. First, we analyze the privacy leakage of a DP mechanism under temporal correlation that can be modeled using Markov Chain. Our analysis reveals that, the event-level privacy loss of a DP mechanism may increase over time. We call the unexpected privacy loss temporal privacy leakage (TPL). Although TPL may increase over time, we find that its supremum may exist in some cases. Second, we design efficient algorithms for calculating TPL. Third, we propose data releasing mechanisms that convert any existing DP mechanism into one against TPL. Experiments confirm that our approach is efficient and effective.
Yang Cao 0011, Masatoshi Yoshikawa, Yonghui Xiao, Li Xiong 0001
IEEE Trans. Knowl. Data Eng.1
2019 Errata on "Quantifying Differential Privacy in Continuous Data Release under Temporal Correlations"
abstract
Presents revisions to the above named paper.
Yang Cao 0011, Masatoshi Yoshikawa, Yonghui Xiao, Li Xiong 0001
IEEE Trans. Knowl. Data Eng.1
2018 ConTPL: Controlling Temporal Privacy Leakage in Differentially Private Continuous Data Release
abstract
In many real-world systems, such as Internet of Thing, sensitive data streams are collected and analyzed continually. To protect privacy, a number of mechanisms are designed to achieve ϵ-differential privacy for processing sensitive streaming data, whose privacy loss is considered to be rigorously controlled within a given parameter ϵ . However, most of the existing studies do not consider the effect of temporal correlations among the continuously generated data on the privacy loss. Our recent work reveals that, the privacy loss of a traditional DP mechanism (e.g., Laplace mechanism) may not be bounded by ϵ due to temporal correlations. We call such unexpected privacy loss Temporal Privacy Leakage (TPL). In this demonstration, we design a system, ConTPL, which is able to automatically convert an existing differentially private streaming data release mechanism into one bounding TPL within a specified level. ConTPL also provides an interactive interface and real-time visualization to help data curator to understand and explore the effect of different parameters on TPL.
Yang Cao 0011, Li Xiong 0001, Masatoshi Yoshikawa, Yonghui Xiao
Proc. VLDB Endow.1
2017 Quantifying Differential Privacy under Temporal Correlations
abstract
Differential Privacy (DP) has received increasing attention as a rigorous privacy framework. Many existing studies employ traditional DP mechanisms (e.g., the Laplace mechanism) as primitives, which assume that the data are independent, or that adversaries do not have knowledge of the data correlations. However, continuous generated data in the real world tend to be temporally correlated, and such correlations can be acquired by adversaries. In this paper, we investigate the potential privacy loss of a traditional DP mechanism under temporal correlations in the context of continuous data release. First, we model the temporal correlations using Markov model and analyze the privacy leakage of a DP mechanism when adversaries have knowledge of such temporal correlations. Our analysis reveals that the privacy loss of a DP mechanism may accumulate and increase over time. We call it temporal privacy leakage. Second, to measure such privacy loss, we design an efficient algorithm for calculating it in polynomial time. Although the temporal privacy leakage may increase over time, we also show that its supremum may exist in some cases. Third, to bound the privacy loss, we propose mechanisms that convert any existing DP mechanism into one against temporal privacy leakage. Experiments with synthetic data confirm that our approach is efficient and effective.
Yang Cao 0011, Masatoshi Yoshikawa, Yonghui Xiao, Li Xiong 0001
ICDE1
2017 LocLok: Location Cloaking with Differential Privacy via Hidden Markov Model
abstract
We demonstrate LocLok, a LOCation-cLOaKing system to protect the locations of a user with differential privacy. LocLok has two features: (a) it protects locations under temporal correlations described through hidden Markov model; (b) it releases the optimal noisy location with the planar isotropic mechanism (PIM), the first mechanism that achieves the lower bound of differential privacy. We show the detailed computation of LocLok with the following components: (a) how to generate the possible locations with Markov model, (b) how to perturb the location with PIM, and (c) how to make inference about the true location in Markov model. An online system with real-word dataset will be presented with the computation details.
Yonghui Xiao, Li Xiong 0001, Yang Cao 0011
Proc. VLDB Endow.4
2015 Differentially Private Real-Time Data Release over Infinite Trajectory Streams
abstract
Recent emerging mobile and wearable technologies make it easy to collect personal spatiotemporal data such as activity trajectories in daily life. Releasing real-time statistics over trajectory streams produced by crowds of people is expected to be valuable for both academia and business, answering questions such as "How many people are in Central Station now?" However, analyzing these raw data will entail risks of compromising individual privacy. ϵ-Differential Privacy has emerged as a de facto standard for private statistics publishing because of its guarantee of being rigorous and mathematically provable. Since user trajectories will be generated infinitely, it is difficult to protect every trajectory under ϵ-differential privacy. To this end, we propose a flexible privacy model of ℓ-trajectory privacy to ensure every length of ℓ trajectories under protection of ϵ-differential privacy. Then we hierarchically design algorithms to satisfy ℓ-trajectory privacy. Experiments using four real-life datasets show that our proposed algorithms are effective and efficient.
Yang Cao 0011, Masatoshi Yoshikawa
MDM (2)1
2013 A User-Friendly Patent Search Paradigm
abstract
As an important operation for finding existing relevant patents and validating a new patent application, patent search has attracted considerable attention recently. However, many users have limited knowledge about the underlying patents, and they have to use a try-and-see approach to repeatedly issue different queries and check answers, which is a very tedious process. To address this problem, in this paper, we propose a new user-friendly patent search paradigm, which can help users find relevant patents more easily and improve user search experience. We propose three effective techniques, error correction, topic-based query suggestion, and query expansion, to improve the usability of patent search. We also study how to efficiently find relevant answers from a large collection of patents. We first partition patents into small partitions based to their topics and classes. Then, given a query, we find highly relevant partitions and answer the query in each of such highly relevant partitions. Finally, we combine the answers of each partition and generate top-$(k)$ answers of the patent-search query.
Yang Cao 0011, Ju Fan, Guoliang Li 0001
IEEE Trans. Knowl. Data Eng.1
2008 Enhancing text clustering by leveraging Wikipedia semantics
abstract
Most traditional text clustering methods are based on "bag of words" (BOW) representation based on frequency statistics in a set of documents. BOW, however, ignores the important information on the semantic relationships between key terms. To overcome this problem, several methods have been proposed to enrich text representation with external resource in the past, such as WordNet. However, many of these approaches suffer from some limitations: 1) WordNet has limited coverage and has a lack of effective word-sense disambiguation ability; 2) Most of the text representation enrichment strategies, which append or replace document terms with their hypernym and synonym, are overly simple. In this paper, to overcome these deficiencies, we first propose a way to build a concept thesaurus based on the semantic relations (synonym, hypernym, and associative relation) extracted from Wikipedia. Then, we develop a unified framework to leverage these semantic relations in order to enhance traditional content similarity measure for text clustering. The experimental results on Reuters and OHSUMED datasets show that with the help of Wikipedia thesaurus, the clustering performance of our method is improved as compared to previous methods. In addition, with the optimized weights for hypernym, synonym, and associative concepts that are tuned with the help of a few labeled data users provided, the clustering performance can be further improved.
Jian Hu 0001, Lujun Fang, Yang Cao 0011, Hua-Jun Zeng, Hua Li 0001, Qiang Yang 0001, Zheng Chen 0001
SIGIR3