VLDB 2026 Research / reviewers in the wild / expert
Jaideep Vaidya
dblp:61/3091
· DBLP profile ↗
39ranked-venue papers in the field
12as first author
6since 2021 · last 2025
0000-0002-7420-6947ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 18 (4 first)Data Mining & Knowledge Discovery · 14 (7 first)Other / Interdisciplinary · 4 (1 first)Information Retrieval & Web Search · 2Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cafe: Improved Federated Data Imputation by Leveraging Missing Data HeterogeneityabstractFederated learning (FL), a decentralized machine learning approach, offers great performance while alleviating autonomy and confidentiality concerns. Despite FL's popularity, how to deal with missing values in a federated manner is not well understood. In this work, we initiate a study of federated imputation of missing values, particularly in complex scenarios, where missing data heterogeneity exists and the state-of-the-art (SOTA) approaches for federated imputation suffer from significant loss in imputation quality. We propose Cafe, a personalized FL approach for missing data imputation. Cafe is inspired from the observation that heterogeneity can induce differences in observable and missing data distribution across clients, and that these differences can be leveraged to improve the imputation quality. Cafe computes personalized weights that are automatically calibrated for the level of heterogeneity, which can remain unknown, to develop personalized imputation models for each client. An extensive empirical evaluation over a variety of settings demonstrates that Cafe matches the performance of SOTA baselines in homogeneous settings while significantly outperforming the baselines in heterogeneous settings. Sitao Min, Hafiz Salman Asif, Xinyue Wang 0003, Jaideep Vaidya |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Data Synthesis Reinvented: Preserving Missing Patterns for Enhanced AnalysisabstractSynthetic data is being widely used as a replacement or enhancement for real data in fields as diverse as healthcare, telecommunications, and finance. Unlike real data, which represents actual people and objects, synthetic data is generated from an estimated distribution that retains key statistical properties of the real data. This makes synthetic data attractive for sharing while addressing privacy, confidentiality, and autonomy concerns. Real data often contains missing values that hold important information about individual, system, or organizational behavior. Standard synthetic data generation methods eliminate missing values as part of their pre-processing steps and thus completely ignore this valuable source of information. Instead, we propose methods to generate synthetic data that preserve both the observable and missing data distributions; consequently, retaining the valuable information encoded in the missing patterns of the real data. Our approach handles various missing data scenarios and can easily integrate with existing data generation methods. Extensive empirical evaluations on diverse datasets demonstrate the effectiveness of our approach as well as the value of preserving missing data distribution in synthetic data. Xinyue Wang 0003, Hafiz Salman Asif, Jaideep Vaidya |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Preserving Missing Data Distribution in Synthetic DataabstractData from Web artifacts and from the Web is often sensitive and cannot be directly shared for data analysis. Therefore, synthetic data generated from the real data is increasingly used as a privacy-preserving substitute. In many cases, real data from the web has missing values where the missingness itself possesses important informational content, which domain experts leverage to improve their analysis. However, this information content is lost if either imputation or deletion is used before synthetic data generation. In this paper, we propose several methods to generate synthetic data that preserve both the observable and the missing data distributions. An extensive empirical evaluation over a range of carefully fabricated and real world datasets demonstrates the effectiveness of our approach. Xinyue Wang 0003, Hafiz Salman Asif, Jaideep Vaidya |
WWW | 3 |
| 2023 | Identifying Anomalies While Preserving PrivacyabstractIdentifying anomalies in data is vital in many domains, including medicine, finance, and national security. However, privacy concerns pose a significant roadblock to carrying out such an analysis. Since existing privacy definitions do not allow good accuracy when doing outlier analysis, the notion of sensitive privacy has been recently proposed to deal with this problem. Sensitive privacy makes it possible to analyze data for anomalies with practically meaningful accuracy while providing a strong guarantee similar to differential privacy, which is the prevalent privacy standard today. In this work, we relate sensitive privacy to other important notions of data privacy so that one can port the technical developments and private mechanism constructions from these related concepts to sensitive privacy. Sensitive privacy critically depends on the underlying anomaly model. We develop a novel n-step lookahead mechanism to efficiently answer arbitrary outlier queries, which provably guarantees sensitive privacy if we restrict our attention to common a class of anomaly models. We also provide general constructions to give sensitively private mechanisms for identifying anomalies and show the conditions under which the constructions would be optimal. Hafiz Salman Asif, Jaideep Vaidya, Periklis A. Papakonstantinou |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | A Generalized Framework for Preserving Both Privacy and Utility in Data OutsourcingabstractProperty preserving encryption techniques have significantly advanced the utility of encrypted data in data outsourcing. However, while preserving certain properties (e.g., the prefixes or order of the data) in the encrypted data, such encryption schemes are typically limited to specific data types (e.g., IP addresses) or applications (e.g., range queries over order-preserved data), and highly vulnerable to the emerging inference attacks which may greatly limit their applications in practice. In this paper, to the best of our knowledge, we make the first attempt to generalize the prefix-preserving encryption to make it applicable to more general data types (e.g., geo-locations, market basket data, DNA sequences, numerical data and timestamps) and secure against the inference attacks. Furthermore, we present a generalized multi-view outsourcing framework that generates multiple indistinguishable data views in which one view fully preserves the utility for data analysis, and its accurate analysis result can be obliviously retrieved. We empirically evaluate the performance of our outsourcing framework against two common inference attacks on two different real datasets: the check-in location dataset and network traffic dataset. The experimental results demonstrate that our proposed framework preserves both privacy (with bounded leakage and indistinguishable data views) and utility (with 100% analysis accuracy). Shangyu Xie, Meisam Mohammady, Han Wang 0021, Lingyu Wang 0001, Jaideep Vaidya, Yuan Hong 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | A Generalized Framework for Preserving Both Privacy and Utility in Data Outsourcing (Extended Abstract)abstractIn this paper, we propose a prefix-preserving encryption based data outsourcing framework which is applicable to multiple different types of data, such as geo-locations, market basket data, DNA sequences, numerical data and timestamps. It enables accurate data analyses on the encrypted data while ensuring strong privacy against inference attacks. The basic idea is to generates multiple indistinguishable data views in which one view fully preserves the utility for data analysis, and its accurate analysis result can be obliviously retrieved. We empirically evaluate the performance of our outsourcing framework against two common inference attacks on two different real datasets: the check-in location dataset and network traffic dataset, respectively. The experimental results demonstrate that our proposed framework preserves both privacy (with bounded leakage and indistinguishability of data views) and utility. Shangyu Xie, Meisam Mohammady, Han Wang 0021, Lingyu Wang 0001, Jaideep Vaidya, Yuan Hong 0001 |
ICDE | 5 |
| 2020 | Publishing Video Data with Indistinguishable Objectsabstractfor all the predefined sensitive objects (e.g., humans and vehicles) in the video, and then propose a video sanitization technique VERRO that randomly generates utility-driven synthetic videos with indistinguishable objects. Therefore, all the objects can be well protected in the generated utility-driven synthetic videos which can be disclosed to any untrusted video recipient. We have conducted extensive experiments on three real videos captured for pedestrians on the streets. The experimental results demonstrate that the generated synthetic videos lie close to the original video for retaining good utility while ensuring rigorous privacy guarantee. Han Wang 0021, Yuan Hong 0001, Yu Kong 0001, Jaideep Vaidya |
EDBT | 4 |
| 2018 | Differentially Private Outlier Detection in a Collaborative EnvironmentabstractOutlier detection is one of the most important data analytics tasks and is used in numerous applications and domains. The goal of outlier detection is to find abnormal entities that are significantly different from the remaining data. Often the underlying data is distributed across different organizations. If outlier detection is done locally, the results obtained are not as accurate as when outlier detection is done collaboratively over the combined data. However, the data cannot be easily integrated into a single database due to privacy and legal concerns. In this paper, we address precisely this problem. We first define privacy in the context of collaborative outlier detection. We then develop a novel method to find outliers from both horizontally partitioned and vertically partitioned categorical data in a privacy-preserving manner. Our method is based on a scalable outlier detection technique that uses attribute value frequencies. We provide an end-to-end privacy guarantee by using the differential privacy model and secure multiparty computation techniques. Experiments on real data show that our proposed technique is both effective and efficient. Hafiz Salman Asif, Tanay Talukdar, Jaideep Vaidya, Basit Shafiq, Nabil R. Adam |
Int. J. Cooperative Inf. Syst. | 3 |
| 2014 | Efficient Integrity Verification for Outsourced Collaborative FilteringabstractCollaborative filtering (CF) over large datasets requires significant computing power. Due to this data owning organizations often outsource the computation of CF (including some abstraction of the data itself) to a public cloud infrastructure. However, this leads to the question of how to verify the integrity of the outsourced computation. In this paper, we develop verification mechanisms for two popular item based collaborative filtering techniques. We further analyze the cheating behavior of the cloud from the game-theoretic perspective. Coupled with the right incentives, we can ensure that the computation is incentive compatible thus ensuring that a rational adversary will not cheat. Leveraging this, we can develop efficient and effective mechanisms to address the problem of integrity in outsourcing. Jaideep Vaidya, Ibrahim Yakut, Anirban Basu 0001 |
ICDM | 1 |
| 2013 | Differentially Private Naive Bayes ClassificationabstractPrivacy and security concerns often prevent the sharing of users' data or even of the knowledge gained from it, thus deterring valuable information from being utilized. Privacy-preserving knowledge discovery, if done correctly, can alleviate this problem. One of the most important and widely used data mining techniques is that of classification. We consider the model where a single provider has centralized access to a dataset and would like to release a classifier while protecting privacy to the best extent possible. Recently, the model of differential privacy has been developed which provides a strong privacy guarantee even if adversaries hold arbitrary prior knowledge. In this paper, we apply this rigorous privacy model to develop a Naive Bayes classifier, which is often used as a baseline and consistently provides reasonable classification performance. We experimentally evaluate the proposed approach, and discuss how it could be potentially deployed in PaaS clouds. Jaideep Vaidya, Basit Shafiq, Anirban Basu 0001, Yuan Hong 0001 |
Web Intelligence | 1 |
| 2012 | Differentially private search log sanitization with optimal output utilityabstractWeb search logs contain extremely sensitive data, as evidenced by the recent AOL incident. However, storing and analyzing search logs can be very useful for many purposes (i.e. investigating human behavior). Thus, an important research question is how to privately sanitize search logs. Several search log anonymization techniques have been proposed with concrete privacy models. However, in all of these solutions, the output utility of the techniques is only evaluated rather than being maximized in any fashion. Indeed, for effective search log anonymization, it is desirable to derive the outputs with optimal utility while meeting the privacy standard. In this paper, we propose utility-maximizing sanitization based on the rigorous privacy standard of differential privacy, in the context of search logs. Specifically, we utilize optimization models to maximize the output utility of the sanitization for different applications, while ensuring that the production process satisfies differential privacy. An added benefit is that our novel randomization strategy maintains the schema integrity in the output search logs. A comprehensive evaluation on real search logs validates the approach and demonstrates its robustness and scalability. Yuan Hong 0001, Jaideep Vaidya, Haibing Lu, Mingrui Wu |
EDBT | 2 |
| 2012 | Boolean Matrix Decomposition Problem: Theory, Variations and Applications to Data EngineeringabstractWith the ubiquitous nature and sheer scale of data collection, the problem of data summarization is most critical for effective data management. Classical matrix decomposition techniques have often been used for this purpose, and have been the subject of much study. In recent years, several other forms of decomposition, including Boolean Matrix Decomposition have become of significant practical interest. Since much of the data collected is categorical in nature, it can be viewed in terms of a Boolean matrix. Boolean matrix decomposition (BMD), wherein a boolean matrix is expressed as a product of two Boolean matrices, can be used to provide concise and interpretable representations of Boolean data sets. The decomposed matrices give the set of meaningful concepts and their combination which can be used to reconstruct the original data. Such decompositions are useful in a number of application domains including role engineering, text mining as well as knowledge discovery from databases. In this seminar, we look at the theory underlying the BMD problem, study some of its variants and solutions, and examine different practical applications. Jaideep Vaidya |
ICDE | 1 |
| 2012 | Anonymizing set-valued data by nonreciprocal recodingabstractToday there is a strong interest in publishing set-valued data in a privacy-preserving manner. Such data associate individuals to sets of values (e.g., preferences, shopping items, symptoms, query logs). In addition, an individual can be associated with a sensitive label (e.g., marital status, religious or political conviction). Anonymizing such data implies ensuring that an adversary should not be able to (1) identify an individual's record, and (2) infer a sensitive label, if such exists. Existing research on this problem either perturbs the data, publishes them in disjoint groups disassociated from their sensitive labels, or generalizes their values by assuming the availability of a generalization hierarchy. In this paper, we propose a novel alternative. Our publication method also puts data in a generalized form, but does not require that published records form disjoint groups and does not assume a hierarchy either; instead, it employs generalized bitmaps and recasts data values in a nonreciprocal manner; formally, the bipartite graph from original to anonymized records does not have to be composed of disjoint complete subgraphs. We configure our schemes to provide popular privacy guarantees while resisting attacks proposed in recent research, and demonstrate experimentally that we gain a clear utility advantage over the previous state of the art. Mingqiang Xue, Panagiotis Karras, Chedy Raïssi, Jaideep Vaidya, Kian-Lee Tan |
KDD | 4 |
| 2011 | Weighted Rank-One Binary Matrix FactorizationabstractMining discrete patterns in binary data is important for many data analysis tasks, such as data sampling, compression, and clustering. An example is that replacing individual records with their patterns would greatly reduce data size and simplify subsequent data analysis tasks. As a straightforward approach, rank-one binary matrix approximation has been actively studied recently for mining discrete patterns from binary data. It factorizes a binary matrix into the multiplication of one binary pattern vector and one binary presence vector, while minimizing mismatching entries. However, this approach suffers from two serious problems. First, if all records are replaced with their respective patterns, the noise could make as much as 50% in the resulting approximate data. This is because the approach simply assumes that a pattern is present in a record as long as their matching entries are more than their mismatching entries. Second, two error types, 1-becoming-0 and 0-becoming-1, are treated evenly, while in many application domains they are discriminated. To address the two issues, we propose weighted rank-one binary matrix approximation. It enables the tradeoff between the accuracy and succinctness in approximate data and allows users to impose their personal preferences on the importance of different error types. The decision problem, however, as proved in the paper is NP-complete. To solve it, several different mathematical programming formulations are provided, from which 2-approximation algorithms are derived for some special cases. An adaptive tabu search heuristic is presented for solving the general problem, and our experimental study shows the effectiveness of the heuristic. Haibing Lu, Jaideep Vaidya, Vijayalakshmi Atluri, Heechang Shin, Lili Jiang 0001 |
SDM | 2 |
| 2011 | Search Engine Query Clustering Using Top-k Search ResultsabstractClustering of search engine queries has attracted significant attention in recent years. Many search engine applications such as query recommendation require query clustering as a pre-requisite to function properly. Indeed, clustering is necessary to unlock the true value of query logs. However, clustering search queries effectively is quite challenging, due to the high diversity and arbitrary input by users. Search queries are usually short and ambiguous in terms of user requirements. Many different queries may refer to a single concept, while a single query may cover many concepts. Existing prevalent clustering methods, such as K-Means or DBSCAN cannot assure good results in such a diverse environment. Agglomerative clustering gives good results but is computationally quite expensive. This paper presents a novel clustering approach based on a key insight -- search engine results might themselves be used to identify query similarity. We propose a novel similarity metric for diverse queries based on the ranked URL results returned by a search engine for queries. This is used to develop a very efficient and accurate algorithm for clustering queries. Our experimental results demonstrate more accurate clustering performance, better scalability and robustness of our approach against known baselines. Yuan Hong 0001, Jaideep Vaidya, Haibing Lu |
Web Intelligence | 2 |
| 2010 | Ensuring Privacy and Security for LBS through Trajectory PartitioningabstractThe concept of location k-anonymity has been proposed to address the privacy issue of location based services (LBS). Under this notion of anonymity, the adversary only has the knowledge that the LBS request originates from a region containing at least k people, and therefore cannot individually distinguish the requestor. However, new types of LBS services such as continuous nearest neighbor searches require the knowledge of the user's trajectory, which can lead to a privacy breach. The longer the adversary can track the user's trajectory, the stronger the possibility that the user's sensitive information is revealed. To alleviate this problem, we propose algorithms to optimally partition a continuous request into multiple LBS requests with shorter trajectories. This results in increased privacy due to the unlinking of different requests over time and has the added benefit of improving the overall quality of service since the anonymized regions are now smaller. Our experimental results show that significant privacy and QoS benefits can be achieved with nominal computational overhead. Heechang Shin, Jaideep Vaidya, Vijayalakshmi Atluri, Sungyong Choi |
Mobile Data Management | 2 |
| 2010 | Reachability Analysis in Privacy-Preserving Perturbed GraphsabstractMany real world phenomena can be naturally modeled as graph structures whose nodes representing entities and whose edges representing interactions or relationships between entities. The analysis of the graph data have many practical implications. However, the release of the data often poses considerable privacy risk to the individuals involved. In this paper, we address the edge privacy problem in graphs. In particular, we explore random perturbation for privacy preservation in graph data, and propose an iterative derivation process to analyze node reachability within the graph. We specifically focus on deriving the probability that the shortest path linking two nodes in a directed graph is of a particular length. This allows us to determine the expected length of the shortest path between two nodes, and determine whether they are linked or not. The performance of the proposed method is demonstrated via extensive experiments on both real and synthetic datasets. Xiaoyun He, Jaideep Vaidya, Basit Shafiq, Nabil R. Adam, Xiaodong Lin 0004 |
Web Intelligence | 2 |
| 2010 | Spatial neighborhood based anomaly detection in sensor datasets
Vandana Pursnani Janeja, Nabil R. Adam, Vijayalakshmi Atluri, Jaideep Vaidya |
Data Min. Knowl. Discov. | 4 |
| 2010 | Efficient privacy-preserving similar document detection
Mummoorthy Murugesan, Wei Jiang 0026, Chris Clifton, Luo Si, Jaideep Vaidya |
VLDB J. | 5 |
| 2009 | Effective anonymization of query logsabstractUser search query logs have proven to be very useful, but have vast potential for misuse. Several incidents have shown that simple removal of identifiers is insufficient to protect the identity of users. Publishing such inadequately anonymized data can cause severe breach of privacy. While significant effort has been expended on coming up with anonymity models and techniques for microdata, there is little corresponding work for query log data. Query logs are different in several important aspects, such as the diversity of queries and the causes of privacy breach. This necessitates the need to design privacy models and techniques specific to this environment. This paper takes a first cut at tackling this challenge. Our main contribution is to define effective anonymization models for query log data along with proposing techniques to achieve such anonymization. We analyze the inherent utility and privacy tradeoff, and experimentally validate the performance of our techniques. Yuan Hong 0001, Xiaoyun He, Jaideep Vaidya, Nabil R. Adam, Vijayalakshmi Atluri |
CIKM | 3 |
| 2009 | An efficient online auditing approach to limit private data disclosureabstractIn a database system, disclosure of confidential private data may occur if users can put together the answers of past queries. Traditional access control mechanisms cannot guard against such breaches to private data. Online auditing techniques have been advanced to limit such disclosure of private data. Essentially, before answering any query, these techniques inspect the answers of the past queries to determine whether answering this query would compromise the stated data disclosure policies. While the primary requirement for online auditing is high efficiency, existing auditing approaches are expensive with respect to both computational time and space. Specifically, this cost is excessive in the general case of auditing arbitrary aggregate queries over real-valued confidential attributes with respect to interval-based privacy disclosure. Haibing Lu, Yingjiu Li, Vijayalakshmi Atluri, Jaideep Vaidya |
EDBT | 4 |
| 2009 | Extended Boolean Matrix DecompositionabstractWith the vast increase in collection and storage of data, the problem of data summarization is most critical for effective data management. Since much of this data is categorical in nature, it can be viewed in terms of a Boolean matrix. Boolean matrix decomposition (BMD) has been used to provide concise and interpretable representations of Boolean data sets. A Boolean matrix can be expressed as a product of two Boolean matrices, where the first matrix represents a set of meaningful concepts, and the second describes how the observed data can be expressed as combinations of those concepts. Typically, the combination is only in terms of the set union. In other words, a successful Boolean matrix decomposition gives a set of concepts and shows how every column of the input data can be expressed as a union of some subset of those concepts. However, this way of modeling only incompletely represents real data semantics. Essentially, it ignores a critical component -- the set difference operation: a column can be expressed as the combination of union of certain concepts as well as the exclusion of other concepts. This has two significant benefits. First, the total number of concepts required to describe the data may itself be reduced. Second, a more succinct summarization may be found for every column. In this paper, we propose the extended Boolean matrix decomposition (EBMD) problem, which aims to factor Boolean matrices using both the set union and set difference operations. We study several variants of the problem, show that they are NP-hard, and propose efficient heuristics to solve them. Extensive experimental results demonstrate the power of EBMD. Haibing Lu, Jaideep Vaidya, Vijayalakshmi Atluri, Yuan Hong 0001 |
ICDM | 2 |
| 2009 | Efficient Privacy-Preserving Link Discovery
Xiaoyun He, Jaideep Vaidya, Basit Shafiq, Nabil R. Adam, Evimaria Terzi, Tyrone Grandison |
PAKDD | 2 |
| 2009 | An Efficient Approximate Protocol for Privacy-Preserving Association Rule Mining
Murat Kantarcioglu, Robert Nix, Jaideep Vaidya |
PAKDD | 3 |
| 2009 | Preserving Privacy in Social Networks: A Structure-Aware ApproachabstractGraph structured data can be ubiquitously found in the real world. For example, social networks can easily be represented as graphs where the graph connotes the complex sets of relationships between members of social systems. While their analysis could be beneficial in many aspects, publishing certain types of social networks raises significant privacy concerns. This brings the problem of graph anonymization into sharp focus. Unlike relational data, the true information in graph structured data is encoded within the structure and graph properties. Motivated by this, we propose a structure aware anonymization approach that maximally preserves the structure of the original network as well as its structural properties while anonymizing it. Instead of anonymizing each node one by one independently, our approach treats each partitioned substructural component of the network as one single unit to be anonymized. This maximizes utility while enabling anonymization. We apply our method to both synthetic and real datasets and demonstrate its effectiveness and practical usefulness. Xiaoyun He, Jaideep Vaidya, Basit Shafiq, Nabil R. Adam, Vijayalakshmi Atluri |
Web Intelligence | 2 |
| 2009 | Privacy-Preserving Kth Element Score over Vertically Partitioned DataabstractGiven a large integer data set shared vertically by two parties, we consider the problem of securely computing a score separating the kth and the (k + 1) to compute such a score while revealing little additional information. The proposed protocol is implemented using the Fairplay system and experimental results are reported. We show a real application of this protocol as a component used in the secure processing of top-k queries over vertically partitioned data. Jaideep Vaidya, Chris Clifton |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2009 | Privacy-preserving indexing of documents on the network
Mayank Bawa, Roberto J. Bayardo, Rakesh Agrawal 0001, Jaideep Vaidya |
VLDB J. | 4 |
| 2008 | Optimal Boolean Matrix Decomposition: Application to Role EngineeringabstractA decomposition of a binary matrix into two matrices gives a set of basis vectors and their appropriate combination to form the original matrix. Such decomposition solutions are useful in a number of application domains including text mining, role engineering as well as knowledge discovery. While a binary matrix can be decomposed in several ways, however, certain decompositions better characterize the semantics associated with the original matrix in a succinct but comprehensive way. Indeed, one can find different decompositions optimizing different criteria matching various semantics. In this paper, we first present a number of variants to the optimal Boolean matrix decomposition problem that have pragmatic implications. We then present a unified framework for modeling the optimal binary matrix decomposition and its variants using binary integer programming. Such modeling allows us to directly adopt the huge body of heuristic solutions and tools developed for binary integer programming. Although the proposed solutions are applicable to any domain of interest, for providing more meaningful discussions and results, in this paper, we present the binary matrix decomposition problem in a role engineering context, whose goal is to discover an optimal and correct set of roles from existing permissions, referred to as the role mining problem (RMP). This problem has gained significant interest in recent years as role based access control has become a popular means of enforcing security in databases. We consider several variants of the above basic RMP, including the min-noise RMP, delta-approximate RMP and edge-RMP. Solutions to each of them aid security administrators in specific scenarios. We then model these variants as Boolean matrix decomposition and present efficient heuristics to solve them. Haibing Lu, Jaideep Vaidya, Vijayalakshmi Atluri |
ICDE | 2 |
| 2008 | A Profile Anonymization Model for Privacy in a Personalized Location Based Service EnvironmentabstractLocation based services (LBS) aim at delivering point of need information. Personalization and customization of such services, based on the profiles of mobile users, would significantly increase the value of these services. Since profiles may include sensitive information of mobile users and moreover can help identify a person, customization is allowed only when the security and privacy policies dictated by them are respected. While LBS are often presumed as untrusted entities, the location services that capture and maintain mobile users' location to enable communication are considered trusted, and therefore can capture and manage the profile information. In this paper, we address the problem of privacy preservation via anonymization. Prior research in this area attempts to ensure k-anonymity by generalizing the location. However, a person may still be identified based on his/her profile if the profiles of all k people are not the same. We extend the notion of k-anonymity by proposing a profile based k-anonymization model that guarantees anonymity even when profiles of mobile users are known to untrusted entities. Specifically, our proposed approaches generalize both location and profiles to the extent specified by the user. We support three types of queries - mobile users requesting stationary resources, stationary users requesting mobile resources, and mobile users requesting mobile resources. We propose a novel unified index structure, called the (PTPR- tree), which organizes both the locations of mobile users as well as their profiles using a single index, and as a result, offers significant performance gain during anonymization as well as query processing. Heechang Shin, Vijayalakshmi Atluri, Jaideep Vaidya |
MDM | 3 |
| 2008 | Privacy-preserving SVM classification
Jaideep Vaidya, Hwanjo Yu, Xiaoqian Jiang |
Knowl. Inf. Syst. | 1 |
| 2008 | Privacy-preserving decision trees over vertically partitioned dataabstractPrivacy and security concerns can prevent sharing of data, derailing data-mining projects. Distributed knowledge discovery, if done correctly, can alleviate this problem. We introduce a generalized privacy-preserving variant of the ID3 algorithm for vertically partitioned data distributed over two or more parties. Along with a proof of security, we discuss what would be necessary to make the protocols completely secure. We also provide experimental results, giving a first demonstration of the practical complexity of secure multiparty computation-based data mining. Jaideep Vaidya, Chris Clifton, Murat Kantarcioglu, A. Scott Patterson |
ACM Trans. Knowl. Discov. Data | 1 |
| 2008 | Privacy-preserving Naïve Bayes classification
Jaideep Vaidya, Murat Kantarcioglu, Chris Clifton |
VLDB J. | 1 |
| 2006 | Privacy-Preserving SVM Classification on Vertically Partitioned Data
Hwanjo Yu, Jaideep Vaidya, Xiaoqian Jiang |
PAKDD | 2 |
| 2005 | Knowledge Discovery from Transportation Network DataabstractTransportation and logistics are a major sector of the economy, however data analysis in this domain has remained largely in the province of optimization. The potential of data mining and knowledge discovery techniques is largely untapped. Transportation networks are naturally represented as graphs. This paper explores the problems in mining of transportation network graphs: we hope to find how current techniques both succeed and fail on this problem, and from the failures, we hope to present new challenges for data mining. Experimental results from applying both existing graph mining and conventional data mining techniques to real transportation network data are provided, including new approaches to making these techniques applicable to the problems. Reasons why these techniques are not appropriate are discussed. We also suggest several challenging problems to precipitate research and galvanize future work in this area. Wei Jiang 0026, Jaideep Vaidya, Zahir Balaporia, Chris Clifton, Brett Banich |
ICDE | 2 |
| 2005 | Privacy-Preserving Top-K QueriesabstractThe primary contribution of this paper is a secure method for doing top-k selection from vertically partitioned data. This has particular relevance to privacy-sensitive searches, and meshes well with privacy policies such as k-anonymity. We have demonstrated how secure primitives from the literature can be composed with efficient query processing algorithms, with the result having provable security properties. The paper also shows a trade-off between efficiency and disclosure. It is worth exploring whether one could have a suite of algorithms to optimize these tradeoffs, e.g., algorithms that guarantee k-anonymity with efficiency based on the choice of k rather than the guarantees of secure multiparty computation. Jaideep Vaidya, Chris Clifton |
ICDE | 1 |
| 2004 | Privacy-Preserving Outlier DetectionabstractOutlier detection can lead to the discovery of truly unexpected knowledge in many areas such as electronic commerce, credit card fraud and especially national security. We look at the problem of finding outliers in large distributed databases where privacy/security concerns restrict the sharing of data. Both homogeneous and heterogeneous distribution of data is considered. We propose techniques to detect outliers in such scenarios while giving formal guarantees on the amount of information disclosed. Jaideep Vaidya, Chris Clifton |
ICDM | 1 |
| 2004 | Privacy Preserving Naïve Bayes Classifier for Vertically Partitioned DataabstractPrivacy-Preserving Data Mining -- developing models without seeing the data -- is receiving growing attention. This paper assumes a privacy-preserving distributed data mining scenario: data sources collaborate to develop a global model, but must not disclose their data to others. Nave Bayes is often used as a baseline classifier, consistently providing reasonable classification performance. This paper brings privacy-preservation to Nave Bayes classification on vertically partitioned data. Jaideep Vaidya, Chris Clifton |
SDM | 1 |
| 2003 | Privacy-preserving k-means clustering over vertically partitioned dataabstractPrivacy and security concerns can prevent sharing of data, derailing data mining projects. Distributed knowledge discovery, if done correctly, can alleviate this problem. The key is to obtain valid results, while providing guarantees on the (non)disclosure of data. We present a method for k-means clustering when different sites contain different attributes for a common set of entities. Each site learns the cluster of each entity, but learns nothing about the attributes at other sites. Jaideep Vaidya, Chris Clifton |
KDD | 1 |
| 2002 | Privacy preserving association rule mining in vertically partitioned dataabstractPrivacy considerations often constrain data mining projects. This paper addresses the problem of association rule mining where transactions are distributed across sources. Each site holds some attributes of each transaction, and the sites wish to collaborate to identify globally valid association rules. However, the sites must not reveal individual transaction data. We present a two-party algorithm for efficiently discovering frequent itemsets with minimum support levels, without either site revealing individual transaction values. Jaideep Vaidya, Chris Clifton |
KDD | 1 |