VLDB 2026 Research / reviewers in the wild / expert
Chris Clifton
dblp:c/CClifton · also Christopher W. Clifton
· DBLP profile ↗
80ranked-venue papers
15as first author
7since 2021 · last 2024
0000-0001-7274-1471ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 50 · 6 first-author · 2 since 2021Security and privacy · 19 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 18 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Systems, architecture and hardware · 2 · 2 first-authorComputer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Do Crowdsourced Fairness Preferences Correlate with Risk Perceptions?abstractWith the increasing prevalence of automatic decision-making systems, concerns regarding the fairness of these systems also arise. Without a universally agreed-upon definition of fairness, given an automated decision-making scenario, researchers often adopt a crowdsourced approach to solicit people’s preferences across multiple fairness definitions. However, it is often found that crowdsourced fairness preferences are highly context-dependent, making it intriguing to explore the driving factors behind these preferences. One plausible hypothesis is that people’s fairness preferences reflect their perceived risk levels for different decision-making mistakes, such that the fairness definition that equalizes across groups the type of mistakes that are perceived as most serious will be preferred. To test this conjecture, we conduct a human-subject study (N = 213) to study people’s fairness perceptions in three societal contexts. In particular, these three societal contexts differ on the expected level of risk associated with different types of decision mistakes, and we elicit both people’s fairness preferences and risk perceptions for each context. Our results show that people can often distinguish between different levels of decision risks across different societal contexts. However, we find that people’s fairness preferences do not vary significantly across the three selected societal contexts, except for within a certain subgroup of people (e.g., people with a certain racial background). As such, we observe minimal evidence suggesting that people’s risk perceptions of decision mistakes correlate with their fairness preference. These results highlight that fairness preferences are highly subjective and nuanced, and they might be primarily affected by factors other than the perceived risks of decision mistakes. Chowdhury Mohammad Rakin Haider, Chris Clifton, Ming Yin 0001 |
IUI | 2 |
| 2023 | On Improving Fairness of AI Models with Synthetic Minority Oversampling TechniquesabstractBiased AI models result in unfair decisions. In response, a number of algorithmic solutions have been engineered to mitigate bias, among which the Synthetic Minority Oversampling Technique (SMOTE) has been studied, to an extent. Although the SMOTE technique and its variants have great potentials to help improve fairness, there is little theoretical justification for its success. In addition, formal error and fairness bounds are not clearly given. This paper attempts to address both issues. We prove and demonstrate that synthetic data generated by oversampling underrepresented groups can mitigate algorithmic bias in AI models, while keeping the predictive errors bounded. We further compare this technique to the existing state-of-the-art fair AI techniques on five datasets using a variety of fairness metrics. We show that this approach can effectively improve fairness even when there is a significant amount of label and selection bias, regardless of the baseline AI algorithm. Yan Zhou 0001, Murat Kantarcioglu, Chris Clifton |
SDM | 3 |
| 2023 | Statistical limitations of sensitive itemset hiding methods
Shalini Jangra, Durga Toshniwal, Chris Clifton |
Appl. Intell. | 3 |
| 2023 | LuMaMi28: Real-Time Millimeter-Wave Multi-User MIMO Systems With Antenna SelectionabstractThis paper presents LuMaMi28, a real-time 28 GHz multi-user (MU) multiple-input multiple-output (MIMO) testbed. In this testbed, the base station has 16 transceiver chains with a fully-digital beamforming architecture (with different pre-coding algorithms) and simultaneously supports multiple user equipments (UEs) with spatial multiplexing. The UEs are equipped with a beam-switchable antenna array for real-time antenna selection where the one with the highest channel magnitude, out of four pre-defined beams, is selected. For the beam-switchable antenna array, we consider two kinds of UE antennas, with different beam-width and different peak-gain. Based on this testbed, we provide measurement results for millimeter-wave (mmWave) MU-MIMO performance in different real-life scenarios with static and mobile UEs. We explore the potential benefit of the mmWave MU-MIMO systems with antenna selection based on measured channel data, and discuss the performance results through real-time measurements. MinKeun Chung, Liang Liu 0002, Andreas Johansson, Sara Willhammar, Zhinong Ying, Olof Zander, Kamal Samanta, Chris Clifton, Toshiyuki Koimori, Shinya Morita, Satoshi Taniguchi, Fredrik Tufvesson, Ove Edfors |
IEEE Trans. Wirel. Commun. | 9 |
| 2022 | Unfair AI: It Isn't Just Biased DataabstractConventional wisdom holds that discrimination in machine learning is a result of historical discrimination: biased training data leads to biased models. We show that the reality is more nuanced; machine learning can be expected to induce types of bias not found in the training data. In particular, if different groups have different optimal models, and the optimal model for one group has higher accuracy, the optimal accuracy joint model will induce disparate impact even when the training data does not display disparate impact. We argue that due to systemic bias, this is a likely situation, and simply ensuring training data appears unbiased is insufficient to ensure fair machine learning. Chowdhury Mohammad Rakin Haider, Chris Clifton, Yan Zhou 0001 |
ICDM | 2 |
| 2022 | Differentially Private k-Nearest Neighbor Missing Data ImputationabstractUsing techniques employing smooth sensitivity , we develop a method for \( k \) -nearest neighbor missing data imputation with differential privacy. This requires bounding the number of data incomplete tuples that can have their data complete “donor” changed by making a single addition or deletion to the dataset. The multiplicity of a single individual’s impact on an imputed dataset necessarily means our mechanisms require the addition of more noise than mechanisms that ignore missing data, but we show empirically that this is significantly outweighed by the bias reduction from imputing missing data. Chris Clifton, Eric J. Hanson, Keith Merrill, Shawn Merrill |
ACM Trans. Priv. Secur. | 1 |
| 2021 | Differentially Private Naïve Bayes Classifier Using Smooth SensitivityabstractThere is increasing awareness of the need to protect individual privacy in the training data used to develop machine learning models. Differential Privacy is a strong concept of protecting individuals. Naïve Bayes is a popular machine learning algorithm, used as a baseline for many tasks. In this work, we have provided a differentially private Naïve Bayes classifier that adds noise proportional to the smooth sensitivity of its parameters. We compare our results to Vaidya, Shafiq, Basu, and Hong [1] which scales noise to the global sensitivity of the parameters. Our experimental results on real-world datasets show that smooth sensitivity significantly improves accuracy while still guaranteeing ɛ-differential privacy. Farzad Zafarani, Chris Clifton |
Proc. Priv. Enhancing Technol. | 2 |
| 2020 | A Partitioned Recoding Scheme for Privacy Preserving Data Publishing
Chris Clifton, Eric J. Hanson, Keith Merrill, Shawn Merrill, Amjad Zahraa |
PSD | 1 |
| 2018 | Support vector classification with ℓ-diversity
Koray Mancuhan, Chris Clifton |
Comput. Secur. | 2 |
| 2017 | Statistical Learning Theory Approach for Data Classification with ℓ-diversityabstractCorporations are retaining ever-larger corpuses of personal data; the frequency or breaches and corresponding privacy impact have been rising accordingly. One way to mitigate this risk is through use of anonymized data, limiting the exposure of individual data to only where it is absolutely needed. This would seem particularly appropriate for data mining, where the goal is generalizable knowledge rather than data on specific individuals. In practice, corporate data miners often insist on original data, for fear that they might ”miss something” with anonymized or differentially private approaches. This paper provides a theoretical justification for the use of anonymized data. Specifically, we show that a support vector classifier trained on anatomized data satisfying ℓ-diversity should be expected to do as well as on the original data. Anatomy preserves all data values, but introduces uncertainty in the mapping between identifying and sensitive values, thus satisfying ℓ-diversity. The theoretical effectiveness of the proposed approach is validated using several publicly available datasets, showing that we outperform the state of the art for support vector classification using training data protected by k-anonymity, and are comparable to learning on the original data. Koray Mancuhan, Chris Clifton |
SDM | 2 |
| 2016 | Differentially Private Significance Testing on Paired-Sample DataabstractRigorous data mining results require measures of the statistical significance of the outcomes. The complexity of the data and models makes this a challenge; methods to protect privacy further complicate the issue. We demonstrate how to estimate statistical significance of results in the context of a social network analysis problem; the impact of the noise required to provide differential privacy is included in the significance measure. As a result, providing privacy does not complicate the use of the analysis. While demonstrated for social network analysis, the approach is general. The Wilcoxon signed-rank test used is appropriate for a wide variety of data with “before” and “after” measurements, and adapts well to differential privacy. We demonstrate on publicly available data with known privacy issues, showing that some apparently large differences are not significant, some small differences are, and that when the analysis is done using differential privacy, the same results can been achieved while protecting individual privacy. Christine Task, Chris Clifton |
SDM | 2 |
| 2015 | Laplace noise generation for two-party computational differential privacyabstractComputing a differentially private function using secure function evaluation prevents private information leakage both in the process, and from information present in the function output. However, the very secrecy provided by secure function evaluation poses new challenges if any of the parties are malicious. We first show how to build a two party differentially private secure protocol in the presence of malicious adversaries. We then relax the utility requirement of computational differential privacy to reduce computational cost, still giving security with rational adversaries. Finally, we provide a modified two-party computational differential privacy definition and show correctness and security guarantees in the rational setting. Balamurugan Anandan, Chris Clifton |
PST | 2 |
| 2015 | Anonymizing transactional datasetsabstractIn this paper, we study the privacy breach caused by unsafe correlations in transactional data where individuals have multiple tuples in a dataset. We provide two safety constraints to guarantee safe correlation of the data: (1) the safe grouping constraint to ensure that quasi-identifier and sensitive partitions are bounded by l-diversity and (2) the schema decomposition constraint to eliminate non-arbitrary correlations between non-sensitive and sensitive values to protect privacy and at the same time increase the aggregate analysis. In our technique, values are grouped together in unique partitions that enforce l-diversity at the level of individuals. We also propose an association preserving technique to increase the ability to learn/analyze from the anonymized data. To evaluate our approach, we conduct a set of experiments to determine the privacy breach and investigate the anonymization cost of safe grouping and preserving associations. Bechara al Bouna, Chris Clifton, Qutaibah M. Malluhi |
J. Comput. Secur. | 2 |
| 2014 | Privacy Beyond ConfidentialityabstractThe computer science community has had a growing research focus in Privacy over the last decade. Much of this has really focused on confidentiality: Anonymization, computing on encrypted data, access control policy, etc. This talk will look at a variety of research results in this area, including "weaker" approaches than the absolutes typically considered in the security community, and how they all come down to the same basic concept of providing confidentiality. Privacy is much more complex. People are often willing to allow use of their data -- but not just for anything. This talk will look at such other privacy issues, such as harm to individuals and society from the fear of disclosure or misuse of private data. The talk will conclude with ideas for new research directions in privacy. Chris Clifton |
CCS | 1 |
| 2014 | Top-k frequent itemsets via differentially private FP-treesabstractFrequent itemset mining is a core data mining task and has been studied extensively. Although by their nature, frequent itemsets are aggregates over many individuals and would not seem to pose a privacy threat, an attacker with strong background information can learn private individual information from frequent itemsets. This has lead to differentially private frequent itemset mining, which protects privacy by giving inexact answers. We give an approach that first identifies top-k frequent itemsets, then uses them to construct a compact, differentially private FP-tree. Once the noisy FP-tree is built, the (privatized) support of all frequent itemsets can be derived from it without access to the original data. Experimental results show that the proposed algorithm gives substantially better results than prior approaches, especially for high levels of privacy. Chris Clifton |
KDD | 2 |
| 2013 | Using Safety Constraint for Transactional Dataset Anonymization
Bechara al Bouna, Chris Clifton, Qutaibah M. Malluhi |
DBSec | 2 |
| 2013 | Updating outsourced anatomized private databasesabstractWe introduce operations to safely update an anatomized database. The result is a database where the view of the server satisfies standards such as k-anonymity or l-diversity, but the client is able to query and modify the original data. By exposing data where possible, the server can perform value-added services such as data analysis not possible with fully encrypted data, while still being unable to violate privacy constraints. Update is a key challenge with this model; naïve application of insertion and deletion operations reveals the actual data to the server. This paper shows how data can be safely inserted, deleted, and updated. The key ideas are that data is inserted or updated into an encrypted temporary table until enough data is available to safely decrypt, and that sensitive information of deleted tuples is left behind to ensure privacy of both deleted and undeleted individuals. This approach is proven effective in maintaining the privacy constraint against an adversarial server. The paper also gives empirical results on how much data remains encrypted, and the resulting quality of the server's (anatomized) view of the data, for various update and delete rates. Ahmet Erhan Nergiz, Chris Clifton, Qutaibah M. Malluhi |
EDBT | 2 |
| 2013 | Privacy through Uncertainty in Location-Based ServicesabstractLocation-Based Services (LBS) are becoming more prevalent. While there are many benefits, there are also real privacy risks. People are unwilling to give up the benefits - but can we reduce privacy risks without giving up on LBS entirely? This paper explores the possibility of introducing uncertainty into location information when using an LBS, so as to reduce privacy risk while maintaining good quality of service. This paper also explores the current uses of uncertainty information in a selection of mobile applications. Shawn Merrill, Nilgun Basalp, Joachim Biskup, Erik Buchmann, Chris Clifton, Bart Kuijpers, Walied Othman, Erkay Savas |
MDM (2) | 5 |
| 2012 | A Guide to Differential Privacy Theory in Social Network AnalysisabstractPrivacy of social network data is a growing concern which threatens to limit access to this valuable data source. Analysis of the graph structure of social networks can provide valuable information for revenue generation and social science research, but unfortunately, ensuring this analysis does not violate individual privacy is difficult. Simply anonymizing graphs or even releasing only aggregate results of analysis may not provide sufficient protection. Differential privacy is an alternative privacy model, popular in data-mining over tabular data, which uses noise to obscure individuals' contributions to aggregate results and offers a very strong mathematical guarantee that individuals' presence in the data-set is hidden. Analyses that were previously vulnerable to identification of individuals and extraction of private data may be safely released under differential-privacy guarantees. We review two existing standards for adapting differential privacy to network data and analyse the feasibility of several common social-network analysis techniques under these standards. Additionally, we propose out-link privacy, a novel standard for differential privacy over network data, and introduce two powerful out-link private algorithms for common network analysis techniques that were infeasible to privatize under previous differential privacy standards. Christine Task, Chris Clifton |
ASONAM | 2 |
| 2012 | Differential identifiabilityabstractA key challenge in privacy-preserving data mining is ensuring that a data mining result does not inherently violate privacy. ε-Differential Privacy appears to provide a solution to this problem. However, there are no clear guidelines on how to set ε to satisfy a privacy policy. We give an alternate formulation, Differential Identifiability, parameterized by the probability of individual identification. This provides the strong privacy guarantees of differential privacy, while letting policy makers set parameters based on the established privacy concept of individual identifiability. Chris Clifton |
KDD | 2 |
| 2011 | Query Processing in Private Data Outsourcing Using Anonymization
Ahmet Erhan Nergiz, Chris Clifton |
DBSec | 2 |
| 2011 | How Much Is Enough? Choosing ε for Differential Privacy
Chris Clifton |
ISC | 2 |
| 2011 | Classifier evaluation and attribute selection against active adversariesabstractMany data mining applications, such as spam filtering and intrusion detection, are faced with active adversaries. In all these applications, the future data sets and the training data set are no longer from the same population, due to the transformations employed by the adversaries. Hence a main assumption for the existing classification techniques no longer holds and initially successful classifiers degrade easily. This becomes a game between the adversary and the data miner: The adversary modifies its strategy to avoid being detected by the current classifier; the data miner then updates its classifier based on the new threats. In this paper, we investigate the possibility of an equilibrium in this seemingly never ending game, where neither party has an incentive to change. Modifying the classifier causes too many false positives with too little increase in true positives; changes by the adversary decrease the utility of the false negative items that are not detected. We develop a game theoretic framework where equilibrium behavior of adversarial classification applications can be analyzed, and provide solutions for finding an equilibrium point. A classifier’s equilibrium performance indicates its eventual success or failure. The data miner could then select attributes based on their equilibrium performance, and construct an effective classifier. A case study on online lending data demonstrates how to apply the proposed game theoretic framework to a real application. Murat Kantarcioglu, Bowei Xi, Chris Clifton |
Data Min. Knowl. Discov. | 3 |
| 2010 | Search-log anonymization and advertisement: are they mutually exclusive?abstractThe revenue of search-engine providers strongly depends on targeted advertisement. Targeted advertisement is becoming more reliant on personal data. This puts user privacy at risk. One way to improve privacy is to anonymize search logs, but this reduces usefulness for ad placement. Further, the usefulness depends on the target function used for the anonymization. This paper is the first to study this tradeoff systematically. We quantify the usefulness of an anonymized search log for advertisement purposes, by estimating outcomes such as the number of clicks on ads or the number of ad impressions possible after anonymization. A main result is that anonymized search logs are still useful for advertisement purposes, but the extent strongly depends on the target function. Thorben Burghardt, Klemens Böhm, Achim Guttmann, Chris Clifton |
CIKM | 4 |
| 2010 | d-Presence without Complete World KnowledgeabstractAdvances in information technology, and its use in research, are increasing both the need for anonymized data and the risks of poor anonymization. We presented a new privacy metric, δ-presence, that clearly links the quality of anonymization to the risk posed by inadequate anonymization. It was shown that existing anonymization techniques are inappropriate for situations where δ-presence is a good metric (specifically, where knowing an individual is in the database poses a privacy risk). This article addresses a practical problem with, extending to situations where the data anonymizer is not assumed to have complete world knowledge. The algorithms are evaluated in the context of a real-world scenario, demonstrating practical applicability of the approach. Mehmet Ercan Nergiz, Chris Clifton |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2010 | Efficient privacy-preserving similar document detection
Mummoorthy Murugesan, Wei Jiang 0026, Chris Clifton, Luo Si, Jaideep Vaidya |
VLDB J. | 3 |
| 2009 | Assured Information Sharing Life CycleabstractThis paper describes our approach to assured information sharing. The research is being carried out under a MURI 9Multiuniversiyt Research Initiative) project funded by the Air Force Office of Scientific Research (AFOSR). The main objective of our project is: define, design and develop an Assured Information Sharing Lifecycle (AISL) that realizes the DoD's information sharing value chain. In this paper we describe the problem faced by the Department of Defense and our solution to developing an AISL System. Tim Finin, Anupam Joshi, Hillol Kargupta, Yelena Yesha, Joel Sachs, Elisa Bertino, Ninghui Li 0001, Chris Clifton, Eugene H. Spafford, Bhavani Thuraisingham, Murat Kantarcioglu, Alain Bensoussan 0001, Nathan Berg, Latifur Khan, Jiawei Han 0001, ChengXiang Zhai, Ravi S. Sandhu, Shouhuai Xu, Jim Massaro, Lada A. Adamic |
ISI | 8 |
| 2009 | Providing Privacy through Plausibly Deniable SearchabstractQuery-based web search is an integral part of many people's daily activities. Most do not realize that their search history can be used to identify them (and their interests). In July 2006, AOL released an anonymized search query log of some 600K randomly selected users. While valuable as a research tool, the anonymization was insufficient: individuals were identified from the contents of the queries alone [2]. Government requests for such logs increases the concern. To address this problem, we propose a client-centered approach of plausibly deniable search. Each user query is substituted with a standard, closely-related query intended to fetch the desired results. In addition, a set of k-1 cover queries are issued; these have characteristics similar to the standard query but on unrelated topics. The system ensures that any of these k queries will produce the same set of k queries, giving k possible topics the user could have been searching for. We use a Latent Semantic Indexing (LSI) based approach to generate queries, and evaluate on the DMOZ [10] webpage collection to show effectiveness of the proposed approach. Mummoorthy Murugesan, Chris Clifton |
SDM | 2 |
| 2009 | Multirelational k-Anonymityabstractk-Anonymity protects privacy by ensuring that data cannot be linked to a single individual. In a k-anonymous data set, any identifying information occurs in at least k tuples. Much research has been done to modify a single-table data set to satisfy anonymity constraints. This paper extends the definitions of k-anonymity to multiple relations and shows that previously proposed methodologies either fail to protect privacy or overly reduce the utility of the data in a multiple relation setting. We also propose two new clustering algorithms to achieve multirelational anonymity. Experiments show the effectiveness of the approach in terms of utility and efficiency. Mehmet Ercan Nergiz, Chris Clifton, Ahmet Erhan Nergiz |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2009 | Privacy-Preserving Kth Element Score over Vertically Partitioned DataabstractGiven a large integer data set shared vertically by two parties, we consider the problem of securely computing a score separating the kth and the (k + 1) to compute such a score while revealing little additional information. The proposed protocol is implemented using the Fairplay system and experimental results are reported. We show a real application of this protocol as a component used in the secure processing of top-k queries over vertically partitioned data. Jaideep Vaidya, Chris Clifton |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2008 | Collaborative Search and User Privacy: How Can They Be Reconciled?
Thorben Burghardt, Erik Buchmann, Klemens Böhm, Chris Clifton |
CollaborateCom | 4 |
| 2008 | Similar Document Detection with Limited Information DisclosureabstractSimilar document detection plays important roles in many applications, such as file management, copyright protection, and plagiarism prevention. Existing protocols assume that the contents of files stored on a server (or multiple servers) are directly accessible. This assumption limits more practical applications, e.g., detecting plagiarized documents between two conferences, where submissions are confidential. We propose novel protocols to detect similar documents between two entities where documents cannot be openly shared with each other. We also conduct experiments to show the practical value of the proposed protocols. Wei Jiang 0026, Mummoorthy Murugesan, Chris Clifton, Luo Si |
ICDE | 3 |
| 2008 | Transforming semi-honest protocols to ensure accountability
Wei Jiang 0026, Chris Clifton, Murat Kantarcioglu |
Data Knowl. Eng. | 2 |
| 2008 | Privacy-preserving decision trees over vertically partitioned dataabstractPrivacy and security concerns can prevent sharing of data, derailing data-mining projects. Distributed knowledge discovery, if done correctly, can alleviate this problem. We introduce a generalized privacy-preserving variant of the ID3 algorithm for vertically partitioned data distributed over two or more parties. Along with a proof of security, we discuss what would be necessary to make the protocols completely secure. We also provide experimental results, giving a first demonstration of the practical complexity of secure multiparty computation-based data mining. Jaideep Vaidya, Chris Clifton, Murat Kantarcioglu, A. Scott Patterson |
ACM Trans. Knowl. Discov. Data | 2 |
| 2008 | Privacy-preserving Naïve Bayes classification
Jaideep Vaidya, Murat Kantarcioglu, Chris Clifton |
VLDB J. | 3 |
| 2007 | Identifying Rare Classes with Sparse Training Data
Mingwu Zhang, Wei Jiang 0026, Chris Clifton, Sunil Prabhakar 0001 |
DEXA | 3 |
| 2007 | MultiRelational k-Anonymityabstractk-anonymity protects privacy by ensuring that data cannot be linked to a single individual. In a k-anonymous dataset, any identifying information occurs in at least k tuples. Much research has been done to modify a single table dataset to satisfy anonymity constraints. This paper extends the definitions of k-anonymity to multiple relations and shows that previously proposed methodologies either fail to protect privacy, or overly reduce the utility of the data, in a multiple relation setting. A new clustering algorithm is proposed to achieve multirelational anonymity. Mehmet Ercan Nergiz, Chris Clifton, Ahmet Erhan Nergiz |
ICDE | 2 |
| 2007 | AC-Framework for Privacy-Preserving CollaborationabstractThe secure multi-party computation (SMC) model provides means for balancing the use and confidentiality of distributed data. Increasing security concerns have led to a surge in work on practical secure multi-party computation protocols. However, most are only proven secure under the semi-honest model, and security under this adversary model is insufficient for most applications in the field of privacy-preserving data mining. In this paper, we present the full spectrum of the accountable computing (AC) framework, which is sufficient or practical for many applications without the complexity and cost of an SMC-protocol under the malicious model. Furthermore, to show the applicability of the AC-framework, we present an application under this framework regarding privacy-preserving mining frequent itemsets. Wei Jiang 0026, Chris Clifton |
SDM | 2 |
| 2007 | Hiding the presence of individuals from shared databasesabstractAdvances in information technology, and its use in research, are increasing both the need for anonymized data and the risks of poor anonymization. We present a metric, δ-presence, that clearly links the quality of anonymization to the risk posed by inadequate anonymization. We show that existing anonymization techniques are inappropriate for situations where δ-presence is a good metric (specifically, where knowing an individual is in the database poses a privacy risk), and present algorithms for effectively anonymizing to meet δ-presence. The algorithms are evaluated in the context of a real-world scenario, demonstrating practical applicability of the approach. Mehmet Ercan Nergiz, Maurizio Atzori, Chris Clifton |
SIGMOD Conference | 3 |
| 2007 | Thoughts on k-anonymization
Mehmet Ercan Nergiz, Chris Clifton |
Data Knowl. Eng. | 2 |
| 2007 | TKDE Guidelines for Survey PapersabstractProvides instructions and guidelines to prospective authors who wish to submit manuscripts. Chris Clifton, Xindong Wu 0001, Christos Faloutsos |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2006 | A secure distributed framework for achieving k-anonymity
Wei Jiang 0026, Chris Clifton |
VLDB J. | 2 |
| 2005 | Privacy-Preserving Distributed k-Anonymity
Wei Jiang 0026, Chris Clifton |
DBSec | 2 |
| 2005 | Security Issues in Querying Encrypted Data
Murat Kantarcioglu, Chris Clifton |
DBSec | 2 |
| 2005 | Privacy-Preserving Decision Trees over Vertically Partitioned Data
Jaideep Vaidya, Chris Clifton |
DBSec | 2 |
| 2005 | Knowledge Discovery from Transportation Network DataabstractTransportation and logistics are a major sector of the economy, however data analysis in this domain has remained largely in the province of optimization. The potential of data mining and knowledge discovery techniques is largely untapped. Transportation networks are naturally represented as graphs. This paper explores the problems in mining of transportation network graphs: we hope to find how current techniques both succeed and fail on this problem, and from the failures, we hope to present new challenges for data mining. Experimental results from applying both existing graph mining and conventional data mining techniques to real transportation network data are provided, including new approaches to making these techniques applicable to the problems. Reasons why these techniques are not appropriate are discussed. We also suggest several challenging problems to precipitate research and galvanize future work in this area. Wei Jiang 0026, Jaideep Vaidya, Zahir Balaporia, Chris Clifton, Brett Banich |
ICDE | 4 |
| 2005 | Privacy-Preserving Top-K QueriesabstractThe primary contribution of this paper is a secure method for doing top-k selection from vertically partitioned data. This has particular relevance to privacy-sensitive searches, and meshes well with privacy policies such as k-anonymity. We have demonstrated how secure primitives from the literature can be composed with efficient query processing algorithms, with the result having provable security properties. The paper also shows a trade-off between efficiency and disclosure. It is worth exploring whether one could have a suite of algorithms to optimize these tradeoffs, e.g., algorithms that guarantee k-anonymity with efficiency based on the choice of k rather than the guarantees of secure multiparty computation. Jaideep Vaidya, Chris Clifton |
ICDE | 2 |
| 2005 | Dependable Real-Time Data MiningabstractIn this paper we discuss the need for real-time data mining for many applications in government and industry and describe resulting research issues. We also discuss dependability issues including incorporating security, integrity, timeliness and fault tolerance into data mining. Several different data mining outcomes are described with regard to their implementation in a real-time environment. These outcomes include clustering, association-rule mining, link analysis and anomaly detection. The paper describes how they would be used together in various parallel-processing architectures. Stream mining is discussed with respect to the challenges of performing data mining on stream data from sensors. The paper concludes with a summary and discussion of directions in this emerging area. Bhavani Thuraisingham, Latifur Khan, Chris Clifton, John A. Maurer, Marion G. Ceruti |
ISORC | 3 |
| 2005 | Secure set intersection cardinality with application to association rule miningabstractThere has been concern over the apparent conflict between privacy and data mining. There is no inherent conflict, as most types of data mining produce summary results that do not reveal information about individuals. The process of data mining may use private data, leading to the potential for priv acy breaches. Secure Multiparty Computation shows that results can be produced without revealing the data used to generate them. The problem is that general techniques for secure multiparty computation do not scale to data-mining size computations. This paper presents an efficient protocol for securely determining the size of set intersection, and shows how this can be used to generate association rules where multiple parties have different (and private) information about the same set of individuals. Jaideep Vaidya, Chris Clifton |
J. Comput. Secur. | 2 |
| 2005 | Privacy-preserving clustering with distributed EM mixture modeling
Xiaodong Lin 0004, Chris Clifton, Michael Yu Zhu |
Knowl. Inf. Syst. | 2 |
| 2004 | Privacy-Preserving Outlier DetectionabstractOutlier detection can lead to the discovery of truly unexpected knowledge in many areas such as electronic commerce, credit card fraud and especially national security. We look at the problem of finding outliers in large distributed databases where privacy/security concerns restrict the sharing of data. Both homogeneous and heterogeneous distribution of data is considered. We propose techniques to detect outliers in such scenarios while giving formal guarantees on the amount of information disclosed. Jaideep Vaidya, Chris Clifton |
ICDM | 2 |
| 2004 | When do data mining results violate privacy?abstractPrivacy-preserving data mining has concentrated on obtaining valid results when the input data is private. An extreme example is Secure Multiparty Computation-based methods, where only the results are revealed. However, this still leaves a potential privacy breach: Do the results themselves violate privacy? This paper explores this issue, developing a framework under which this question can be addressed. Metrics are proposed, along with analysis that those metrics are consistent in the face of apparent problems. Murat Kantarcioglu, Jiashun Jin, Chris Clifton |
KDD | 3 |
| 2004 | Privately Computing a Distributed k-nn Classifier
Murat Kantarcioglu, Chris Clifton |
PKDD | 2 |
| 2004 | Privacy Preserving Naïve Bayes Classifier for Vertically Partitioned DataabstractPrivacy-Preserving Data Mining -- developing models without seeing the data -- is receiving growing attention. This paper assumes a privacy-preserving distributed data mining scenario: data sources collaborate to develop a global model, but must not disclose their data to others. Nave Bayes is often used as a baseline classifier, consistently providing reasonable classification performance. This paper brings privacy-preservation to Nave Bayes classification on vertically partitioned data. Jaideep Vaidya, Chris Clifton |
SDM | 2 |
| 2004 | TopCat: Data Mining for Topic Identification in a Text CorpusabstractTopCat (topic categories) is a technique for identifying topics that recur in articles in a text corpus. Natural language processing techniques are used to identify key entities in individual articles, allowing us to represent an article as a set of items. This allows us to view the problem in a database/data mining context: Identifying related groups of items. We present a novel method for identifying related items based on traditional data mining techniques. Frequent itemsets are generated from the groups of items, followed by clusters formed with a hypergraph partitioning scheme. We present an evaluation against a manually categorized ground truth news corpus; it shows this technique is effective in identifying topics in collections of news articles. Chris Clifton, Robert Cooley, Jason Rennie |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2004 | Privacy-Preserving Distributed Mining of Association Rules on Horizontally Partitioned DataabstractData mining can extract important knowledge from large data collections ut sometimes these collections are split among various parties. Privacy concerns may prevent the parties from directly sharing the data and some types of information about the data. We address secure mining of association rules over horizontally partitioned data. The methods incorporate cryptographic techniques to minimize the information shared, while adding little overhead to the mining task. Murat Kantarcioglu, Chris Clifton |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2003 | Privacy-preserving k-means clustering over vertically partitioned dataabstractPrivacy and security concerns can prevent sharing of data, derailing data mining projects. Distributed knowledge discovery, if done correctly, can alleviate this problem. The key is to obtain valid results, while providing guarantees on the (non)disclosure of data. We present a method for k-means clustering when different sites contain different attributes for a common set of entities. Each site learns the cluster of each entity, but learns nothing about the attributes at other sites. Jaideep Vaidya, Chris Clifton |
KDD | 2 |
| 2003 | Privacy-Enhanced Data Management for Next-Generation e-Commerce
Chris Clifton, Irini Fundulaki, Richard Hull 0001, Bharat Kumar, Daniel F. Lieuwen, Arnaud Sahuguet |
VLDB | 1 |
| 2003 | Change Detection in Overhead Imagery Using Neural Networks
Chris Clifton |
Appl. Intell. | 1 |
| 2002 | Privacy preserving association rule mining in vertically partitioned dataabstractPrivacy considerations often constrain data mining projects. This paper addresses the problem of association rule mining where transactions are distributed across sources. Each site holds some attributes of each transaction, and the sites wish to collaborate to identify globally valid association rules. However, the sites must not reveal individual transaction data. We present a two-party algorithm for efficiently discovering frequent itemsets with minimum support levels, without either site revealing individual transaction values. Jaideep Vaidya, Chris Clifton |
KDD | 2 |
| 2001 | Real-Time Data Mining of Multimedia ObjectsabstractWhereas much of the previous work on data mining has focused on mining data in relational databases, we discuss mining objects. Object models are very popular for representing multimedia data, and therefore we need to mine object databases to extract useful information from the large quantities of multimedia data. We first describe the motivation for multimedia data mining with examples and then discuss object mining with focus on text, image, video and audio mining. We also address the need for real time data mining for multimedia applications. Bhavani Thuraisingham, Chris Clifton, John A. Maurer, Marion G. Ceruti |
ISORC | 2 |
| 2000 | SEMINT: A tool for identifying attribute correspondences in heterogeneous databases using neural networks
Wen-Syan Li, Chris Clifton |
Data Knowl. Eng. | 2 |
| 2000 | Using Sample Size to Limit Exposure to Data MiningabstractData mining introduces new problems in database security. The basic problem of using non-sensitive data to infer sensitive data is made more difficult by the “probabilistic” inferences possible with data mining. This paper shows how lower bounds from Chris Clifton |
J. Comput. Secur. | 1 |
| 2000 | Database Integration Using Neural Networks: Implementation and Experiences
Wen-Syan Li, Chris Clifton, Shu-Yao Liu |
Knowl. Inf. Syst. | 2 |
| 1999 | Protecting Against Data Mining through Samples
Chris Clifton |
DBSec | 1 |
| 1999 | Panel on Intrusion Detection
Ming-Yuh Huang, D. Shayne Pitcock, Chris Clifton, Tsau Young Lin |
DBSec | 4 |
| 1999 | TopCat: Data Mining for Topic Identification in a Text Corpus
Chris Clifton, Robert Cooley |
PKDD | 1 |
| 1998 | Data Mining on TextabstractData mining technology is giving us the ability to extract meaningful patterns from large quantities of structured data. Information retrieval systems have made large quantities of textual data available. Extracting meaningful patterns from this data is difficult. Current tools for mining structured data are inappropriate for free text. We outline problems involved in Knowledge Discovery in Text, and present an architecture for extracting patterns that hold across multiple documents. The capabilities that such a system could provide are illustrated. Chris Clifton, Rick Steinheiser |
COMPSAC | 1 |
| 1998 | Query Flocks: A Generalization of Association-Rule MiningabstractAssociation-rule mining has proved a highly successful technique for extracting useful information from very large databases. This success is attributed not only to the appropriateness of the objectives, but to the fact that a number of new query-optimization ideas, such as the “a-priori” trick, make association-rule mining run much faster than might be expected. In this paper we see that the same tricks can be extended to a much more general context, allowing efficient mining of very large databases for many different kinds of patterns. The general idea, called “query flocks,” is a generate-and-test model for data-mining problems. We show how the idea can be used either in a general-purpose mining system or in a next generation of conventional query optimizers. Shalom Tsur, Jeffrey D. Ullman, Serge Abiteboul, Chris Clifton, Rajeev Motwani 0001, Svetlozar Nestorov, Arnon Rosenthal |
SIGMOD Conference | 4 |
| 1998 | Multidatabase Query Processing with Uncertainty in Global Keys and Attribute ValuesabstractSemantic integration and data integration are two main processes that multidatabase systems need to employ in order to support interoperability. Both these processes involve uncertainty when attribute correspondences and global IDs are unknown or imprecise. The role-set approach is a new conceptual framework for data integration in multidatabase systems that maintains the materialization autonomy of local database systems by presenting the answer to a query as a set of sets representing the distinct intersections between the relations corresponding to the various roles played by an entity. In this article, we present an approach for dynamic database integration and query processing in the absence of information about attribute correspondences and global IDs. We define different types of equivalence conditions for the construction of global IDs. We propose a strategy based on ranked role-sets that makes use of an automated semantic integration procedure based on neural networks to determine candidate global IDs. The data integration and query processing steps then produce a number of role-sets, ranked by the similarity of the candidate IDs. © 1998 John Wiley & Sons, Inc. Peter Scheuermann, Wen-Syan Li, Chris Clifton |
J. Am. Soc. Inf. Sci. | 3 |
| 1997 | Security Issues in Data Warehousing and Data Mining: Panel Discussion
Bhavani Thuraisingham, Linda Schlipper, Pierangela Samarati, Tsau Young Lin, Sushil Jajodia, Chris Clifton |
DBSec | 6 |
| 1996 | Dynamic Integration and Query Processing with Ranked Role SetsabstractThe role-set approach is a new conceptual framework for data integration in multidatabase systems that maintains the materialization autonomy of local database systems and provides users with more accurate information. The role-set approach presents the answer to a query as a set of relations where the distinct intersections between the relations correspond to the various roles played by an entity. The authors show how the basic role-based approach can be extended in the absence of information about the multidatabase keys (global IDs). They propose a strategy based on ranked role-sets that makes use of a semantic integration procedure based on neural networks to determine candidate global IDs. The data integration and query processing steps then produce a number of role-sets, ranked by the similarity of the candidate IDs. Peter Scheuermann, Wen-Syan Li, Chris Clifton |
CoopIS | 3 |
| 1995 | How are We Going to Pay for This? Fee-for-Service in Distributed Systems: Research and Policy Issues (Panel)abstractWith the increasing array of information and services being supported by distributed computing, we face a new challenge: How do we handle charges? Providers of information will want to receive "royalties", providers of computing services will want to receive payment for use of the service, and providers of the network will want to receive payment for the transmission. At some point, the users of the provided services will have to pay for this. What do we need to do in designing our distributed computing systems to support some form of charge/payment? This paper discusses these issues. Chris Clifton, Peter Gemmel, Ed Means, Matt J. Merges, J. D. Tygar |
ICDCS | 1 |
| 1995 | Semint: A System Prototype for Semantic Integration in Heterogeneous DatabasesabstractNo abstract available. Wen-Syan Li, Chris Clifton |
SIGMOD Conference | 2 |
| 1995 | HyperFile: A Data and Query Model for Documents
Chris Clifton, Hector Garcia-Molina, David Bloom |
VLDB J. | 1 |
| 1994 | Semantic Integration in Heterogeneous Databases Using Neural Networks
Wen-Syan Li, Chris Clifton |
VLDB | 2 |
| 1993 | Information Brokers: Sharing Knowledge in a Heterogeneous Distributed System
Daniel Barbará, Chris Clifton |
DEXA | 2 |
| 1993 | The Gold MailerabstractThe Gold Mailer, a system that provides users with an integrated way to send and receive messages using different media, efficiently store and retrieve these messages, and access a variety of sources of other useful information, is described. The mailer solves the problems of information overload, organization of messages and multiple interfaces. By providing good storage and retrieval facilities, it can be used as a powerful information processing engine covering a range of useful office information. The Gold Mailer's query language, indexing engine, file organization, data structures, and support of mail message data and multimedia documents are discussed.> Daniel Barbará, Chris Clifton, Fred Douglis, Hector Garcia-Molina, Ben Kao, Sharad Mehrotra, Jens Tellefsen, Rosemary Walsh |
ICDE | 2 |
| 1991 | Distributed processing of filtering queries in HyperFileabstractA language has been developed for queries which serves as an extension of the browsing model of hypertext systems. The query language and data model fit naturally into a distributed environment. A simple and efficient method is discussed for processing distributed queries in this language. Results of experiments run on a distributed data server using this algorithm are presented.> Chris Clifton, Hector Garcia-Molina |
ICDCS | 1 |
| 1990 | Indexing in a Hypertext Database
Chris Clifton, Hector Garcia-Molina |
VLDB | 1 |