EDBT 2026 Demo / reviewers in the wild / expert
Dinusha Vatsalan
dblp:89/7958
· DBLP profile ↗
34ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0001-6713-7667ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 25 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 2 since 2021Security and privacy · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Membership Inference Attack Vulnerabilities of Record Linkage ModelsabstractRecord linkage plays a crucial role in integrating health, legal, and administrative data. In domains such as healthcare, unstructured records, including clinical notes, contain rich information, making their integration valuable for applications like clinical trials. While deep learning models have improved linkage quality, their privacy risks remain under-explored. We present what is, to our knowledge, the first systematic study of membership inference attack vulnerabilities in record linkage models trained on de-identified texts. Unlike traditional classifiers, linkage models operate on record pairs, prompting a rethinking of what constitutes membership leakage. Does it occur if only a record pair was seen together during training or individually? Or if neither was seen, but the pair resembles known patterns? We introduce a black-box attack based on semantic perturbation sensitivity, requiring no access to the model's internals. Our findings expose a previously unaddressed membership inference vulnerability in record linkage models: black-box attacks, even with a simple threshold-based attack model, achieved up to 92% precision and AUC 0.79. Across record linkage models trained on MIMIC-IV and PMC-Patients datasets, we observe that perturbing training-seen phrases causes significantly larger confidence shifts (e.g., Δ s = -0.648) and higher change in predicted label (label flip from 0 = non-match to 1 = match and vice versa) (e.g., 68.75%) compared to unseen variants. These preliminary results reveal membership inference vulnerabilities in text-based linkage systems, highlighting the need for deeper investigation into privacy risks, motivating new lines of privacy defences for pairwise models. Piyumi Seneviratne, Dinusha Vatsalan, Mohamed Ali Kâafar |
CIKM | 2 |
| 2025 | "Do It to Know It": Reshaping the Privacy Mindset of Computer Science UndergraduatesabstractSoftware applications, while being an integral part of the modern world, pose significant threats to end-user privacy. Thus, computer professionals require knowledge and skills to develop privacy-aware software. However, undergraduate computing degree programs often lack privacy-focused curricula that can cultivate this ability in the future workforce. Therefore, we designed a privacy curriculum informed by the common challenges that computing professionals face when developing privacy-embedded software. It guides students in realising the need for privacy, identifying privacy protection mechanisms and programming Privacy Enhancing Technologies (PETs). We piloted the curriculum for third-year Computer Science undergraduates at the University of Auckland, New Zealand. The curriculum was evaluated using course assessments and surveys conducted before and after the lessons. Overall, the students improved their understanding of privacy, especially technical aspects. Most of them valued the applied learning experience of the programming lessons yet showed distinct views on task completion difficulty and motivation to do programming. Students recognised that privacy should be integral to their skill set by confirming the importance and relevance of the lessons. However, their perceived responsibility in privacy protection varied depending on their intention to take proactive measures. Based on the results, the paper suggests improvements to the proposed curriculum. Maisha Boteju, Danielle Lottridge, Thilina Ranbaduge, Dinusha Vatsalan, Ni Ding |
Proc. Priv. Enhancing Technol. | 4 |
| 2024 | Privacy Preserving Release of Mobile Sensor DataabstractSensors embedded in mobile smart devices can monitor users’ activity with high accuracy to provide a variety of services to end-users ranging from precise geolocation, health monitoring, and handwritten word recognition. However, this involves the risk of accessing and potentially disclosing sensitive information of individuals to the apps that may lead to privacy breaches. In this paper, we aim to minimize privacy leakages that may lead to user identification on mobile devices through user tracking and distinguishability while preserving the functionality of apps and services. We propose a privacy-preserving mechanism that effectively handles the sensor data fluctuations (e.g., inconsistent sensor readings while walking, sitting, and running at different times) by formulating the data as time-series modeling and forecasting. The proposed mechanism uses correlated noise-series against noise filtering attacks from an adversary, which aims to filter out the noise from the perturbed data to re-identify the original data. Unlike existing solutions, our mechanism keeps running in isolation without the interaction of a user or a service provider. We perform rigorous experiments on three benchmark datasets and show that our proposed mechanism limits user tracking and distinguishability threats to a significant extent compared to the original data while maintaining a reasonable level of utility of functionalities. In general, we show that our obfuscation mechanism reduces the user trackability threat by 60% across all the datasets while maintaining the utility loss below 0.3 Mean Absolute Error (MAE). More specifically, we observe that 80% of users achieve a 100% untrackability rate in the Swipes dataset across all noise scales. In the handwriting dataset, distinguishability is 17% for 60% of the users. Overall, our mechanism provides a utility error (MAE) of only 0.12 for 60% of users, and this increases to 0.2 for 100% users when correction thresholds are altered. Rahat Masood, Wing Yan Cheng, Dinusha Vatsalan, Deepak Mishra 0001, Hassan Jameel Asghar, Mohamed Ali Kâafar |
ARES | 3 |
| 2024 | On Adversarial Training with Incorrect Labels
Benjamin Zi Hao Zhao, Junda Lu 0001, Xiaowei Zhou 0003, Dinusha Vatsalan, Muhammad Ikram 0001, Mohamed Ali Kâafar |
WISE (4) | 4 |
| 2024 | Privacy-Preserving Deep Learning Based Record LinkageabstractDeep learning-based linkage of records across different databases is becoming increasingly useful in data integration and mining applications to discover new insights from multiple data sources. However, due to privacy and confidentiality concerns, organisations often are unwilling or allowed to share their sensitive data with any external parties, thus making it challenging to build/train deep learning models for record linkage across different organisations' databases. To overcome this limitation, we propose the first deep learning-based multi-party privacy-preserving record linkage (PPRL) protocol that can be used to link sensitive databases held by multiple different organisations. In our approach, each database owner first trains a local deep learning model, which is then uploaded to a secure environment and securely aggregated to create a global model. The global model is then used by a linkage unit to distinguish unlabelled record pairs as matches and non-matches. We utilise differential privacy to achieve provable privacy protection against re-identification attacks. We evaluate the linkage quality and scalability of our approach using several large real-world databases, showing that it can achieve high linkage quality while providing sufficient privacy protection against existing attacks. Thilina Ranbaduge, Dinusha Vatsalan, Ming Ding 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Privacy-Preserving Record Linkage for Cardinality CountingabstractSeveral applications require counting the number of distinct items in the data, which is known as the cardinality counting problem. Example applications include health applications such as rare disease patients counting for adequate awareness and funding, and counting the number of cases of a new disease for outbreak detection, marketing applications such as counting the visibility reached for a new product, and cybersecurity applications such as tracking the number of unique views of social media posts. The data needed for the counting is however often personal and sensitive, and need to be processed using privacy-preserving techniques. The quality of data in different databases, for example typos, errors and variations, poses additional challenges for accurate cardinality estimation. While privacy-preserving cardinality counting has gained much attention in the recent times and a few privacy-preserving algorithms have been developed for cardinality estimation, no work has so far been done on privacy-preserving cardinality counting using record linkage techniques with fuzzy matching and provable privacy guarantees. We propose a novel privacy-preserving record linkage algorithm using unsupervised clustering techniques to link and count the cardinality of individuals in multiple datasets without compromising their privacy or identity. In addition, existing Elbow methods to find the optimal number of clusters as the cardinality are far from accurate as they do not take into account the purity and completeness of generated clusters. We propose a novel method to find the optimal number of clusters in unsupervised learning. Our experimental results on real and synthetic datasets are highly promising in terms of significantly smaller error rate of less than 0.1 with a privacy budget ϵ = 1.0 compared to the state-of-the-art fuzzy matching and clustering method. Nan Wu 0013, Dinusha Vatsalan, Mohamed Ali Kâafar, Sanath Kumar Ramesh |
AsiaCCS | 2 |
| 2023 | Local Differentially Private Fuzzy Counting in Stream Data Using Probabilistic Data StructuresabstractPrivacy-preserving estimation of counts of items in streaming data finds applications in several real-world scenarios including word auto-correction and traffic management applications. Recent works of RAPPOR [1] and Apple's count-mean sketch (CMS) algorithm [2] propose privacy preserving mechanisms for count estimation in large volumes of data using probabilistic data structures like counting Bloom filter and CMS. However, these existing methods fall short in providing a sound solution for real-time streaming data applications. Since the size of the data structure in these methods is not adaptive to the volume of the streaming data, the utility (accuracy of the count estimate) can suffer over time due to increased false positive rates. Further, the lookup operation needs to be highly efficient to answer count estimate queries in real-time. More importantly, the local Differential privacy mechanisms used in these approaches to provide privacy guarantees come at a large cost to utility (impacting the accuracy of count estimation). In this work, we propose a novel (local) Differentially private mechanism that provides high utility for the streaming data count estimation problem with similar or even lower privacy budgets while providing: a) fuzzy counting to report counts of related or similar items (for instance to account for typing errors and data variations), and b) improved querying efficiency to reduce the response time for real-time querying of counts. Our algorithm uses a combination of two probabilistic data structures Cuckoo filter and Bloom filter. We provide formal proofs for privacy and utility guarantees and present extensive experimental evaluation of our algorithm using real and synthetic English words datasets for both the exact and fuzzy counting scenarios. Our privacy preserving mechanism substantially outperforms the prior work in terms of lower querying time, significantly higher utility (accuracy of count estimation) under similar or lower privacy guarantees, at the cost of communication overhead. Dinusha Vatsalan, Raghav Bhaskar, Mohamed Ali Kâafar |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Fairness and Cost Constrained Privacy-Aware Record LinkageabstractRecord linkage algorithms match and link records from different databases that refer to the same real-world entity based on direct and/or quasi-identifiers, such as name, address, age, and gender, available in the records. Since these identifiers generally contain personal identifiable information (PII) about the entities, record linkage algorithms need to be developed with privacy constraints. Known as privacy-preserving record linkage (PPRL), many research studies have been conducted to perform the linkage on encoded and/or encrypted identifiers. Differential privacy (DP) combined with computationally efficient encoding methods, e.g. Bloom filter encoding, has been used to develop PPRL with provable privacy guarantees. The standard DP notion does not however address other constraints, among which the most important ones are fairness-bias and cost of linkage in terms of number of record pairs to be compared. In this work, we propose new notions of fairness-constrained DP and fairness and cost-constrained DP for PPRL and develop a framework for PPRL with these new notions of DP combined with Bloom filter encoding. We provide theoretical proofs for the new DP notions for fairness and cost-constrained PPRL and experimentally evaluate them on two datasets containing person-specific data. Our experimental results show that with these new notions of DP, PPRL with better performance (compared to the standard DP notion for PPRL) can be achieved with regard to privacy, cost and fairness constraints. Nan Wu 0013, Dinusha Vatsalan, Sunny Verma, Mohamed Ali Kâafar |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Privacy Preserving Text Data Encoding and Topic ModellingabstractTextual data, such as clinical notes, product or movie reviews in online stores, transcripts, chat records, and business documents, are widely collected nowadays and can be used to support a large spectrum of Big Data applications. At the same time, textual data, collected about individuals or from individuals, can be susceptible to inference attacks that may leak private and/or sensitive information about individuals.The increasing concerns of privacy risks in textual data preclude sharing or exchanging textual data across different parties/organizations for various applications such as record linkage, similar entity matching, natural language processing (NLP), or machine learning on large collections of textual data. This has led to the development of privacy preserving techniques for applying matching, machine learning or NLP techniques on textual data that contain personal and sensitive information about individuals. While cryptographic techniques are highly secure and accurate, they incur significant amount of computational cost for encoding and matching data – especially textual data – due to the complex nature of text.In this paper, we propose an efficient textual data encoding and matching algorithm using probabilistic techniques based on counting Bloom filters combined with Differential privacy. We apply our algorithm to a popular use case scenario that involves privacy preserving topic modeling – a widely used NLP technique – in order to identify common or collective topics in texts across multiple parties without learning the individual topics of each party, and show its effectiveness in supporting this application. Finally, through extensive experimental evaluation on three large text datasets against a state-of-the-art probabilistic encoding algorithm for privacy preserving LDA topic modelling, we show that our method provides a better privacy-utility trade-off at the cost of more computation complexity and memory space, while still being computationally efficient (log-linear complexity in the size of documents) for Big data compared to cryptographic techniques that have quadratic complexity. Dinusha Vatsalan, Raghav Bhaskar, Aris Gkoulalas-Divanis, Dimitrios Karapiperis |
IEEE BigData | 1 |
| 2021 | Modern Privacy-Preserving Record Linkage Techniques: An OverviewabstractRecord linkage is the challenging task of deciding which records, coming from disparate data sources, refer to the same entity. Established back in 1946 by Halbert L. Dunn [1], the area of record linkage has received tremendous attention over the years due to its numerous real-world applications, and has led to a plethora of technologies, methods, metrics, and systems. A major direction in record linkage regards methods for linking records in a privacy-preserving manner, where sensitive and personally identifiable information in the records is not leaked as part of the linkage process. In this article, we provide an overview of the large body of research literature in privacy-preserving record linkage, discuss the different generations of techniques that have been proposed, their advantages and limitations, and present a taxonomy as well as an extensive survey on the latest generation of methods. We conclude this work with a roadmap to the new generation of analytics-driven techniques that aims to address some of the major challenges in the field. Aris Gkoulalas-Divanis, Dinusha Vatsalan, Dimitrios Karapiperis, Murat Kantarcioglu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Privacy-Preserving Techniques for Protecting Large-Scale Data of Cyber-Physical SystemsabstractAs Cyber-Physical Systems (CPSs), such as power and gas networks, generate heterogeneous and large-scale data sources from devices and networks, they need efficient privacy-preserving techniques to protect data and systems from cyber attacks. To safeguard CPSs from potential cyber threats, it is vital to identify vulnerabilities of CPSs' components to prevent Advanced Persistent Threats (APTs) and protect their generated data using privacy-preserving techniques. This paper aims to review the current state of privacy-preserving techniques for protecting CPSs and their networks against cyber attacks. Concepts of Privacy preservation and CPSs are discussed, illustrating CPSs' components and how they could be hacked using cyber and physical hacking scenarios. Then, types of privacy preservation, including perturbation, authentication, machine learning (ML), cryptography and blockchain, are discussed to demonstrate how they would be applied to protect the original data in CPSs and their networks. Finally, we explain existing challenges, solutions and future research directions of privacy preservation in CPSs. Marwa Keshk, Nour Moustafa, Elena Sitnikova, Benjamin P. Turnbull, Dinusha Vatsalan |
MSN | 5 |
| 2020 | Incremental clustering techniques for multi-party Privacy-Preserving Record Linkage
Dinusha Vatsalan, Peter Christen, Erhard Rahm |
Data Knowl. Eng. | 1 |
| 2020 | Sequence Data Matching and Beyond: New Privacy-Preserving Primitives Based on Bloom FiltersabstractBloom filter encoding has widely been used as an efficient masking technique for privacy-preserving matching functions. The existing matching techniques, however, are limited to relatively simple types such as string, categorical and signal numerical values. In this paper, we propose a new scheme that significantly extends the class of matching primitives that are based on privacy-preserving Bloom filter mechanism. These primitives include sequence data matching and popular distance-based machine learning algorithms such as KNN and SVM. Our scheme hash-maps a sequence data vector into the Bloom filter space while checking the similarity of the data points efficiently with negligible utility loss by adding a timestamp (bit) for each element in the data represented with its neighboring values. Furthermore, it includes a Laplace-like perturbation method on the constructed Bloom filters to address the weakness of deterministic probability led by encoding techniques. As a result, the proposed work guarantee the private data records are difficult to be discriminated due to collisions and differential privacy. The experimental results on three real-scenario based datasets illustrate that our method can achieve a significantly better trade-off between utility and privacy than the state-of-the-art differential privacy-based method by adding Laplace noise to the data directly. Wanli Xue, Dinusha Vatsalan, Wen Hu 0001, Aruna Seneviratne |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | A Privacy-Preserving-Framework-Based Blockchain and Deep Learning for Protecting Smart Power NetworksabstractModern power systems depend on cyber-physical systems to link physical devices and control technologies. A major concern in the implementation of smart power networks is to minimize the risk of data privacy violation (e.g., by adversaries using data poisoning and inference attacks). In this article, we propose a privacy-preserving framework to achieve both privacy and security in smart power networks. The framework includes two main modules: a two-level privacy module and an anomaly detection module. In the two-level privacy module, an enhanced-proof-of-work-technique-based blockchain is designed to verify data integrity and mitigate data poisoning attacks, and a variational autoencoder is simultaneously applied for transforming data into an encoded format for preventing inference attacks. In the anomaly detection module, a long short-term memory deep learning technique is used for training and validating the outputs of the two-level privacy module using two public datasets. The results highlight that the proposed framework can efficiently protect data of smart power networks and discover abnormal behaviors, in comparison to several state-of-the-art techniques. Marwa Keshk, Benjamin P. Turnbull, Nour Moustafa, Dinusha Vatsalan, Kim-Kwang Raymond Choo |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | Repairing of Record Linkage: Turning Errors into InsightabstractLinking records from different data sources, referred to as record linkage, is a longstanding but not yet satisfactorily resolved question in many fields of science. For practitioners, it is difficult to ensure the quality of linkage at the time of applying linkage techniques in real world applications. Instead, linkage errors are often detected later on, mostly by users of the applications. This not only requires us to repair errors, but also provides us with opportunities to observe the linkage quality and uncover why such errors occur. In viewing that record linkage is a complex and evolving process, we study how to acquire insights from linkage errors for achieving high-quality linkage. We propose a generic repairing framework which allows us to start with imperfect linkage models, and dynamically repair linkage models and errors for improved linkage quality. We have evaluated our repairing framework over three real-world datasets and the experimental results show that the performance of the proposed tree-structured classifier SVM-tree outperforms the baseline methods. Quyen Bui-Nguyen, Qing Wang 0002, Jingyu Shao, Dinusha Vatsalan |
EDBT | 4 |
| 2019 | Precise and Fast Cryptanalysis for Bloom Filter Based Privacy-Preserving Record LinkageabstractBeing able to identify records that correspond to the same entity across diverse databases is an increasingly important step in many data analytics projects. Research into privacy-preserving record linkage (PPRL) aims to develop techniques that can link records across databases such that besides the record pairs classified as matches no sensitive information about the entities in these databases is revealed. A popular technique used in PPRL is to encode sensitive values into Bloom filters (bit vectors), which has the advantage of allowing approximate matching using character q-grams. PPRL based on Bloom filter encoding has been shown to be accurate and scalable to large databases, and is thus now being used in real-world PPRL systems in Australia, Canada, and the UK. However, recent studies have shown that Bloom filters used for PPRL are vulnerable to cryptanalysis attacks that can re-identify some of the sensitive values encoded in these Bloom filters. While previous such attack methods were slow and required knowledge of various encoding parameters, we present a novel efficient attack which exploits how attribute values are encoded into Bloom filters. Our attack method does not require knowledge of the encoding function or its parameter settings used. It is able to correctly re-identify with high precision q-grams that could not have been hashed to certain Bloom filter bit positions, and using these re-identified q-grams it can then re-identify attribute values with high precision. Our method is significantly faster than earlier PPRL cryptanalysis attacks, and in our experimental evaluation, it is able to successfully re-identify attribute values from large real-world databases in a few minutes. Peter Christen, Thilina Ranbaduge, Dinusha Vatsalan, Rainer Schnell |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2018 | A Scalable and Efficient Subgroup Blocking Scheme for Multidatabase Record Linkage
Thilina Ranbaduge, Dinusha Vatsalan, Peter Christen |
PAKDD (3) | 2 |
| 2018 | Incognito: A Method for Obfuscating Web DataabstractUsers leave a trail of their personal data, interests, and intents while surfing or sharing information on the Web. Web data could therefore reveal some private/sensitive information about users based on inference analysis. The possible identification of information corresponding to a single individual by an inference attack holds true even if the user identifiers are encoded or removed in the Web data. Several works have been done on improving privacy of Web data through obfuscation methods~\citeHow09,Dom09,Sha05,Che14. However, these methods are neither comprehensive, generic to be applicable to any Web data, nor effective against adversarial attacks. To this end, we propose a privacy-aware obfuscation method for Web data addressing these identified drawbacks of existing methods. We use probabilistic methods to predict privacy risk of Web data that incorporates all key privacy aspects, which are uniqueness, uniformity, and linkability of Web data. The Web data with high predicted risk are then obfuscated by our method to minimize the privacy risk using semantically similar data. Our method is resistant against adversary who has knowledge about the datasets and model learned risk probabilities using differential privacy-based noise addition. Experimental study conducted on two real Web datasets validates the significance and efficacy of our method. Our results indicate that the average privacy risk reaches to 100% with a minimum of 10 sensitive Web entries, while at most 0% privacy risk could be attained with our obfuscation method at the cost of average utility loss of 64.3%. Rahat Masood, Dinusha Vatsalan, Muhammad Ikram 0001, Mohamed Ali Kâafar |
WWW | 2 |
| 2017 | Efficient Cryptanalysis of Bloom Filters for Privacy-Preserving Record Linkage
Peter Christen, Rainer Schnell, Dinusha Vatsalan, Thilina Ranbaduge |
PAKDD (1) | 3 |
| 2017 | Improving Temporal Record Linkage Using Regression Classification
Yichen Hu, Qing Wang 0002, Dinusha Vatsalan, Peter Christen |
PAKDD (1) | 3 |
| 2016 | Efficient Record Linkage Using a Compact Hamming SpaceabstractRecord linkage, the process of identifying similar records that correspond to the same real-world entities across databases, is a well-established research problem in the database, data mining, and information retrieval communities. Computing distances between string values of records is the key component in order to determine the similarity of the represented entities. Due to the typically large volumes of records, a two-step process is followed. A blocking mechanism is first applied for grouping similar records together, and then a matching mechanism is performed for comparing the records which have been inserted into the same block. However, there does not exist any efficient blocking/matching mechanism which provides theoretical guarantees for identifying similar records which consist of strings. Towards this end, we put forth the novel notion of embedding string-based records into a Hamming space, where such a mechanism exists. The size of these embeddings is kept as small as needed in order to guarantee the correspondence of distances in that space to the types of errors that exist between strings, e.g., a missing or a modified character. We build embeddings whose size is 120 bits for representing accurately four fields of a publicly available data set. We also present a distance threshold-aware blocking technique for higher accuracy rates compared to blocking approaches which ignore the specified threshold. Our empirical study conducted on real-world data sets shows the efficacy achieved by our embedding method as compared to several existing solutions. Dimitrios Karapiperis, Dinusha Vatsalan, Vassilios S. Verykios, Peter Christen |
EDBT | 2 |
| 2016 | Scalable Block Scheduling for Efficient Multi-database Record LinkageabstractRecord linkage (RL) is a task in data integration that aims to identify matching records that refer to the same entity from different databases. When records from more than two databases are to be linked RL is significantly challenged by the intrinsic exponential growth in the number of potential record comparisons to be conducted. We propose a scalable meta blocking protocol to be used for Multi-Database RL (MDRL) to significantly reduce the complexity of the matching (comparison and classification) phase. Our approach uses a graph structure to schedule the comparison of pairs of blocks with the aim of minimizing the number of repeated and superfluous comparisons between records. We provide an analysis of our approach and conduct an empirical study on large real-world databases. Thilina Ranbaduge, Dinusha Vatsalan, Peter Christen |
ICDM | 2 |
| 2016 | Hashing-Based Distributed Multi-party Blocking for Privacy-Preserving Record Linkage
Thilina Ranbaduge, Dinusha Vatsalan, Peter Christen, Vassilios S. Verykios |
PAKDD (2) | 2 |
| 2016 | Privacy-preserving matching of similar patients
Dinusha Vatsalan, Peter Christen |
J. Biomed. Informatics | 1 |
| 2015 | Large-Scale Multi-party Counting Set Intersection Using a Space Efficient Global Synopsis
Dimitrios Karapiperis, Dinusha Vatsalan, Vassilios S. Verykios, Peter Christen |
DASFAA (2) | 2 |
| 2015 | Efficient Entity Resolution with Adaptive and Interactive Training Data SelectionabstractEntity resolution (ER) is the task of deciding which records in one or more databases refer to the same real-world entities. A crucial step in ER is the accurate classification of pairs of records into matches and non-matches. In most practical ER applications, obtaining training data %of high quality is costly and time consuming. Various techniques have been proposed for ER to interactively generate training data and learn an accurate classifier. We propose an approach for training data selection for ER that exploits the cluster structure of the weight vectors (similarities) calculated from compared record pairs. Our approach adaptively selects an optimal number of informative training examples for manual labeling based on a user defined sampling error margin, and recursively splits the set of weight vectors to find pure enough subsets for training. We consider two aspects of ER that are highly significant in practice: a limited budget for the number of manual labeling that can be done, and a noisy oracle where manual labels might be incorrect. Experiments on four real public data sets show that our approach can significantly reduce manual labeling efforts for training an ER classifier while achieving matching quality comparative to fully supervised classifiers. Peter Christen, Dinusha Vatsalan, Qing Wang 0002 |
ICDM | 2 |
| 2015 | Clustering-Based Scalable Indexing for Multi-party Privacy-Preserving Record Linkage
Thilina Ranbaduge, Dinusha Vatsalan, Peter Christen |
PAKDD (2) | 2 |
| 2015 | Efficient Interactive Training Selection for Large-Scale Entity Resolution
Qing Wang 0002, Dinusha Vatsalan, Peter Christen |
PAKDD (2) | 2 |
| 2014 | Scalable Privacy-Preserving Record Linkage for Multiple DatabasesabstractPrivacy-preserving record linkage (PPRL) is the process of identifying records that correspond to the same real-world entities across several databases without revealing any sensitive information about these entities. Various techniques have been developed to tackle the problem of PPRL, with the majority of them only considering linking two databases. However, in many real-world applications data from more than two sources need to be linked. In this paper we consider the problem of linking data from three or more sources in an efficient and secure way. We propose a protocol that combines the use of Bloom filters, secure summation, and Dice coefficient similarity calculation with the aim to identify all records held by the different data sources that have a similarity above a certain threshold. Our protocol is secure in that no party learns any sensitive information about the other parties' data, but all parties learn which of their records have a high similarity with records held by the other parties. We evaluate our protocol on a large dataset showing the scalability, linkage quality, and privacy of our protocol. Dinusha Vatsalan, Peter Christen |
CIKM | 1 |
| 2013 | Flexible and extensible generation and corruption of personal dataabstractWith much of today's data being generated by people or referring to people, researchers increasingly require data that contain personal identifying information to evaluate their new algorithms. In areas such as record matching and de-duplication, fraud detection, cloud computing, and health informatics, issues such as data entry errors, typographical mistakes, noise, or recording variations, can all significantly affect the outcomes of data integration, processing, and mining projects. However, privacy concerns make it challenging to obtain real data that contain personal details. An alternative to using sensitive real data is to create synthetic data which follow similar characteristics. The advantages of synthetic data are that (1) they can be generated with well defined characteristics; (2) it is known which records represent an individual created entity (this is often unknown in real data); and (3) the generated data and the generator program itself can be published. We present a sophisticated data generation and corruption tool that allows the creation of various types of data, ranging from names and addresses, dates, social security and credit card numbers, to numerical values such as salary or blood pressure. Our tool can model dependencies between attributes, and it allows the corruption of values in various ways. We describe the overall architecture and main components of our tool, and illustrate how a user can easily extend this tool with novel functionalities. Peter Christen, Dinusha Vatsalan |
CIKM | 2 |
| 2013 | GeCo: an online personal data generator and corruptorabstractWe demonstrate GeCo, an online personal data GEnerator and COrruptor that facilitates the creation of realistic personal data ranging from names, addresses, and dates, to social security and credit card numbers, as well as numerical values such as salary or blood pressure. Using an intuitive Web interface, a user can create records containing such data according to their needs, and apply various corruption functions to generate duplicates of these records. Synthetic personal data are increasingly required in areas such as record de-duplication, fraud detection, cloud computing, and health informatics, where data quality issues can significantly affect the outcomes of data integration, processing, and mining projects. Privacy concerns, however, often make it difficult for researchers to obtain real data that contain personal details. Compared to other data generators that have to be downloaded, installed and customized,GeCo allows the creation of personal data with much less effort. In this demonstration we show (1) how different types of attributes, and dependencies between them, can be specified; (2) how the generated data can be modified using various types of corruption functions; and (3) how a user can contribute to GeCo by providing attribute generation functions and look-up files. We believe GeCo will be a valuable tool for researchers that require realistic personal data to evaluate their algorithms with regard to efficiency and effectiveness. Khoi-Nguyen Tran, Dinusha Vatsalan, Peter Christen |
CIKM | 2 |
| 2013 | Efficient two-party private blocking based on sorted nearest neighborhood clusteringabstractIntegrating data from diverse sources with the aim to identify similar records that refer to the same real-world entities without compromising privacy of these entities is an emerging research problem in various domains. This problem is known as privacy-preserving record linkage (PPRL). Scalability of PPRL is a main challenge due to growing data size in real-world applications. Private blocking techniques have been used in PPRL to address this challenge by reducing the number of record pair comparisons that need to be conducted. Many of these private blocking techniques require a trusted third party to perform the blocking. One main threat with three-party solutions is the collusion between parties to identify the private data of another party. Dinusha Vatsalan, Peter Christen, Vassilios S. Verykios |
CIKM | 1 |
| 2013 | Sorted Nearest Neighborhood Clustering for Efficient Private Blocking
Dinusha Vatsalan, Peter Christen |
PAKDD (2) | 1 |
| 2013 | A taxonomy of privacy-preserving record linkage techniques
Dinusha Vatsalan, Peter Christen, Vassilios S. Verykios |
Inf. Syst. | 1 |