EDBT 2026 Demo / reviewers in the wild / expert
Mark J. Elliot
dblp:82/6753
· DBLP profile ↗
13ranked-venue papers
5as first author
3since 2021 · last 2024
0000-0002-3142-4493ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 8 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multi-objective evolutionary GAN for tabular data synthesisabstractSynthetic data has a key role to play in data sharing by statistical agencies and other generators of statistical data products. Generative Adversarial Networks (GANs), typically applied to image synthesis, are also a promising method for tabular data synthesis. However, there are unique challenges in tabular data compared to images, eg tabular data may contain both continuous and discrete variables and conditional sampling, and, critically, the data should possess high utility and low disclosure risk (the risk of re-identifying a population unit or learning something new about them), providing an opportunity for multi-objective (MO) optimization. Inspired by MO GANs for images, this paper proposes a smart MO evolutionary conditional tabular GAN (SMOE-CTGAN). This approach models conditional synthetic data by applying conditional vectors in training, and uses concepts from MO optimisation to balance disclosure risk against utility. Our results indicate that SMOE-CTGAN is able to discover synthetic datasets with different risk and utility levels for multiple national census datasets. We also find a sweet spot in the early stage of training where a competitive utility and extremely low risk are achieved, by using an Improvement Score. The full code can be downloaded from github1. Nian Ran, Bahrul Ilmi Nasution, Claire Little, Richard Allmendinger 0001, Mark J. Elliot |
GECCO | 5 |
| 2024 | The Production of Bespoke Synthetic Teaching Datasets Without Access to the Original Data
Mark J. Elliot, Claire Little, Richard Allmendinger 0001 |
PSD | 1 |
| 2022 | Comparing the Utility and Disclosure Risk of Synthetic Data with Samples of Microdata
Claire Little, Mark J. Elliot, Richard Allmendinger 0001 |
PSD | 2 |
| 2018 | Breaking the Activation Function Bottleneck through Adaptive ParameterizationabstractStandard neural network architectures are non-linear only by virtue of a simple element-wise activation function, making them both brittle and excessively large. In this paper, we consider methods for making the feed-forward layer more flexible while preserving its basic structure. We develop simple drop-in replacements that learn to adapt their parameterization conditional on the input, thereby increasing statistical efficiency significantly. We present an adaptive LSTM that advances the state of the art for the Penn Treebank and Wikitext-2 word-modeling tasks while using fewer parameters and converging in half as many iterations. Sebastian Flennerhag, Hujun Yin, John A. Keane, Mark J. Elliot |
NeurIPS | 4 |
| 2018 | The Application of Genetic Algorithms to Data Synthesis: A Comparison of Three Crossover Methods
Yingrui Chen, Mark J. Elliot, Duncan Smith |
PSD | 2 |
| 2018 | Differential Correct Attribution Probability for Synthetic Data: An Exploration
Jennifer Taub, Mark J. Elliot, Maria Pampaka, Duncan Smith |
PSD | 2 |
| 2018 | Functional anonymisation: Personal data and the data environment
Mark J. Elliot, Kieron O'Hara, Charles D. Raab, Christine M. O'Keefe, Elaine Mackey, Chris Dibben, Heather Gowans, Kingsley Purdam, Karen McCullagh |
Comput. Law Secur. Rev. | 1 |
| 2018 | Are 'pseudonymised' data always personal data? Implications of the GDPR for administrative data research in the UKabstractThere has naturally been a good deal of discussion of the forthcoming General Data Protection Regulation. One issue of interest to all data controllers, and of particular concern for researchers, is whether the GDPR expands the scope of personal data through the introduction of the term ‘pseudonymisation’ in Article 4(5). If all data which have been ‘pseudonymised’ in the conventional sense of the word (e.g. key-coded) are to be treated as personal data, this would have serious implications for research. Administrative data research, which is carried out on data routinely collected and held by public authorities, would be particularly affected as the sharing of de-identified data could constitute the unconsented disclosure of identifiable information. Instead, however, we argue that the definition of pseudonymisation in Article 4(5) GDPR will not expand the category of personal data, and that there is no intention that it should do so. The definition of pseudonymisation under the GDPR is not intended to determine whether data are personal data; indeed it is clear that all data falling within this definition are personal data. Rather, it is Recital 26 and its requirement of a ‘means reasonably likely to be used’ which remains the relevant test as to whether data are personal. This leaves open the possibility that data which have been ‘pseudonymised’ in the conventional sense of key-coding can still be rendered anonymous. There may also be circumstances in which data which have undergone pseudonymisation within one organisation could be anonymous for a third party. We explain how, with reference to the data environment factors as set out in the UK Anonymisation Network's Anonymisation Decision-Making Framework. Miranda Mourby, Elaine Mackey, Mark J. Elliot, Heather Gowans, Susan E. Wallace, Jessica Bell, Hannah Smith, Stergios Aidinlis, Jane Kaye |
Comput. Law Secur. Rev. | 3 |
| 2014 | Measuring Disclosure Risk with Entropy in Population Based Frequency Tables
Laszlo Antal, Natalie Shlomo, Mark J. Elliot |
Privacy in Statistical Databases | 3 |
| 2010 | Data Environment Analysis and the Key Variable Mapping System
Mark J. Elliot, Susan Lomax, Elaine Mackey, Kingsley Purdam |
Privacy in Statistical Databases | 1 |
| 2009 | Factors affecting the performance of parallel mining of minimal unique itemsets on diverse architecturesabstractAbstract Three parallel implementations of a divide‐and‐conquer search algorithm (called SUDA2) for finding minimal unique itemsets (MUIs) are compared in this paper. The identification of MUIs is used by national statistics agencies for statistical disclosure assessment. The first parallel implementation adapts SUDA2 to a symmetric multi‐processor cluster using the message passing interface (MPI), which we call an MPI cluster; the second optimizes the code for the Cray MTA2 (a shared‐memory, multi‐threaded architecture) and the third uses a heterogeneous ‘group’ of workstations connected by LAN. Each implementation considers the parallel structure of SUDA2, and how the subsearch computation times and sequence of subsearches affect load balancing. All three approaches scale with the number of processors, enabling SUDA2 to handle larger problems than before. For example, the MPI implementation is able to achieve nearly two orders of magnitude improvement with 132 processors. Performance results are given for a number of data sets. Copyright © 2009 John Wiley & Sons, Ltd. David J. Haglin, Kenneth R. Mayes, Anna M. Manning, John Feo, John R. Gurd, Mark J. Elliot, John A. Keane |
Concurr. Comput. Pract. Exp. | 6 |
| 2008 | Statistical disclosure control architectures for patient records in biomedical information systems
Mark J. Elliot, Kingsley Purdam, Duncan Smith |
J. Biomed. Informatics | 1 |
| 2002 | A Computational Algorithm for Handling the Special Uniques ProblemabstractMany organizations require detailed individual-level information, much of which has been collected under guarantees of confidentiality. However, simple anonymization procedures, i.e. removing names and addresses, are insufficient for this to be ensured. The records belonging to certain individuals have a high probability of being identified (as their contents, or attributes, are unusual) and therefore have the potential to be recognized spontaneously - such records are referred to as special uniques. Consider, for example, a sixteen-year-old widow in a population survey. Confidentiality of a given dataset cannot be enabled until all special unique records are identified and either disguised or removed. However, to the knowledge of the authors, no exhaustive automated analysis of this nature has been conducted due to the demanding levels of computation and data storage that are required. This paper introduces a new algorithm that locates 'risky' records in discrete data by first identifying all unique attribute sets (up to a user-specified maximum size) and secondly by grading the 'risk' of each record by considering the number and distribution of unique attribute sets within each record. Empirical tests indicate that the algorithm is highly effective at picking out 'risky' records from large samples of data. Mark J. Elliot, Anna M. Manning, Rupert W. Ford |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |