EDBT 2026 Demo / reviewers in the wild / expert
Teddy Cunningham
dblp:298/7835
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2022
0000-0002-8829-2532ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | RAGUEL: Recourse-Aware Group Unfairness EliminationabstractWhile machine learning and ranking-based systems are in widespread use for sensitive decision-making processes (e.g., determining job candidates, assigning credit scores), they are rife with concerns over unintended biases in their outcomes, which makes algorithmic fairness (e.g., demographic parity, equal opportunity) an objective of interest. 'Algorithmic recourse' offers feasible recovery actions to change unwanted outcomes through the modification of attributes. We introduce the notion of ranked group-level recourse fairness, and develop a 'recourse-aware ranking' solution that satisfies ranked recourse fairness constraints while minimizing the cost of suggested modifications. Our solution suggests interventions that can reorder the ranked list of database records and mitigate group-level unfairness; specifically, disproportionate representation of sub-groups and recourse cost imbalance. This re-ranking identifies the minimum modifications to data points, with these attribute modifications weighted according to their ease of recourse. We then present an efficient block-based extension that enables re-ranking at any granularity (e.g., multiple brackets of bank loan interest rates, multiple pages of search engine results). Evaluation on real datasets shows that, while existing methods may even exacerbate recourse unfairness, our solution – RAGUEL – significantly improves recourse-aware fairness. RAGUEL outperforms alternatives at improving recourse fairness, through a combined process of counterfactual generation and re-ranking, whilst remaining efficient for large-scale datasets. Aparajita Haldar, Teddy Cunningham, Hakan Ferhatosmanoglu |
CIKM | 2 |
| 2022 | Sharing and Generating Privacy-Preserving Spatio-Temporal Data Using Real-World KnowledgeabstractPrivacy-preserving spatio-temporal data sharing is vital in many machine learning and analysis tasks, such as managing disease spread or tailoring public services to a population's travel patterns. Current methods for data release are insufficiently accurate to provide meaningful utility, and they carry a high risk of deanonymization or membership inference attacks. These limitations and public concern over privacy and data protection has limited the extent to which data is shared. This work presents approaches generating and publishing spatio-temporal data, such as geographic locations and trajectories, with differential privacy. In the first solution, differentially private spatial data is generated using kernel density estimation and a road network-aware approach. In the second solution, a local differentially private mechanism is developed by perturbing hierarchically-structured, overlapping n-grams of trajectory data. Both of the solutions incorporate publicly available information, such as the road network or categories of places of interests, to enhance the utility of the output data without negatively affecting privacy or efficiency. Experiments with real-world data demonstrate that the private data can perform as well as the non-private data in a range of practical data science tasks. Teddy Cunningham |
MDM | 1 |
| 2021 | Privacy-Preserving Synthetic Location Data in the Real WorldabstractSharing sensitive data is vital in enabling many modern data analysis and machine learning tasks. However, current methods for data release are insufficiently accurate or granular to provide meaningful utility, and they carry a high risk of deanonymization or membership inference attacks. In this paper, we propose a differentially private synthetic data generation solution with a focus on the compelling domain of location data. We present two methods with high practical utility for generating synthetic location data from real locations, both of which protect the existence and true location of each individual in the original dataset. Our first, partitioning-based approach introduces a novel method for privately generating point data using kernel density estimation, in addition to employing private adaptations of classic statistical techniques, such as clustering, for private partitioning. Our second, network-based approach incorporates public geographic information, such as the road network of a city, to constrain the bounds of synthetic data points and hence improve the accuracy of the synthetic data. Both methods satisfy the requirements of differential privacy, while also enabling accurate generation of synthetic data that aims to preserve the distribution of the real locations. We conduct experiments using three large-scale location datasets to show that the proposed solutions generate synthetic location data with high utility and strong similarity to the real datasets. We highlight some practical applications for our work by applying our synthetic data to a range of location analytics queries, and we demonstrate that our synthetic data produces near-identical answers to the same queries compared to when real data is used. Our results show that the proposed approaches are practical solutions for sharing and analyzing sensitive location data privately. Teddy Cunningham, Graham Cormode, Hakan Ferhatosmanoglu |
SSTD | 1 |
| 2021 | Real-World Trajectory Sharing with Local Differential PrivacyabstractSharing trajectories is beneficial for many real-world applications, such as managing disease spread through contact tracing and tailoring public services to a population's travel patterns. However, public concern over privacy and data protection has limited the extent to which this data is shared. Local differential privacy enables data sharing in which users share a perturbed version of their data, but existing mechanisms fail to incorporate user-independent public knowledge (e.g., business locations and opening times, public transport schedules, geo-located tweets). This limitation makes mechanisms too restrictive, gives unrealistic outputs, and ultimately leads to low practical utility. To address these concerns, we propose a local differentially private mechanism that is based on perturbing hierarchically-structured, overlapping n -grams (i.e., contiguous subsequences of length n ) of trajectory data. Our mechanism uses a multi-dimensional hierarchy over publicly available external knowledge of real-world places of interest to improve the realism and utility of the perturbed, shared trajectories. Importantly, including real-world public data does not negatively affect privacy or efficiency. Our experiments, using real-world data and a range of queries, each with real-world application analogues, demonstrate the superiority of our approach over a range of alternative methods. Teddy Cunningham, Graham Cormode, Hakan Ferhatosmanoglu, Divesh Srivastava |
Proc. VLDB Endow. | 1 |