VLDB 2026 Research / reviewers in the wild / expert
Thomas Drake
dblp:72/973
· DBLP profile ↗
5ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
2 papers |
Privacy and data protection · 100% | |
| Artificial intelligence
1 paper |
Trustworthy machine learning · 38% Multi-agent systems · 19% Efficient and distributed learning · 19% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Privacy and data protection
differential privacy |
0.8 | 2 | 2020 | Privacy- and Utility-Preserving Textual Analysis via Calibrated Multivariate Perturbations · WSDM 2020 Leveraging Hierarchical Representations for Preserving Privacy and Utility in Text · ICDM 2019 |
Privacy and data protection › location privacy
geo-indistinguishability |
0.4 | 1 | 2020 | Privacy- and Utility-Preserving Textual Analysis via Calibrated Multivariate Perturbations · WSDM 2020 |
Privacy and data protection
privacy-preserving data analysis |
0.4 | 1 | 2020 | Privacy- and Utility-Preserving Textual Analysis via Calibrated Multivariate Perturbations · WSDM 2020 |
Privacy and data protection › data confidentiality › content privacy
text privacy |
0.4 | 1 | 2019 | Leveraging Hierarchical Representations for Preserving Privacy and Utility in Text · ICDM 2019 |
Machine learning › Efficient and distributed learning
active learning |
0.3 | 1 | 2018 | Leveraging Crowdsourcing Data for Deep Active Learning An Application: Learning Intents in Alexa · WWW 2018 |
Machine learning › Trustworthy machine learning › Data-centric AI
annotator expertise modeling |
0.3 | 1 | 2018 | Leveraging Crowdsourcing Data for Deep Active Learning An Application: Learning Intents in Alexa · WWW 2018 |
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models
bayesian deep learning |
0.3 | 1 | 2018 | Leveraging Crowdsourcing Data for Deep Active Learning An Application: Learning Intents in Alexa · WWW 2018 |
Machine learning › Trustworthy machine learning
crowdsourced annotation |
0.3 | 1 | 2018 | Leveraging Crowdsourcing Data for Deep Active Learning An Application: Learning Intents in Alexa · WWW 2018 |
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction
human-in-the-loop learning |
0.3 | 1 | 2018 | Leveraging Crowdsourcing Data for Deep Active Learning An Application: Learning Intents in Alexa · WWW 2018 |
Privacy and data protection › privacy evaluation
privacy-utility tradeoff |
0.1 | 1 | 2019 | Leveraging Hierarchical Representations for Preserving Privacy and Utility in Text · ICDM 2019 |
Natural language and speech › Question answering and dialogue systems
intent detection |
0.1 | 1 | 2018 | Leveraging Crowdsourcing Data for Deep Active Learning An Application: Learning Intents in Alexa · WWW 2018 |
Methods — techniques the papers use, named apart from their topics
word embeddings · 0.4calibrated noise · 0.4hyperbolic word embedding · 0.4authorship attribution · 0.4uncertainty estimation · 0.3low-rank matrix factorization · 0.3bayesian deep learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Privacy- and Utility-Preserving Textual Analysis via Calibrated Multivariate PerturbationsabstractAccurately learning from user data while providing quantifiable privacy guarantees provides an opportunity to build better ML models while maintaining user trust. This paper presents a formal approach to carrying out privacy preserving text perturbation using the notion of d_χ-privacy designed to achieve geo-indistinguishability in location data. Our approach applies carefully calibrated noise to vector representation of words in a high dimension space as defined by word embedding models. We present a privacy proof that satisfies d_χ-privacy where the privacy parameter $\varepsilon$ provides guarantees with respect to a distance metric defined by the word embedding space. We demonstrate how $\varepsilon$ can be selected by analyzing plausible deniability statistics backed up by large scale analysis on GloVe and fastText embeddings. We conduct privacy audit experiments against $2$ baseline models and utility experiments on 3 datasets to demonstrate the tradeoff between privacy and utility for varying values of varepsilon on different task types. Our results demonstrate practical utility (< 2% utility loss for training binary classifiers) while providing better privacy guarantees than baseline models. Oluwaseyi Feyisetan, Borja Balle, Thomas Drake, Tom Diethe |
WSDM | 3 |
| 2019 | Leveraging Hierarchical Representations for Preserving Privacy and Utility in TextabstractGuaranteeing a certain level of user privacy in an arbitrary piece of text is a challenging issue. However, with this challenge comes the potential of unlocking access to vast data stores for training machine learning models and supporting data driven decisions. We address this problem through the lens of dx-privacy, a generalization of Differential Privacy to non Hamming distance metrics. In this work, we explore word representations in Hyperbolic space as a means of preserving privacy in text. We provide a proof satisfying dx-privacy, then we define a probability distribution in Hyperbolic space and describe a way to sample from it in high dimensions. Privacy is provided by perturbing vector representations of words in high dimensional Hyperbolic space to obtain a semantic generalization. We conduct a series of experiments to demonstrate the tradeoff between privacy and utility. Our privacy experiments illustrate protections against an authorship attribution algorithm while our utility experiments highlight the minimal impact of our perturbations on several downstream machine learning models. Compared to the Euclidean baseline, we observe > 20x greater guarantees on expected privacy against comparable worst case statistics. Oluwaseyi Feyisetan, Tom Diethe, Thomas Drake |
ICDM | 3 |
| 2018 | Leveraging Crowdsourcing Data for Deep Active Learning An Application: Learning Intents in AlexaabstractThis paper presents a generic Bayesian framework that enables any deep learning model to actively learn from targeted crowds. Our framework inherits from recent advances in Bayesian deep learning, and extends existing work by considering the targeted crowdsourcing approach, where multiple annotators with unknown expertise contribute an uncontrolled amount (often limited) of annotations. Our framework leverages the low-rank structure in annotations to learn individual annotator expertise, which then helps to infer the true labels from noisy and sparse annotations. It provides a unified Bayesian model to simultaneously infer the true labels and train the deep learning model in order to reach an optimal learning efficacy. Finally, our framework exploits the uncertainty of the deep learning model during prediction as well as the annotators» estimated expertise to minimize the number of required annotations and annotators for optimally training the deep learning model. We evaluate the effectiveness of our framework for intent classification in Alexa (Amazon»s personal assistant), using both synthetic and real-world datasets. Experiments show that our framework can accurately learn annotator expertise, infer true labels, and effectively reduce the amount of annotations in model training as compared to state-of-the-art approaches. We further discuss the potential of our proposed framework in bridging machine learning and crowdsourcing towards improved human-in-the-loop systems. Jie Yang 0028, Thomas Drake, Andreas Damianou, Yoelle Maarek |
WWW | 2 |
| 2009 | Observations of Thermal Variations in the Mixed Layer Depth of the Equatorial AtlanticabstractA set of Argo temperature data collected in the equatorial Atlantic [0°-5°N, 55°W-10°E] was used to estimate the mixed layer depth (MLD) and associated thermal variability for the period between January 2002 to April 2009. MLD climatology were estimated from 0.3°×0.3° median binned temperature profile using temperature difference criterion with a reference layer at 10m depth. At the 30m depth, 22°C cold water flows from the south onto the continental shelfs of Ghana-Cote D'Ivoire indicating the potential source of nutrient rich bottom water that nourishes the MLD and drives biological production. The MLD was shallow at the east and relatively deeper at the western end of the equatorial Atlantic. Variability within the MLD can be associated with variations in the westward flow of the warm and saline Equatorial Undercurrent. Further warming of the equatorial Atlantic has a potential of increasing the mixed layer depth and affecting upper surface ocean processes. Kwame Agyekum, George Wiafe, Bob Houghton, Shaun Dolk, Thomas Drake, Augustus Vogel |
IGARSS (1) | 5 |
| 1997 | From Lectures to the World Wide Web: Some Instructional Design Considerations
Anju Relan, Thomas Drake |
AMIA | 2 |