VLDB 2026 Research / reviewers in the wild / expert
Simon Preston
dblp:166/1397 · also S. P. Preston
· DBLP profile ↗
2ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Representation and self-supervised learning · 58% Trustworthy machine learning · 25% Probabilistic and Bayesian machine learning · 17% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › causal representation learning
identifiability |
0.4 | 1 | 2019 | Invariance and identifiability issues for word embeddings · NeurIPS 2019 |
Machine learning › Trustworthy machine learning
invariance |
0.4 | 1 | 2019 | Invariance and identifiability issues for word embeddings · NeurIPS 2019 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.4 | 1 | 2019 | Invariance and identifiability issues for word embeddings · NeurIPS 2019 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process |
0.2 | 1 | 2016 | Event Series Prediction via Non-Homogeneous Poisson Process Modelling · ICDM 2016 |
Data mining
temporal data mining |
0.2 | 1 | 2016 | Event Series Prediction via Non-Homogeneous Poisson Process Modelling · ICDM 2016 |
Machine learning › Representation and self-supervised learning
transformation invariance |
0.1 | 1 | 2019 | Invariance and identifiability issues for word embeddings · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
parallel forecasting · 0.5mixture model · 0.5numerical examples · 0.4formal analysis · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Invariance and identifiability issues for word embeddingsabstractWord embeddings are commonly obtained as optimisers of a criterion function f of a text corpus, but assessed on word-task performance using a different evaluation function g of the test data. We contend that a possible source of disparity in performance on tasks is the incompatibility between classes of transformations that leave f and g invariant. In particular, word embeddings defined by f are not unique; they are defined only up to a class of transformations to which f is invariant, and this class is larger than the class to which g is invariant. One implication of this is that the apparent superiority of one word embedding over another, as measured by word task performance, may largely be a consequence of the arbitrary elements selected from the respective solution sets. We provide a formal treatment of the above identifiability issue, present some numerical examples, and discuss possible resolutions. Rachel Carrington, Karthik Bharath, Simon Preston |
NeurIPS | 3 |
| 2016 | Event Series Prediction via Non-Homogeneous Poisson Process ModellingabstractData streams whose events occur at random arrival times rather than at the regular, tick-tock intervals of traditional time series are increasingly prevalent. Event series are continuous, irregular and often highly sparse, differing greatly in nature to the regularly sampled time series traditionally the concern of hard sciences. As mass sets of such data have become more common, so interest in predicting future events in them has grown. Yet repurposing of traditional forecasting approaches has proven ineffective, in part due to issues such as sparsity, but often due to inapplicable underpinning assumptions such as stationarity and ergodicity. In this paper we derive a principled new approach to forecasting event series that avoids such assumptions, based upon: 1. The processing of event series datasets in order to produce a first parameterized mixture model of non-homogeneous Poisson processes, and 2. Application of a technique called parallel forecasting that uses these processes' rate functions to directly generate accurate temporal predictions for new query realizations. This approach uses forerunners of a stochastic process to shed light on the distribution of future events, not for themselves, but for realizations that subsequently follow in their footsteps. James Goulding, Simon Preston, Gavin Smith |
ICDM | 2 |