VLDB 2026 Research / reviewers in the wild / expert
Gavin Smith
dblp:58/2979
· DBLP profile ↗
13ranked-venue papers in the field
2as first author
4since 2021 · last 2025
0000-0001-5679-6309ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 8 (1 first)Database Systems & Data Management · 2Data Mining & Knowledge Discovery · 1Information Retrieval & Web Search · 1 (1 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automatic Lifestate Identification for High-Dimensional Time Series DataabstractTime series summarisation methods that account for temporal structure in high-dimensional data are important for analysis in a wide variety of domains, yet current statistical and machine-learning tools offer limited support for this task. We introduce ALI (Automatic Lifestate Identification), a parameter-free algorithm that, across high-dimensional time series, automatically identifies and clusters temporal patterns shared across entities (e.g. customers), effectively capturing latent states without requiring manual tuning. For example, in retail settings, these 'latent states' represent distinct, persistent behavioural modes, such as periods of high-value customer engagement, shifts in purchasing habits due to life events, or stable patterns of product consumption. ALI utilises an Expectation Maximisation algorithm and achieves a time complexity of O(h n log(n)) per iteration, where h represents the number of latent states and n the total length of concatenated time series. Empirical results show fast convergence. While no previous work directly addresses the specific problem motivated in this work, on synthetic data with known ground truth, ALI outperforms existing methods at a number of comparable subtasks, including identifying the generating (k,h) pairs (number of segments k, and number of latent states h) on single time series. Applied to large scale retail transactions, ALI recovers interpretable lifestates that align with known events (e.g., transition to parenthood), yielding compact, practitioner friendly summaries without manual tuning. Samuel Smith, Gavin Smith |
IEEE Big Data | 2 |
| 2022 | Privacy-preserving & machine-learned catchment models for national dietary surveillance via digital footprint dataabstractBig data from food retail stores is increasingly being used for population dietary surveillance, epidemiological studies of diet-related diseases, and evaluations of public health interventions. However, for retail data to be useful it is necessary to understand the spatio-temporal variation of when and where food is purchased and consumed. While some customers willingly share home location data with retailers as part of loyalty programs such data is typically too fine-grained/sensitive to be applied for research purposes. The aim of this study was to analyse differences between privacy-preserving models and actual retail catchments, and investigate if machine learning techniques could improve the accuracy of such catchment models. Based on a UK-wide sample of 4 million grocery store loyalty card holders, covering 485 million transactions over 29 months (2019-2021) and distributed across 33,000 neighbourhoods (Lower Super Output Areas, or LSOA), the study demonstrates how models trained on geolocated data perform at predicting, per store, catchment areas which contain 50, 80, and 95% of its customers’ primary location. Through comparative assessment of machine learning approaches, we find better performance from tree-based models (RF, XGB) with the best performance from an XGB model achieving an R2of 0.72 and MAE of 1.06. To conclude, we review variable importance measures using SHAP values and discuss the relative merits of including specific features when modeling catchment areas. Gavin Long, Gavin Smith, Georgiana Nica-Avram, Gregor Engelmann, James Goulding |
IEEE Big Data | 3 |
| 2022 | Bundle entropy as an optimized measure of consumers' systematic product choice combinations in mass transactional dataabstractUnderstanding and measuring the predictability of consumer purchasing (basket) behaviour is of significant value. While predictability measures such as entropy have been well studied and leveraged in other sectors, their development and application to very large multi-dimensional data sets present in the retailing sector are less common. While a small number of methods exist, we demonstrate they fail to accord with intuition, leading to the potential for misunderstandings between those who conduct the analysis and those who act on the insights. We delineate the requirements for such a measure in this domain to demonstrate these issues in context. A novel measure is then developed based on entropy to directly measure the predictability of basket composition. The measure is designated as bundle entropy (zero denotes a bundle’s total predictability, one the total unpredictability). We empirically compare the proposed bundle entropy against existing measures using two large-scale real-world transactional data sets, each including more than 2,000 households (frequent shoppers) over two years. First, we demonstrate how the proposed measure is the only measure that behaves according to the desired properties. Second, we show empirically that bundle entropy differs noticeably from the other measures. Finally, we consider some use case analyses and discuss the utility of the proposed measure in practice. Roberto Mansilla, Gavin Smith, James Goulding |
IEEE Big Data | 2 |
| 2021 | Using Model Class Reliance to Measure Group Effects on Non-Adherence to Asthma MedicationabstractAsthma affects an estimated 300 million people across the world. Despite being a highly treatable condition using preventative inhalers, mortality rates remain unacceptably high, with lack of adherence to medications cited as a major cause. While various drivers for non-adherence have been considered in isolation, interactions between demographic, behavioural and situational factors have never been modelled in concert mostly due to the limited ability of traditional methods to group such a large variety of features. This was addressed in this paper through a non-linear modelling approach, leveraging a novel dataset obtained via online surveying of asthma patients. Application of traditional variable importance methods to examine explanatory factors, however, is not possible. This is due to the presence of high multicollinearity in the data, a highly common occurrence in big data, or any datasets which include a large number of input features. This results in insights being obfuscated by extensive shared information and non-linear interactions occurring across variables. To mitigate this, we introduce the first Grouped Feature approach to Model Class Reliance (Group-MCR), that is able to quantify the importance of specific variable sets in underpinning explanations. Cross-validated models achieve 71% accuracy, with Group-MCR revealing the importance of perceptual factors. Out of all the perceptual factors denial proves to be most predictive of non-adherence to asthma medication, indicating that public health interventions should not only target the physical aspects of asthma, but additionally focus on patients’ beliefs and perceptions as valuable parts of their treatment. Vanja Ljevar, James Goulding, Gavin Smith, Alexa Spence |
IEEE BigData | 3 |
| 2020 | Exploration of links between anxiety purchases, deprivation and personality traitsabstractThe links between anxiety (and negative mental heath outcomes in general) and socio-economic conditions have been the subject of a large number of studies. However, the underlying mechanisms that affect this relationship have not been fully elucidated, nor have they been extended to consider potential mediators in the form of individual differences/psychological traits. Interrogating over 8-million customers' loyalty card transactions from a major health retailer, this paper investigated two ideas: the potential to use health product purchases to detect anxiety distributions across geo-spatial regions in England; and exploring the relationship between these product health purchases, deprivation levels and personality traits. Specifically, this analysis examined the co-variation between: anxiety purchases within district level geo-spatial regions in the UK; mean deprivation levels; and personality traits, across those districts. Contrary to previous findings, results demonstrated a negative correlation between anxiety related purchases and deprivation, and a positive correlation between anxiety related purchases and conscientiousness. This indicated the complex nature of the various forms of anxiety and its underlying causes and drivers - but also highlighted the challenges faced by different demographics in treating its symptoms. Vanja Ljevar, James Goulding, Gavin Smith |
IEEE BigData | 3 |
| 2020 | Perception detection using TwitterabstractPatients' perceptions about their condition have a strong impact on not only adherence to medication, but also on how they view themselves in the light of their condition. Research implies that Twitter is a particularly rich source of perceptions, as patients frequently use internet for information sharing and support. However, Twitter contains a lot of noise in the form of tweets that do not relate to perceptions, but are rather generated to advertise research and corporate news and this kind of information could `pollute' perception analysis. This study examined methods that could be used to extract perception tweets, on the example of tweets related to asthma. We first demonstrated differences between perception and non-perception tweets in terms of their linguistic features, and then focused on filtering perceptions using the classification process. Results demonstrated that there is a significant difference between perceptions and non-perceptions: perception tweets are shorter, have less capital letters, less punctuation signs and less hashtags. These features also performed well in predicting perceptions. However, the bag of words approach had better results in distinguishing between perception and non-perception tweets and the best results were obtained using word-based frequency vectorization and by training a neural network based classifier. Future research could explore the synergy of these approaches. Vanja Ljevar, James Goulding, Alexa Spence, Gavin Smith |
IEEE BigData | 4 |
| 2018 | The Unbanked and Poverty: Predicting area-level socio-economic vulnerability from M-Money transactionsabstractEmerging economies around the world are often characterized by governments and institutions struggling to keep key demographic data streams up to date. A demographic of interest particularly linked to social vulnerability is that of poverty and socio-economic status. The combination of mass call detail records (CDR) data with machine learning has recently been proposed as a way to obtain this data without the expense required by traditional census and household survey methods. Based on a sample of 330k mobile phone subscribers resident in Dar es Salaam, Tanzania (7.6m M-Money records, 450.2m call and SMS event logs) this paper demonstrates the improvements that can be made via an alternate data stream: M-Money transaction records. An alternative to traditional banking services, particularly utilized by citizens unable to obtain a bank account, M-Money transactions provide a currently unexplored but potentially more powerful data set held by the same telecommunication companies.Comparing directly to CDR as used in prior work the results show that M-Money provides an increase in socio-demographic classification accuracy (average F1 score) from 65.9% (0.63) to 71.3% (0.7) at much finer-grained spatial regions than previously examined. Notably, the combined use of M-Money and CDR data only increases prediction accuracy (average F1 score) from 71.3% (0.7) to 72.3% (0.71), providing evidence that M-Money is informationally subsuming CDR data. The reasons for this and the importance/contributions of individual features are subsequently investigated. Gregor Engelmann, Gavin Smith, James Goulding |
IEEE BigData | 2 |
| 2017 | Data-driven estimation of building interior plansabstractThis work investigates constructing plans of building interiors using learned building measurements. In particular, we address the problem of accurately estimating dimensions of rooms when measurements of the interior space have not been captured. Our approach focuses on learning the geometry, orientation and occurrence of rooms from a corpus of real-world building plan data to form a predictive model. The trained predictive model may then be queried to generate estimates of room dimensions and orientations. These estimates are then integrated with the overall building footprint and iteratively improved using a two-stage optimisation process to form complete interior plans.The approach is presented as a semi-automatic method for constructing plans which can cope with a limited set of known information and constructs likely representations of building plans through modelling of soft and hard constraints. We evaluate the method in the context of estimating residential house plans and demonstrate that predictions can effectively be used for constructing plans given limited prior knowledge about the types of rooms and their topology. Julian F. Rosser, Gavin Smith, Jeremy G. Morley |
Int. J. Geogr. Inf. Sci. | 2 |
| 2016 | Event Series Prediction via Non-Homogeneous Poisson Process ModellingabstractData streams whose events occur at random arrival times rather than at the regular, tick-tock intervals of traditional time series are increasingly prevalent. Event series are continuous, irregular and often highly sparse, differing greatly in nature to the regularly sampled time series traditionally the concern of hard sciences. As mass sets of such data have become more common, so interest in predicting future events in them has grown. Yet repurposing of traditional forecasting approaches has proven ineffective, in part due to issues such as sparsity, but often due to inapplicable underpinning assumptions such as stationarity and ergodicity. In this paper we derive a principled new approach to forecasting event series that avoids such assumptions, based upon: 1. The processing of event series datasets in order to produce a first parameterized mixture model of non-homogeneous Poisson processes, and 2. Application of a technique called parallel forecasting that uses these processes' rate functions to directly generate accurate temporal predictions for new query realizations. This approach uses forerunners of a stochastic process to shed light on the distribution of future events, not for themselves, but for realizations that subsequently follow in their footsteps. James Goulding, Simon Preston, Gavin Smith |
ICDM | 3 |
| 2015 | A novel symbolization technique for time-series outlier detectionabstractThe detection of outliers in time series data is a core component of many data-mining applications and broadly applied in industrial applications. In large data sets algorithms that are efficient in both time and space are required. One area where speed and storage costs can be reduced is via symbolization as a pre-processing step, additionally opening up the use of an array of discrete algorithms. With this common pre-processing step in mind, this work highlights that (1) existing symbolization approaches are designed to address problems other than outlier detection and are hence sub-optimal and (2) use of off-the-shelf symbolization techniques can therefore lead to significant unnecessary data corruption and potential performance loss when outlier detection is a key aspect of the data mining task at hand. Addressing this a novel symbolization method is motivated specifically targeting the end use application of outlier detection. The method is empirically shown to outperform existing approaches. Gavin Smith, James Goulding |
IEEE BigData | 1 |
| 2013 | Improving route prediction through user journey detectionabstractThe positioning datasets that underpin route prediction models arrive as time series or point process logs. However, their use for prediction requires them to be split into meaningful segments, conceptualised as travelling periods or 'journeys', to form a set of training inputs. Despite significant research into route prediction, this important pre-processing step has traditionally occurred in an ad-hoc fashion, using arbitrary connectivity or movement thresholds. There has been little consideration to date of the impact of this on prediction, a fact rectified in this work. Mark Dimond, Gavin Smith, James Goulding |
SIGSPATIAL/GIS | 2 |
| 2012 | Evaluating implicit judgments from image search clickthrough dataabstractThe interactions of users with search engines can be seen as implicit relevance feedback by the user on the results offered to them. In particular, the selection of results by users can be interpreted as a confirmation of the relevance of those results, and used to reorder or prioritize subsequent search results. This collection of search/result pairings is called clickthrough data, and many uses for it have been proposed. However, the reliability of clickthrough data has been challenged and it has been suggested that clickthrough data are not a completely accurate measure of relevance between search term and results. This paper reports on an experiment evaluating the reliability of clickthrough data as a measure of the mutual relevance of search term and result. The experiment comprised a user study involving over 67 participants and determines the reliability of image search clickthrough data, using factors identified in previous similar studies. A major difference in this work to previous work is that the source of clickthrough data comes from image searches, rather than the traditional text page searches. Image search clickthrough data were rarely examined in prior works but has differences that impact the accuracy of clickthrough data. These differences include a more complete representation of the results in image search, allowing users to scrutinize the results more closely before selecting them, as well as presenting the results in a less obviously ordered way. The experiment reported here demonstrates that image clickthrough data can be more reliable as a relevance feedback measure than has been the case with traditional text‐based search. There is also evidence that the precision of the search system influences the accuracy of click data when users make searches in an information‐seeking capacity. Gavin Smith, Chris Brien, Helen Ashman |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2009 | Are Clickthroughs Useful for Image Labelling?abstractIn this paper we look at how images can be labelled as a result of click throughs from searches. One approach acts as a filter on image searches specifically, while the other approach propagates labels to images from their containing pages, where those pages were labelled themselves using clickthrough as a filter on text search. Then the paper reports on an experiment where users ranked for relevance six methods for labelling images, comparing the two clickthrough-based methods with flickr's amateur explicit labelling, Getty's professional explicit labelling, Google's standard image search, and the new Google Image Labeller. As well as comparing the accuracy of the proposed image labelling methods and discovering that automatic methods outperform explicit human labelling methods, the experiment suggests clickthrough data is reliable with very few clicks for image classification purposes. Helen Ashman, Michael Antunovic, Christoph Donner, Rebecca Frith, Eric Rebelos, Jan-Felix Schmakeit, Gavin Smith, Mark Truran |
Web Intelligence | 7 |