VLDB 2026 Research / reviewers in the wild / expert
James Goulding
dblp:18/2993
· DBLP profile ↗
17ranked-venue papers in the field
1as first author
7since 2021 · last 2023
0000-0002-8892-6398ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 10Data Mining & Knowledge Discovery · 4 (1 first)Database Systems & Data Management · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Who consumes anthocyanins and anthocyanidins? Mining national retail data to reveal the influence of socioeconomic deprivation and seasonality on polyphenol dietary intakeabstractAnthocyanins are a class of polyphenols that have received widespread recent attention due to their potential health benefits. However, estimating the dietary intake of anthocyanins at a population level is a challenging task, due to the difficulty of scaling dietary surveys. Further, there is limited evidence as to who regularly consumes anthocyanins, whether temporally, spatially, or culturally according to levels of socioeconomic deprivation. Leveraging a massive retail loyalty card dataset in the UK, we pair two years of real-world purchasing data for 619,524 regular shoppers and 207 million shopping baskets with anthocyanin estimates drawn from polyphenol databases. We subsequently analyse relative deprivation levels of the neighbourhoods in which shoppers reside, illustrating how anthocyanin intake varies according to affluence. Results indicate that deprivation is linked dramatically with both lower total intake of anthocyanins and lower breadth of dietary sources for them, potentially aggravating the incidence of diet-related diseases in the poorest sections of society. Gavin Long, Roberto Mansilla, Simon Welham, Peter Rose, Michelle Thomas, Gregor Milligan, Elizabeth Dolan, Joanne Parkes, Kuzivakwashe Makokoro, James Goulding |
IEEE Big Data | 11 |
| 2023 | Assessing relative contribution of Environmental, Behavioural and Social factors on Life Satisfaction via mobile app dataabstractLife satisfaction significantly contributes to wellbeing and is linked to positive outcomes for individual people and society more broadly. However, previous research demonstrates that many factors contribute to the life satisfaction of an individual person, including: demography, socioeconomic status, health, deprivation, family life, friendships, social networks, living environment, and the broad range of behaviours enacted by the person, such as helping or volunteering. Consequently, it is challenging to disentangle the factors that contribute most significantly to life satisfaction, and thus more importantly, inform public policies designed to help foster positive wellbeing. We analyse primary survey data $(\mathrm{n}=2849)$ on self-reported life satisfaction in relation to a range of self-reported and observed variables associated with wellbeing. Specifically, we draw on a massive paired dataset related to use of a food sharing application in London, to augment the analysis using additional socioeconomic, environmental, and behavioural variables. Through a random forest machine learning approach and variable importance measures, we evaluate how a range of factors, that are often only evaluated individually, provide relative contributions towards life satisfaction. Result reveal that factors such as employment and social reliance contribute most significantly towards the experience of life satisfaction. Gregor Milligan, Liz Dowthwaite, Elvira Perez, Georgiana Nica-Avram, James Goulding |
IEEE Big Data | 6 |
| 2022 | Towards Idea Mining: Problem-Solution Phrase Extraction from Text
Haixia Liu 0001, Tim J. Brailsford, James Goulding, Tomas Maul, Tao Tan 0002, Debanjan Chaudhuri |
ADMA (2) | 3 |
| 2022 | Privacy-preserving & machine-learned catchment models for national dietary surveillance via digital footprint dataabstractBig data from food retail stores is increasingly being used for population dietary surveillance, epidemiological studies of diet-related diseases, and evaluations of public health interventions. However, for retail data to be useful it is necessary to understand the spatio-temporal variation of when and where food is purchased and consumed. While some customers willingly share home location data with retailers as part of loyalty programs such data is typically too fine-grained/sensitive to be applied for research purposes. The aim of this study was to analyse differences between privacy-preserving models and actual retail catchments, and investigate if machine learning techniques could improve the accuracy of such catchment models. Based on a UK-wide sample of 4 million grocery store loyalty card holders, covering 485 million transactions over 29 months (2019-2021) and distributed across 33,000 neighbourhoods (Lower Super Output Areas, or LSOA), the study demonstrates how models trained on geolocated data perform at predicting, per store, catchment areas which contain 50, 80, and 95% of its customers’ primary location. Through comparative assessment of machine learning approaches, we find better performance from tree-based models (RF, XGB) with the best performance from an XGB model achieving an R2of 0.72 and MAE of 1.06. To conclude, we review variable importance measures using SHAP values and discuss the relative merits of including specific features when modeling catchment areas. Gavin Long, Gavin Smith, Georgiana Nica-Avram, Gregor Engelmann, James Goulding |
IEEE Big Data | 6 |
| 2022 | Bundle entropy as an optimized measure of consumers' systematic product choice combinations in mass transactional dataabstractUnderstanding and measuring the predictability of consumer purchasing (basket) behaviour is of significant value. While predictability measures such as entropy have been well studied and leveraged in other sectors, their development and application to very large multi-dimensional data sets present in the retailing sector are less common. While a small number of methods exist, we demonstrate they fail to accord with intuition, leading to the potential for misunderstandings between those who conduct the analysis and those who act on the insights. We delineate the requirements for such a measure in this domain to demonstrate these issues in context. A novel measure is then developed based on entropy to directly measure the predictability of basket composition. The measure is designated as bundle entropy (zero denotes a bundle’s total predictability, one the total unpredictability). We empirically compare the proposed bundle entropy against existing measures using two large-scale real-world transactional data sets, each including more than 2,000 households (frequent shoppers) over two years. First, we demonstrate how the proposed measure is the only measure that behaves according to the desired properties. Second, we show empirically that bundle entropy differs noticeably from the other measures. Finally, we consider some use case analyses and discuss the utility of the proposed measure in practice. Roberto Mansilla, Gavin Smith, James Goulding |
IEEE Big Data | 4 |
| 2022 | Ill-fated interactions: modeling complaints on a food waste fighting platformabstractThe redistribution of surplus food is a challenging problem, yet a crucial one to address given the urgent nature of climate change. However, designing computer-mediated food sharing systems is made even harder due to failed interactions between users and ensuing complaints, which can dissuade others from participating when shared within a public forum. To examine the phenomenon of complaints within such data, we analyze the public forum of a food sharing platform, OLIO. We characterize complaining behaviour and augment it through qualitative labeling and a machine learning approach to model complaints using affective indicators of dissatisfaction across a corpus of 3,195 forum posts. Results emphasize that linguistic features yield high prediction accuracies, with negative, nonconstructive sentiment being of greatest relevance. We discuss how machine learning can further enrich qualitative understandings and validation of complaints in the sharing economy. Georgiana Nica-Avram, Vanja Ljevar, Ines Branco-Illodo, H. P. Samanthika Gallage, James Goulding |
IEEE Big Data | 6 |
| 2021 | Using Model Class Reliance to Measure Group Effects on Non-Adherence to Asthma MedicationabstractAsthma affects an estimated 300 million people across the world. Despite being a highly treatable condition using preventative inhalers, mortality rates remain unacceptably high, with lack of adherence to medications cited as a major cause. While various drivers for non-adherence have been considered in isolation, interactions between demographic, behavioural and situational factors have never been modelled in concert mostly due to the limited ability of traditional methods to group such a large variety of features. This was addressed in this paper through a non-linear modelling approach, leveraging a novel dataset obtained via online surveying of asthma patients. Application of traditional variable importance methods to examine explanatory factors, however, is not possible. This is due to the presence of high multicollinearity in the data, a highly common occurrence in big data, or any datasets which include a large number of input features. This results in insights being obfuscated by extensive shared information and non-linear interactions occurring across variables. To mitigate this, we introduce the first Grouped Feature approach to Model Class Reliance (Group-MCR), that is able to quantify the importance of specific variable sets in underpinning explanations. Cross-validated models achieve 71% accuracy, with Group-MCR revealing the importance of perceptual factors. Out of all the perceptual factors denial proves to be most predictive of non-adherence to asthma medication, indicating that public health interventions should not only target the physical aspects of asthma, but additionally focus on patients’ beliefs and perceptions as valuable parts of their treatment. Vanja Ljevar, James Goulding, Gavin Smith, Alexa Spence |
IEEE BigData | 2 |
| 2020 | Exploration of links between anxiety purchases, deprivation and personality traitsabstractThe links between anxiety (and negative mental heath outcomes in general) and socio-economic conditions have been the subject of a large number of studies. However, the underlying mechanisms that affect this relationship have not been fully elucidated, nor have they been extended to consider potential mediators in the form of individual differences/psychological traits. Interrogating over 8-million customers' loyalty card transactions from a major health retailer, this paper investigated two ideas: the potential to use health product purchases to detect anxiety distributions across geo-spatial regions in England; and exploring the relationship between these product health purchases, deprivation levels and personality traits. Specifically, this analysis examined the co-variation between: anxiety purchases within district level geo-spatial regions in the UK; mean deprivation levels; and personality traits, across those districts. Contrary to previous findings, results demonstrated a negative correlation between anxiety related purchases and deprivation, and a positive correlation between anxiety related purchases and conscientiousness. This indicated the complex nature of the various forms of anxiety and its underlying causes and drivers - but also highlighted the challenges faced by different demographics in treating its symptoms. Vanja Ljevar, James Goulding, Gavin Smith |
IEEE BigData | 2 |
| 2020 | Perception detection using TwitterabstractPatients' perceptions about their condition have a strong impact on not only adherence to medication, but also on how they view themselves in the light of their condition. Research implies that Twitter is a particularly rich source of perceptions, as patients frequently use internet for information sharing and support. However, Twitter contains a lot of noise in the form of tweets that do not relate to perceptions, but are rather generated to advertise research and corporate news and this kind of information could `pollute' perception analysis. This study examined methods that could be used to extract perception tweets, on the example of tweets related to asthma. We first demonstrated differences between perception and non-perception tweets in terms of their linguistic features, and then focused on filtering perceptions using the classification process. Results demonstrated that there is a significant difference between perceptions and non-perceptions: perception tweets are shorter, have less capital letters, less punctuation signs and less hashtags. These features also performed well in predicting perceptions. However, the bag of words approach had better results in distinguishing between perception and non-perception tweets and the best results were obtained using word-based frequency vectorization and by training a neural network based classifier. Future research could explore the synergy of these approaches. Vanja Ljevar, James Goulding, Alexa Spence, Gavin Smith |
IEEE BigData | 2 |
| 2018 | The Unbanked and Poverty: Predicting area-level socio-economic vulnerability from M-Money transactionsabstractEmerging economies around the world are often characterized by governments and institutions struggling to keep key demographic data streams up to date. A demographic of interest particularly linked to social vulnerability is that of poverty and socio-economic status. The combination of mass call detail records (CDR) data with machine learning has recently been proposed as a way to obtain this data without the expense required by traditional census and household survey methods. Based on a sample of 330k mobile phone subscribers resident in Dar es Salaam, Tanzania (7.6m M-Money records, 450.2m call and SMS event logs) this paper demonstrates the improvements that can be made via an alternate data stream: M-Money transaction records. An alternative to traditional banking services, particularly utilized by citizens unable to obtain a bank account, M-Money transactions provide a currently unexplored but potentially more powerful data set held by the same telecommunication companies.Comparing directly to CDR as used in prior work the results show that M-Money provides an increase in socio-demographic classification accuracy (average F1 score) from 65.9% (0.63) to 71.3% (0.7) at much finer-grained spatial regions than previously examined. Notably, the combined use of M-Money and CDR data only increases prediction accuracy (average F1 score) from 71.3% (0.7) to 72.3% (0.71), providing evidence that M-Money is informationally subsuming CDR data. The reasons for this and the importance/contributions of individual features are subsequently investigated. Gregor Engelmann, Gavin Smith, James Goulding |
IEEE BigData | 3 |
| 2018 | Generating vague neighbourhoods through data mining of passive web dataabstractNeighbourhoods have been described as ‘the building blocks of public services society’. Their subjective nature, however, and the resulting difficulties in collecting data, means that in many countries there are no officially defined neighbourhoods either in terms of names or boundaries. This has implications not only for policy but also business and social decisions as a whole. With the absence of neighbourhood boundaries many studies resort to using standard administrative units as proxies. Such administrative geographies, however, often have a poor fit with those perceived by residents. Our approach detects these important social boundaries by automatically mining the Web en masse for passively declared neighbourhood data within postal addresses. Focusing on the United Kingdom (UK), this research demonstrates the feasibility of automated extraction of urban neighbourhood names and their subsequent mapping as vague entities. Importantly, and unlike previous work, our process does not require any neighbourhood names to be established a priori. Paul Brindley, James Goulding, Max L. Wilson 0001 |
Int. J. Geogr. Inf. Sci. | 2 |
| 2017 | Seasonal Variation in Collective Mood via Twitter Content and Medical Purchases
Fabon Dzogang, James Goulding, Stafford Lightman, Nello Cristianini |
IDA | 2 |
| 2016 | Event Series Prediction via Non-Homogeneous Poisson Process ModellingabstractData streams whose events occur at random arrival times rather than at the regular, tick-tock intervals of traditional time series are increasingly prevalent. Event series are continuous, irregular and often highly sparse, differing greatly in nature to the regularly sampled time series traditionally the concern of hard sciences. As mass sets of such data have become more common, so interest in predicting future events in them has grown. Yet repurposing of traditional forecasting approaches has proven ineffective, in part due to issues such as sparsity, but often due to inapplicable underpinning assumptions such as stationarity and ergodicity. In this paper we derive a principled new approach to forecasting event series that avoids such assumptions, based upon: 1. The processing of event series datasets in order to produce a first parameterized mixture model of non-homogeneous Poisson processes, and 2. Application of a technique called parallel forecasting that uses these processes' rate functions to directly generate accurate temporal predictions for new query realizations. This approach uses forerunners of a stochastic process to shed light on the distribution of future events, not for themselves, but for realizations that subsequently follow in their footsteps. James Goulding, Simon Preston, Gavin Smith |
ICDM | 1 |
| 2015 | A novel symbolization technique for time-series outlier detectionabstractThe detection of outliers in time series data is a core component of many data-mining applications and broadly applied in industrial applications. In large data sets algorithms that are efficient in both time and space are required. One area where speed and storage costs can be reduced is via symbolization as a pre-processing step, additionally opening up the use of an array of discrete algorithms. With this common pre-processing step in mind, this work highlights that (1) existing symbolization approaches are designed to address problems other than outlier detection and are hence sub-optimal and (2) use of off-the-shelf symbolization techniques can therefore lead to significant unnecessary data corruption and potential performance loss when outlier detection is a key aspect of the data mining task at hand. Addressing this a novel symbolization method is motivated specifically targeting the end use application of outlier detection. The method is empirically shown to outperform existing approaches. Gavin Smith, James Goulding |
IEEE BigData | 2 |
| 2015 | Towards Computation of Novel Ideas from Corpora of Scientific Text
Haixia Liu 0001, James Goulding, Tim J. Brailsford |
ECML/PKDD (2) | 2 |
| 2014 | A data driven approach to mapping urban neighbourhoodsabstractNeighbourhoods have been described by the UK Secretary of State for Communities and Local Government as the "building blocks of public service society". Despite this, difficulties in data collection combined with the concept's subjective nature have left most countries lacking official neighbourhood definitions. This issue has implications not only for policy, but for the field of computational social science as a whole (with many studies being forced to use administrative units as proxies despite the fact that these bear little connection to resident perceptions of social boundaries). In this paper we illustrate that the mass linguistic datasets now available on the internet need only be combined with relatively simple linguistic computational models to produce definitions that are not only probabilistic and dynamic, but do not require a priori knowledge of neighbourhood names. Paul Brindley, James Goulding, Max L. Wilson 0001 |
SIGSPATIAL/GIS | 2 |
| 2013 | Improving route prediction through user journey detectionabstractThe positioning datasets that underpin route prediction models arrive as time series or point process logs. However, their use for prediction requires them to be split into meaningful segments, conceptualised as travelling periods or 'journeys', to form a set of training inputs. Despite significant research into route prediction, this important pre-processing step has traditionally occurred in an ad-hoc fashion, using arbitrary connectivity or movement thresholds. There has been little consideration to date of the impact of this on prediction, a fact rectified in this work. Mark Dimond, Gavin Smith, James Goulding |
SIGSPATIAL/GIS | 3 |