EDBT 2026 Demo / reviewers in the wild / expert
Gavin Long
dblp:231/4400
· DBLP profile ↗
2ranked-venue papers in the field
1as first author
2since 2021 · last 2023
0000-0002-3142-2201ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Who consumes anthocyanins and anthocyanidins? Mining national retail data to reveal the influence of socioeconomic deprivation and seasonality on polyphenol dietary intakeabstractAnthocyanins are a class of polyphenols that have received widespread recent attention due to their potential health benefits. However, estimating the dietary intake of anthocyanins at a population level is a challenging task, due to the difficulty of scaling dietary surveys. Further, there is limited evidence as to who regularly consumes anthocyanins, whether temporally, spatially, or culturally according to levels of socioeconomic deprivation. Leveraging a massive retail loyalty card dataset in the UK, we pair two years of real-world purchasing data for 619,524 regular shoppers and 207 million shopping baskets with anthocyanin estimates drawn from polyphenol databases. We subsequently analyse relative deprivation levels of the neighbourhoods in which shoppers reside, illustrating how anthocyanin intake varies according to affluence. Results indicate that deprivation is linked dramatically with both lower total intake of anthocyanins and lower breadth of dietary sources for them, potentially aggravating the incidence of diet-related diseases in the poorest sections of society. Gavin Long, Roberto Mansilla, Simon Welham, Peter Rose, Michelle Thomas, Gregor Milligan, Elizabeth Dolan, Joanne Parkes, Kuzivakwashe Makokoro, James Goulding |
IEEE Big Data | 2 |
| 2022 | Privacy-preserving & machine-learned catchment models for national dietary surveillance via digital footprint dataabstractBig data from food retail stores is increasingly being used for population dietary surveillance, epidemiological studies of diet-related diseases, and evaluations of public health interventions. However, for retail data to be useful it is necessary to understand the spatio-temporal variation of when and where food is purchased and consumed. While some customers willingly share home location data with retailers as part of loyalty programs such data is typically too fine-grained/sensitive to be applied for research purposes. The aim of this study was to analyse differences between privacy-preserving models and actual retail catchments, and investigate if machine learning techniques could improve the accuracy of such catchment models. Based on a UK-wide sample of 4 million grocery store loyalty card holders, covering 485 million transactions over 29 months (2019-2021) and distributed across 33,000 neighbourhoods (Lower Super Output Areas, or LSOA), the study demonstrates how models trained on geolocated data perform at predicting, per store, catchment areas which contain 50, 80, and 95% of its customers’ primary location. Through comparative assessment of machine learning approaches, we find better performance from tree-based models (RF, XGB) with the best performance from an XGB model achieving an R2of 0.72 and MAE of 1.06. To conclude, we review variable importance measures using SHAP values and discuss the relative merits of including specific features when modeling catchment areas. Gavin Long, Gavin Smith, Georgiana Nica-Avram, Gregor Engelmann, James Goulding |
IEEE Big Data | 1 |