EDBT 2026 Demo / reviewers in the wild / expert
Caitlin Brown
dblp:311/0283
· DBLP profile ↗
3ranked-venue papers in the field
1as first author
3since 2021 · last 2025
0000-0001-8234-3004ORCID · reported
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 3 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large Scale Integrated Simulation of Household Vehicle Fleet Composition with Geographically Explicit Synthetic Population
Naomi Panjaitan, Ling Jin 0001, Caitlin Brown, Tin Ho, Anna Spurlock, Thomas Wenzel, Alina Lazar, Qianmiao Chen, Andrew Bae |
IEEE Big Data | 3 |
| 2022 | Simple and Efficient Identification of Personally Identifiable Information on a Public WebsiteabstractPersonally Identifiable Information (PII) is a key concept in privacy regulation. This form of information can provide revealing information about individuals, which may be collected and used for malicious purposes, such as social engineering and identity theft [1]. Consequently, privacy preserving legislations, such as GDPR place the responsibility of appropriately handling PII onto organisations who may process large amounts of personal data as part of their day to day operations. Therefore, it is necessary to develop processes that can provide privacy assurances but also to do so with increased automation and reliability [2]. This work will focus on assessing the ability of the Natural Language Processing tool, sentiment analysis and text classification algorithms to detect PII automatically, reliably and without too much complexity. To achieve this a dataset containing web pages from Newcastle University’s website, with a focus on staff profiles was created and manually labelled to indicate which sentences contained PII. The dataset was then used to train three text classification algorithms: Multinominal Naïve Bayes, Random Forest Classifier and LSTM model in order to predict the labels of an unseen portion of the dataset. The algorithms all performed well at detecting PII, with Random Forest achieving the highest accuracy at 96% and 96% F1-Score. Nevertheless, the models all mislabelled more sentences containing PII as not containing PII, than those which did not contain PII but were labelled as doing so. Caitlin Brown, Charles Morisset |
IEEE Big Data | 1 |
| 2021 | Performance of the Gold Standard and Machine Learning in Predicting Vehicle TransactionsabstractLogistic regression has long been the gold standard for choice modeling in the transportation field. Despite the rising popularity of machine learning (ML), few is applied to predicting the household vehicle transactions. To address the research gap, this paper presents a first use case of ML application to predicting household vehicle transaction decisions by leveraging a newly processed national panel data set. Model performances are reported for four ML models and the traditional multinomial logit model (MNL). Instead of treating the gold standard and ML models as competitors, this paper tries to use ML tools to inform the MNL model building process. We find the two gradient boosting based methods, CatBoost and LightGBM, are the best performing ML models; and improving logistic models with SHAP interpretation tools can achieve similar performance levels to the best performing ML methods. Alina Lazar, Ling Jin 0001, Caitlin Brown, Anna Spurlock, Alex Sim, Kesheng Wu |
IEEE BigData | 3 |