Caitlin Brown

dblp:311/0283 · DBLP profile ↗
← Back
3ranked-venue papers in the field
1as first author
3since 2021 · last 2025
0000-0001-8234-3004ORCID · reported

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3 (1 first)
YearPublicationVenuePosition
2025 Large Scale Integrated Simulation of Household Vehicle Fleet Composition with Geographically Explicit Synthetic Population
Naomi Panjaitan, Ling Jin 0001, Caitlin Brown, Tin Ho, Anna Spurlock, Thomas Wenzel, Alina Lazar, Qianmiao Chen, Andrew Bae
IEEE Big Data3
2022 Simple and Efficient Identification of Personally Identifiable Information on a Public Website
abstract
Personally Identifiable Information (PII) is a key concept in privacy regulation. This form of information can provide revealing information about individuals, which may be collected and used for malicious purposes, such as social engineering and identity theft [1]. Consequently, privacy preserving legislations, such as GDPR place the responsibility of appropriately handling PII onto organisations who may process large amounts of personal data as part of their day to day operations. Therefore, it is necessary to develop processes that can provide privacy assurances but also to do so with increased automation and reliability [2]. This work will focus on assessing the ability of the Natural Language Processing tool, sentiment analysis and text classification algorithms to detect PII automatically, reliably and without too much complexity. To achieve this a dataset containing web pages from Newcastle University’s website, with a focus on staff profiles was created and manually labelled to indicate which sentences contained PII. The dataset was then used to train three text classification algorithms: Multinominal Naïve Bayes, Random Forest Classifier and LSTM model in order to predict the labels of an unseen portion of the dataset. The algorithms all performed well at detecting PII, with Random Forest achieving the highest accuracy at 96% and 96% F1-Score. Nevertheless, the models all mislabelled more sentences containing PII as not containing PII, than those which did not contain PII but were labelled as doing so.
Caitlin Brown, Charles Morisset
IEEE Big Data1
2021 Performance of the Gold Standard and Machine Learning in Predicting Vehicle Transactions
abstract
Logistic regression has long been the gold standard for choice modeling in the transportation field. Despite the rising popularity of machine learning (ML), few is applied to predicting the household vehicle transactions. To address the research gap, this paper presents a first use case of ML application to predicting household vehicle transaction decisions by leveraging a newly processed national panel data set. Model performances are reported for four ML models and the traditional multinomial logit model (MNL). Instead of treating the gold standard and ML models as competitors, this paper tries to use ML tools to inform the MNL model building process. We find the two gradient boosting based methods, CatBoost and LightGBM, are the best performing ML models; and improving logistic models with SHAP interpretation tools can achieve similar performance levels to the best performing ML methods.
Alina Lazar, Ling Jin 0001, Caitlin Brown, Anna Spurlock, Alex Sim, Kesheng Wu
IEEE BigData3