VLDB 2026 Research / reviewers in the wild / expert
Ling Jin 0001
dblp:04/5481-1
· DBLP profile ↗
6ranked-venue papers in the field
1as first author
4since 2021 · last 2025
0000-0002-4381-195XORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 6 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large Scale Integrated Simulation of Household Vehicle Fleet Composition with Geographically Explicit Synthetic Population
Naomi Panjaitan, Ling Jin 0001, Caitlin Brown, Tin Ho, Anna Spurlock, Thomas Wenzel, Alina Lazar, Qianmiao Chen, Andrew Bae |
IEEE Big Data | 2 |
| 2024 | Macroscopic Emission Modeling of Urban Traffic Using Probe Vehicle Data: A Machine Learning ApproachabstractUrban congestions cause inefficient movement of vehicles and exacerbate greenhouse gas emissions and urban air pollution. Macroscopic emission fundamental diagram (eMFD) captures an orderly relationship among emission and aggregated traffic variables at the network level, allowing for real-time monitoring of region-wide emissions and optimal allocation of travel demand to existing networks, reducing urban congestion and associated emissions. However, empirically derived eMFD models are sparse due to historical data limitation. Leveraging a large-scale and granular traffic and emission data derived from probe vehicles, this study is the first to apply machine learning methods to predict the network-wide emission rate to traffic relationship in U.S. urban areas at a large scale. The analysis framework and insights developed in this work generate data-driven eMFDs and a deeper understanding of their location dependence on network, infrastructure, land use, and vehicle characteristics, enabling transportation authorities to measure carbon emissions from urban transport of given travel demand and optimize location-specific traffic management and planning decisions to mitigate network-wide emissions. Mohammed Adlouni, Ling Jin 0001, Xiaodan Xu, Anna Spurlock, Alina Lazar, Kaveh Farokhi Sadabadi, Mahyar Amirgholy, Mona Asudegi |
IEEE Big Data | 2 |
| 2023 | Leveraging Probe Data and Machine Learning to Derive and Interpret Macroscopic Fundamental Diagrams Across U.S. CitiesabstractMacroscopic fundamental diagram (MFD) captures an orderly relationship among traffic flow, density, and speed at the network level. Understanding network-wide traffic through MFDs can optimally allocate demand to existing networks, improving performance by maximizing network production and avoiding congestion. However, due to historical data limitations, empirically derived MFD models are sparse in the literature, especially for the U.S. cities. Leveraging a large-scale and granular census-tract-level flow and density derived from vehicle probe data, this research is the first to develop a machine learning approach to both derive MFD models and interpret their underlying difference among urban networks across the entire United States. Among the four machine learning methods tested here XGBoost is found to deliver the best performance to predict the network traffic flow for given vehicular density and location attributes. Interaction Shapley Additive explanation (SHAP) values are used to interpret the factors, such as land use, transportation infrastructure, and network topology, that influence the flow-density relationships among locations. The analysis framework developed in this work can generate datadriven MFDs and a deeper understanding of their shape dependence on network, infrastructure, and land use characteristics, which can be used by transportation authorities to derive and optimize location-specific MFDs facilitating more informed management and planning decisions at the network level. Ling Jin 0001, Xiaodan Xu, Kaveh Farokhi Sadabadi, Alina Lazar, Duleep Rathgamage Don, Zachary Needell, Anna Spurlock, Mahyar Amirgholy, Mona Asudegi |
IEEE Big Data | 1 |
| 2021 | Performance of the Gold Standard and Machine Learning in Predicting Vehicle TransactionsabstractLogistic regression has long been the gold standard for choice modeling in the transportation field. Despite the rising popularity of machine learning (ML), few is applied to predicting the household vehicle transactions. To address the research gap, this paper presents a first use case of ML application to predicting household vehicle transaction decisions by leveraging a newly processed national panel data set. Model performances are reported for four ML models and the traditional multinomial logit model (MNL). Instead of treating the gold standard and ML models as competitors, this paper tries to use ML tools to inform the MNL model building process. We find the two gradient boosting based methods, CatBoost and LightGBM, are the best performing ML models; and improving logistic models with SHAP interpretation tools can achieve similar performance levels to the best performing ML methods. Alina Lazar, Ling Jin 0001, Caitlin Brown, Anna Spurlock, Alex Sim, Kesheng Wu |
IEEE BigData | 2 |
| 2019 | Machine Learning for Prediction of Mid to Long Term Habitual Transportation Mode UseabstractPrediction of daily transportation mode use (car, public transit, or active travel) is a important task in transportation research. Unlike statistical models that impose a predetermined model structure, machine learning models are learned from the data, making them more flexible with higher prediction accuracy. However, prediction of mid-to long-term habitual modes still largely relies on traditional statistical analysis using small samples of cross-sectional data. Low interpretability of “black-box” machine learning models limits their usefulness for generating behavior insights needed for designing appropriate interventions. This paper, leveraging a set of unique longitudinal life course data, is the first use case to demonstrate machine learning methods applied for both predicting and interpreting regularly used travel modes. We combine sequence clustering and tree-based machine learning methods coupled with TreeExplainer to predict and interpret habitual travel modes using mid-to long-term predictors. Five life course clusters are derived to provide evaluation and interpretation contexts. This allows us to improve upon a recently developed TreeExplainer method to better distinguish predictor importance locally and globally; and predictor interactions across subpopulations within distinctive life history contexts. Our results demonstrate a promising step toward interpretable machine learning applications to mid-to long-term prediction of travel modes for transportation planning. Alina Lazar, Alexandra Ballow, Ling Jin 0001, Anna Spurlock, Alex Sim, Kesheng Wu |
IEEE BigData | 3 |
| 2017 | Data quality challenges with missing values and mixed types in joint sequence analysisabstractThe goal of this paper is to investigate the impact of missing values in categorical time series sequences on common data analysis tasks. Being able to more effectively identify patterns in socio-demographic longitudinal data is an important component in a number of social science settings. However, performing fundamental analytical operations, such as clustering for grouping these data based on similarity patterns, is challenging due to the categorical and multi-dimensional nature of the data, and their corruption by missing and inconsistent values. To study these data quality issues, we employ longitudinal sequence data representations, a similarity measure designed for categorical and longitudinal data, together with state-of-the art clustering methodologies reliant on hierarchical algorithms. The key to quantifying the similarity and difference among data records is a distance metric. Given the categorical nature of our data, we employ an “edit” type distance using Optimal Matching (OM). Because each data record has multiple variables of different types, we investigate the impact of mixing these variables in a single similarity measure. Between variables with binary values and those with multiple nominal values, we find that the ability to overcome missing data problems is harder in the nominal domain versus the binary domain. Additionally, artificial clusters introduced by the alignment of leading missing values can be resolved by tuning the missing value substitution cost parameter. Alina Lazar, Ling Jin 0001, Anna Spurlock, Kesheng Wu, Alex Sim |
IEEE BigData | 2 |