EDBT 2026 Demo / reviewers in the wild / expert
Alina Lazar
dblp:93/5635
· DBLP profile ↗
13ranked-venue papers in the field
6as first author
5since 2021 · last 2025
0000-0002-2096-1541ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 10 (4 first)Other / Interdisciplinary · 3 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large Scale Integrated Simulation of Household Vehicle Fleet Composition with Geographically Explicit Synthetic Population
Naomi Panjaitan, Ling Jin 0001, Caitlin Brown, Tin Ho, Anna Spurlock, Thomas Wenzel, Alina Lazar, Qianmiao Chen, Andrew Bae |
IEEE Big Data | 7 |
| 2024 | Macroscopic Emission Modeling of Urban Traffic Using Probe Vehicle Data: A Machine Learning ApproachabstractUrban congestions cause inefficient movement of vehicles and exacerbate greenhouse gas emissions and urban air pollution. Macroscopic emission fundamental diagram (eMFD) captures an orderly relationship among emission and aggregated traffic variables at the network level, allowing for real-time monitoring of region-wide emissions and optimal allocation of travel demand to existing networks, reducing urban congestion and associated emissions. However, empirically derived eMFD models are sparse due to historical data limitation. Leveraging a large-scale and granular traffic and emission data derived from probe vehicles, this study is the first to apply machine learning methods to predict the network-wide emission rate to traffic relationship in U.S. urban areas at a large scale. The analysis framework and insights developed in this work generate data-driven eMFDs and a deeper understanding of their location dependence on network, infrastructure, land use, and vehicle characteristics, enabling transportation authorities to measure carbon emissions from urban transport of given travel demand and optimize location-specific traffic management and planning decisions to mitigate network-wide emissions. Mohammed Adlouni, Ling Jin 0001, Xiaodan Xu, Anna Spurlock, Alina Lazar, Kaveh Farokhi Sadabadi, Mahyar Amirgholy, Mona Asudegi |
IEEE Big Data | 5 |
| 2024 | Pruned Graph Neural Networks for Efficient Edge Classification and Fast InferenceabstractOptimizing Graph Neural Networks (GNNs) inference is important for various applications, including High-Energy Physics experiments, natural language processing, and point cloud analysis. Effective optimization techniques, such as pruning, model reduction, automatic mixed precision, and advanced compilation strategies, improve inference speed and memory efficiency significantly. These methods are crucial for deploying GNNs in resource-constrained environments, ensuring scalability and real-time responses. In an experimental setting, using the TrackML dataset and applying these techniques to GNN models yielded promising results. Pruning and model reduction, in particular, emerged as the most impactful methods, leading to a 53% reduction in inference time and a 42% decrease in overall memory usage. These improvements demonstrate the potential of customized optimizations in improving GNN performance, making them more feasible for various real-world applications where computational efficiency is very important. Henry Paschke, James Gaboriault-Whitcomb, Carson Zoccole, Alina Lazar |
IEEE Big Data | 4 |
| 2023 | Leveraging Probe Data and Machine Learning to Derive and Interpret Macroscopic Fundamental Diagrams Across U.S. CitiesabstractMacroscopic fundamental diagram (MFD) captures an orderly relationship among traffic flow, density, and speed at the network level. Understanding network-wide traffic through MFDs can optimally allocate demand to existing networks, improving performance by maximizing network production and avoiding congestion. However, due to historical data limitations, empirically derived MFD models are sparse in the literature, especially for the U.S. cities. Leveraging a large-scale and granular census-tract-level flow and density derived from vehicle probe data, this research is the first to develop a machine learning approach to both derive MFD models and interpret their underlying difference among urban networks across the entire United States. Among the four machine learning methods tested here XGBoost is found to deliver the best performance to predict the network traffic flow for given vehicular density and location attributes. Interaction Shapley Additive explanation (SHAP) values are used to interpret the factors, such as land use, transportation infrastructure, and network topology, that influence the flow-density relationships among locations. The analysis framework developed in this work can generate datadriven MFDs and a deeper understanding of their shape dependence on network, infrastructure, and land use characteristics, which can be used by transportation authorities to derive and optimize location-specific MFDs facilitating more informed management and planning decisions at the network level. Ling Jin 0001, Xiaodan Xu, Kaveh Farokhi Sadabadi, Alina Lazar, Duleep Rathgamage Don, Zachary Needell, Anna Spurlock, Mahyar Amirgholy, Mona Asudegi |
IEEE Big Data | 5 |
| 2021 | Performance of the Gold Standard and Machine Learning in Predicting Vehicle TransactionsabstractLogistic regression has long been the gold standard for choice modeling in the transportation field. Despite the rising popularity of machine learning (ML), few is applied to predicting the household vehicle transactions. To address the research gap, this paper presents a first use case of ML application to predicting household vehicle transaction decisions by leveraging a newly processed national panel data set. Model performances are reported for four ML models and the traditional multinomial logit model (MNL). Instead of treating the gold standard and ML models as competitors, this paper tries to use ML tools to inform the MNL model building process. We find the two gradient boosting based methods, CatBoost and LightGBM, are the best performing ML models; and improving logistic models with SHAP interpretation tools can achieve similar performance levels to the best performing ML methods. Alina Lazar, Ling Jin 0001, Caitlin Brown, Anna Spurlock, Alex Sim, Kesheng Wu |
IEEE BigData | 1 |
| 2019 | Federated Wireless Network Intrusion DetectionabstractWi-Fi has become the wireless networking standard that allows short- to medium-range device to connect without wires. For the last 20 year, the Wi-Fi technology has so pervasive that most devices in use today are mobile and connect to the internet through Wi-Fi. Unlike wired network, a wireless network lacks a clear boundary, which leads to significant Wi-Fi network security concerns, especially because the current security measures are prone to several types of intrusion. To address this problem, machine learning and deep learning methods have been successfully developed to identify network attacks. However, collecting data to develop models is expensive and raises privacy concerns. The goal of this paper is to evaluate a federated learning approach that would alleviate such privacy concerns. This initial work on intrusion detection is performed in a simulated environment. Once proven feasible, this process would allow edge devices to collaboratively update global anomaly detection models, without sharing sensitive training data. On a set of tests with the AWID intrusion detection data set, we show that our federated approach is effective in terms of classification accuracy, computation cost, as well as communication cost. Burak Cetin, Alina Lazar, Jinoh Kim, Alex Sim, Kesheng Wu |
IEEE BigData | 2 |
| 2019 | Machine Learning for Prediction of Mid to Long Term Habitual Transportation Mode UseabstractPrediction of daily transportation mode use (car, public transit, or active travel) is a important task in transportation research. Unlike statistical models that impose a predetermined model structure, machine learning models are learned from the data, making them more flexible with higher prediction accuracy. However, prediction of mid-to long-term habitual modes still largely relies on traditional statistical analysis using small samples of cross-sectional data. Low interpretability of “black-box” machine learning models limits their usefulness for generating behavior insights needed for designing appropriate interventions. This paper, leveraging a set of unique longitudinal life course data, is the first use case to demonstrate machine learning methods applied for both predicting and interpreting regularly used travel modes. We combine sequence clustering and tree-based machine learning methods coupled with TreeExplainer to predict and interpret habitual travel modes using mid-to long-term predictors. Five life course clusters are derived to provide evaluation and interpretation contexts. This allows us to improve upon a recently developed TreeExplainer method to better distinguish predictor importance locally and globally; and predictor interactions across subpopulations within distinctive life history contexts. Our results demonstrate a promising step toward interpretable machine learning applications to mid-to long-term prediction of travel modes for transportation planning. Alina Lazar, Alexandra Ballow, Ling Jin 0001, Anna Spurlock, Alex Sim, Kesheng Wu |
IEEE BigData | 1 |
| 2019 | Understanding Data Similarity in Large-Scale Scientific DatasetsabstractToday, scientific experiments and simulations produce massive amounts of heterogeneous data that need to be stored and analyzed. Given that these large datasets are stored in many files, formats and locations, how can scientists find relevant data, duplicates or similarities? In this context, we concentrate on developing algorithms to compare similarity of time series for the purpose of search, classification and clustering. For example, generating accurate patterns from climate related time series is important not only for building models for weather forecasting and climate prediction, but also for modeling and predicting the cycle of carbon, water, and energy. We developed the methodology and ran an exploratory analysis of climatic and ecosystem variables from the FLUXNET2015 dataset. The proposed combination of similarity metrics, nonlinear dimension reduction, clustering methods and validity measures for time series data has never been applied to unlabeled datasets before, and provides a process that can be easily extended to other scientific time series data. The dimensionality reduction step provides a good way to identify the optimum number of clusters, detect outliers and assign initial labels to the time series data. We evaluated multiple similarity metrics, in terms of the internal cluster validity for driver as well as response variables. While the best metric often depends on a number of factor, the Euclidean distance seems to perform well for most variables and also in terms of computational expense. Payton Linton, William Melodia, Alina Lazar, Deborah A. Agarwal, Ludovico Bianchi, Devarshi Ghoshal, Gilberto Zonta Pastorello, Lavanya Ramakrishnan, Kesheng Wu |
IEEE BigData | 3 |
| 2018 | Predicting Network Traffic Using TCP AnomaliesabstractAccurately predicting network traffic volume is beneficial for congestion control, improving routing, allocating network resources and network optimization. Traffic congestion happens when a network device is receiving more data packets than its processing capability. The number of retransmissions per flow, packet duplication and synthetic reordering can seriously degrade the overall TCP performance. An unsupervised/supervised technique to accurately identify TCP anomalies occurring during file transfers based on passive measurements of TCP traffic collected using Tstat is proposed. This method will be validated on real large datasets collected from several data transfer nodes. The preliminary results indicate that the percentage of TCP anomalies correlate well with the average throughput in any given time window. Alina Lazar, Kesheng Wu, Alex Sim |
IEEE BigData | 1 |
| 2017 | Data quality challenges with missing values and mixed types in joint sequence analysisabstractThe goal of this paper is to investigate the impact of missing values in categorical time series sequences on common data analysis tasks. Being able to more effectively identify patterns in socio-demographic longitudinal data is an important component in a number of social science settings. However, performing fundamental analytical operations, such as clustering for grouping these data based on similarity patterns, is challenging due to the categorical and multi-dimensional nature of the data, and their corruption by missing and inconsistent values. To study these data quality issues, we employ longitudinal sequence data representations, a similarity measure designed for categorical and longitudinal data, together with state-of-the art clustering methodologies reliant on hierarchical algorithms. The key to quantifying the similarity and difference among data records is a distance metric. Given the categorical nature of our data, we employ an “edit” type distance using Optimal Matching (OM). Because each data record has multiple variables of different types, we investigate the impact of mixing these variables in a single similarity measure. Between variables with binary values and those with multiple nominal values, we find that the ability to overcome missing data problems is harder in the nominal domain versus the binary domain. Additionally, artificial clusters introduced by the alignment of leading missing values can be resolved by tuning the missing value substitution cost parameter. Alina Lazar, Ling Jin 0001, Anna Spurlock, Kesheng Wu, Alex Sim |
IEEE BigData | 1 |
| 2016 | Analyzing developer sentiment in commit logsabstractThe paper presents an analysis of developer commit logs for GitHub projects. In particular, developer sentiment in commits is analyzed across 28,466 projects within a seven year time frame. We use the Boa infrastructure's online query system to generate commit logs as well as files that were changed during the commit. We analyze the commits in three categories: large, medium, and small based on the number of commits using a sentiment analysis tool. In addition, we also group the data based on the day of week the commit was made and map the sentiment to the file change history to determine if there was any correlation. Although a majority of the sentiment was neutral, the negative sentiment was about 10% more than the positive sentiment overall. Tuesdays seem to have the most negative sentiment overall. In addition, we do find a strong correlation between the number of files changed and the sentiment expressed by the commits the files were part of. Future work and implications of these results are discussed. Vinayak Sinha, Alina Lazar, Bonita Sharif |
MSR | 2 |
| 2014 | Improving the accuracy of duplicate bug report detection using textual similarity measuresabstractThe paper describes an improved method for automatic duplicate bug report detection based on new textual similarity features and binary classification. Using a set of new textual features, inspired from recent text similarity research, we train several binary classification models. A case study was conducted on three open source systems: Eclipse, Open Office, and Mozilla to determine the effectiveness of the improved method. A comparison is also made with current state-of-the-art approaches highlighting similarities and differences. Results indicate that the accuracy of the proposed method is better than previously reported research with respect to all three systems. Alina Lazar, Sarah Ritchey, Bonita Sharif |
MSR | 1 |
| 2014 | Generating duplicate bug datasetsabstractAutomatic identification of duplicate bug reports is an important research problem in the mining software repositories field. This paper presents a collection of bug datasets collected, cleaned and preprocessed for the duplicate bug report identification problem. The datasets were extracted from open-source systems that use Bugzilla as their bug tracking component and contain all the bugs ever submitted. The systems used are Eclipse, Open Office, NetBeans and Mozilla. For each dataset, we store both the initial data and the cleaned data in separate collections in a mongoDB document-oriented database. For each dataset, in addition to the bug data collections downloaded from bug repositories, the database includes a set of all pairs of duplicate bugs together with randomly selected pairs of non-duplicate bugs. Such a dataset is useful as input for classification models and forms a good base to support replications and comparisons by other researchers. We used a subset of this data to predict duplicate bug reports but the same data set may also be used to predict bug priorities and severity. Alina Lazar, Sarah Ritchey, Bonita Sharif |
MSR | 1 |