VLDB 2026 Research / reviewers in the wild / expert
Luan V. Tran
dblp:295/3467
· DBLP profile ↗
8ranked-venue papers
6as first author
2since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 6 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Data mining · 40% Spatial and temporal data management · 25% Data stream processing · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Smart cities and intelligent transportation · 64% Computational social science and digital humanities · 36% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 50% Ubiquitous computing and smart environments · 50% |
Topics — the 10 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
anomaly detection |
0.7 | 2 | 2020 | Real-Time Distance-Based Outlier Detection in Data Streams · Proc. VLDB Endow. 2020 Distance-based Outlier Detection in Data Streams · Proc. VLDB Endow. 2016 |
Data mining › anomaly detection › outlier detection
distance-based outlier detection |
0.7 | 2 | 2020 | Real-Time Distance-Based Outlier Detection in Data Streams · Proc. VLDB Endow. 2020 Distance-based Outlier Detection in Data Streams · Proc. VLDB Endow. 2016 |
Data stream processing › stream mining
streaming outlier detection |
0.7 | 2 | 2020 | Real-Time Distance-Based Outlier Detection in Data Streams · Proc. VLDB Endow. 2020 Distance-based Outlier Detection in Data Streams · Proc. VLDB Endow. 2016 |
Data integration and cleaning › heterogeneous data integration
multimodal data integration |
0.5 | 1 | 2021 | Crosstown Foundry: A Scalable Data-driven Journalism Platform for Hyper-local News · SIGMOD Conference 2021 |
Smart cities and intelligent transportation › public transit
bus travel time prediction |
0.4 | 1 | 2020 | DeepTRANS: A Deep Learning System for Public Bus Travel Time Estimation using Traffic Forecasting · Proc. VLDB Endow. 2020 |
Smart cities and intelligent transportation
public transit |
0.4 | 1 | 2020 | DeepTRANS: A Deep Learning System for Public Bus Travel Time Estimation using Traffic Forecasting · Proc. VLDB Endow. 2020 |
Spatial and temporal data management
trajectory data |
0.4 | 1 | 2020 | DeepTRANS: A Deep Learning System for Public Bus Travel Time Estimation using Traffic Forecasting · Proc. VLDB Endow. 2020 |
Spatial and temporal data management › trajectory analysis
travel time estimation |
0.4 | 1 | 2020 | DeepTRANS: A Deep Learning System for Public Bus Travel Time Estimation using Traffic Forecasting · Proc. VLDB Endow. 2020 |
Collaborative and social computing › cooperative work
task allocation |
0.2 | 1 | 2016 | Real-time task assignment in hyperlocal spatial crowdsourcing under budget constraints · PerCom 2016 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2016 | Distance-based Outlier Detection in Data Streams · Proc. VLDB Endow. 2016 |
Methods — techniques the papers use, named apart from their topics
traffic forecasting · 1.3deep learning · 1.3personalized newsletter generation · 1.0multi-distance indexing · 0.9core point indexing · 0.9comparative evaluation · 0.5online algorithms · 0.2heuristic algorithm · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Clustering Mixed-Type Data with Correlation-Preserving Embedding
Luan V. Tran, Liyue Fan, Cyrus Shahabi |
DASFAA (2) | 1 |
| 2021 | Crosstown Foundry: A Scalable Data-driven Journalism Platform for Hyper-local NewsabstractGenerating hyper-local news at scale is challenging because publicly available data is not provided at the desired spatial and temporal granularity. Besides, there is a lack of automated analytical and publishing tools. Crosstown Foundry, which is being actively developed and used by engineers and journalists, is a novel data-driven system that leverages a massive multi-modal dataset to generate personalized newsletters for Los Angeles County readers. Luciano Nocera, George Constantinou, Luan V. Tran, Seon Ho Kim, Gabriel Kahn, Cyrus Shahabi |
SIGMOD Conference | 3 |
| 2020 | DeepTRANS: A Deep Learning System for Public Bus Travel Time Estimation using Traffic ForecastingabstractIn the public transportation domain, accurate estimation of travel times helps to manage rider expectations as well as to provide a powerful tool for transportation agencies to coordinate the public transport vehicles. Although many statistical and machine learning methods have been proposed to estimate travel times, none of the methods consider utilizing predicted traffic information. Forecasting how congestion is going to evolve is critical for accurate travel time estimations. In this paper, we present DeepTRANS, which incorporates traffic forecasting information to our prior Deep Learning-based Bus Estimated Time of Arrival (ETA) model, increasing its accuracy by 21% in estimating bus travel time. Luan V. Tran, Minyoung Mun, Matthew Lim, Jonah Yamato, Nathan Huh, Cyrus Shahabi |
Proc. VLDB Endow. | 1 |
| 2020 | Real-Time Distance-Based Outlier Detection in Data StreamsabstractReal-time outlier detection in data streams has drawn much attention recently as many applications need to be able to detect abnormal behaviors as soon as they occur. The arrival and departure of streaming data on edge devices impose new challenges to process the data quickly in real-time due to memory and CPU limitations of these devices. Existing methods are slow and not memory efficient as they mostly focus on quick detection of inliers and pay less attention to expediting neighbor searches for outlier candidates. In this study, we propose a new algorithm, CPOD, to improve the efficiency of outlier detections while reducing its memory requirements. CPOD uses a unique data structure called "core point" with multi-distance indexing to both quickly identify inliers and reduce neighbor search spaces for outlier candidates. We show that with six real-world and one synthetic dataset, CPOD is, on average, 10, 19, and 73 times faster than M_MCOD, NETS, and MCOD, respectively, while consuming low memory. Luan V. Tran, Minyoung Mun, Cyrus Shahabi |
Proc. VLDB Endow. | 1 |
| 2019 | Outlier Detection in Non-stationary Data StreamsabstractContinuous outlier detection in data streams is an important topic in data mining and has applications in various domains such as fraud detection, weather analysis, and intrusion detection. The non-stationary characteristic of real-world data streams brings the challenge of updating the outlier detection model in a timely and accurate manner. In this paper, we propose a framework for outlier detection in non-stationary data streams (O-NSD) which detects changes in the underlying data distribution to trigger a model update. We propose an improved distance function between sliding windows which offers a monotonicity property; we develop two accurate change detection algorithms, one of which is parameter-free; and we further propose new evaluation measures that quantify the timeliness of the detected changes. Our extensive experiments with real-world and synthetic datasets show that our change detection algorithms outperform the state-of-the-art solution. In addition, we demonstrate our O-NSD framework with two popular unsupervised outlier classifiers. Empirical results show that our framework offers higher accuracy and requires a much lower running time, compared to retrain-based and incremental update approaches. Luan V. Tran, Liyue Fan, Cyrus Shahabi |
SSDBM | 1 |
| 2018 | A Real-Time Framework for Task Assignment in Hyperlocal Spatial CrowdsourcingabstractSpatial Crowdsourcing (SC) is a novel platform that engages individuals in the act of collecting various types of spatial data. This method of data collection can significantly reduce cost and turnover time and is particularly useful in urban environmental sensing, where traditional means fail to provide fine-grained field data. In this study, we introduce hyperlocal spatial crowdsourcing, where all workers who are located within the spatiotemporal vicinity of a task are eligible to perform the task (e.g., reporting the precipitation level at their area and time). In this setting, there is often a budget constraint, either for every time period or for the entire campaign, on the number of workers to activate to perform tasks. The challenge is thus to maximize the number of assigned tasks under the budget constraint despite the dynamic arrivals of workers and tasks. We introduce a taxonomy of several problem variants, such as budget-per-time-period vs. budget-per-campaign and binary-utility vs. distance-based-utility . We study the hardness of the task assignment problem in the offline setting and propose online heuristics which exploit the spatial and temporal knowledge acquired over time. Our experiments are conducted with spatial crowdsourcing workloads generated by the SCAWG tool, and extensive results show the effectiveness and efficiency of our proposed solutions. Luan V. Tran, Hien To, Liyue Fan, Cyrus Shahabi |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2016 | Real-time task assignment in hyperlocal spatial crowdsourcing under budget constraintsabstractSpatial Crowdsourcing (SC) is a novel platform that engages individuals in the act of collecting various types of spatial data. This method of data collection can significantly reduce cost and turnover time, and is particularly useful in environmental sensing, where traditional means fail to provide fine-grained field data. In this study, we introduce hyperlocal spatial crowdsourcing, where all workers who are located within the spatiotemporal vicinity of a task are eligible to perform the task, e.g., reporting the precipitation level at their area and time. In this setting, there is often a budget constraint, either for every time period or for the entire campaign, on the number of workers to activate to perform tasks. The challenge is thus to maximize the number of assigned tasks under the budget constraint, despite the dynamic arrivals of workers and tasks as well as their co-location relationship. We study two problem variants in this paper: budget is constrained for every timestamp, i.e. fixed, and budget is constrained for the entire campaign, i.e. dynamic. For each variant, we study the complexity of its offline version and then propose several heuristics for the online version which exploit the spatial and temporal knowledge acquired over time. Extensive experiments with real-world and synthetic datasets show the effectiveness and efficiency of our proposed solutions. Hien To, Liyue Fan, Luan V. Tran, Cyrus Shahabi |
PerCom | 3 |
| 2016 | Distance-based Outlier Detection in Data StreamsabstractContinuous outlier detection in data streams has important applications in fraud detection, network security, and public health. The arrival and departure of data objects in a streaming manner impose new challenges for outlier detection algorithms, especially in time and space efficiency. In the past decade, several studies have been performed to address the problem of distance-based outlier detection in data streams (DODDS), which adopts an unsupervised definition and does not have any distributional assumptions on data values. Our work is motivated by the lack of comparative evaluation among the state-of-the-art algorithms using the same datasets on the same platform. We systematically evaluate the most recent algorithms for DODDS under various stream settings and outlier rates. Our extensive results show that in most settings, the MCOD algorithm offers the superior performance among all the algorithms, including the most recent algorithm Thresh_LEAP. Luan V. Tran, Liyue Fan, Cyrus Shahabi |
Proc. VLDB Endow. | 1 |