Luan V. Tran

dblp:295/3467 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
2since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 6 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data mining · 40% Spatial and temporal data management · 25% Data stream processing · 20%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Smart cities and intelligent transportation · 64% Computational social science and digital humanities · 36%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 50% Ubiquitous computing and smart environments · 50%

Topics — the 10 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
anomaly detection
0.722020
Real-Time Distance-Based Outlier Detection in Data Streams · Proc. VLDB Endow. 2020
Distance-based Outlier Detection in Data Streams · Proc. VLDB Endow. 2016
Data mining › anomaly detection › outlier detection
distance-based outlier detection
0.722020
Real-Time Distance-Based Outlier Detection in Data Streams · Proc. VLDB Endow. 2020
Distance-based Outlier Detection in Data Streams · Proc. VLDB Endow. 2016
Data stream processing › stream mining
streaming outlier detection
0.722020
Real-Time Distance-Based Outlier Detection in Data Streams · Proc. VLDB Endow. 2020
Distance-based Outlier Detection in Data Streams · Proc. VLDB Endow. 2016
Data integration and cleaning › heterogeneous data integration
multimodal data integration
0.512021
Crosstown Foundry: A Scalable Data-driven Journalism Platform for Hyper-local News · SIGMOD Conference 2021
Smart cities and intelligent transportation › public transit
bus travel time prediction
0.412020
DeepTRANS: A Deep Learning System for Public Bus Travel Time Estimation using Traffic Forecasting · Proc. VLDB Endow. 2020
Smart cities and intelligent transportation
public transit
0.412020
DeepTRANS: A Deep Learning System for Public Bus Travel Time Estimation using Traffic Forecasting · Proc. VLDB Endow. 2020
Spatial and temporal data management
trajectory data
0.412020
DeepTRANS: A Deep Learning System for Public Bus Travel Time Estimation using Traffic Forecasting · Proc. VLDB Endow. 2020
Spatial and temporal data management › trajectory analysis
travel time estimation
0.412020
DeepTRANS: A Deep Learning System for Public Bus Travel Time Estimation using Traffic Forecasting · Proc. VLDB Endow. 2020
Collaborative and social computing › cooperative work
task allocation
0.212016
Real-time task assignment in hyperlocal spatial crowdsourcing under budget constraints · PerCom 2016
Performance modeling and evaluation
benchmarking
0.112016
Distance-based Outlier Detection in Data Streams · Proc. VLDB Endow. 2016

Methods — techniques the papers use, named apart from their topics

traffic forecasting · 1.3deep learning · 1.3personalized newsletter generation · 1.0multi-distance indexing · 0.9core point indexing · 0.9comparative evaluation · 0.5online algorithms · 0.2heuristic algorithm · 0.2
YearPublicationVenuePosition
2021 Clustering Mixed-Type Data with Correlation-Preserving Embedding
Luan V. Tran, Liyue Fan, Cyrus Shahabi
DASFAA (2)1
2021 Crosstown Foundry: A Scalable Data-driven Journalism Platform for Hyper-local News
abstract
Generating hyper-local news at scale is challenging because publicly available data is not provided at the desired spatial and temporal granularity. Besides, there is a lack of automated analytical and publishing tools. Crosstown Foundry, which is being actively developed and used by engineers and journalists, is a novel data-driven system that leverages a massive multi-modal dataset to generate personalized newsletters for Los Angeles County readers.
Luciano Nocera, George Constantinou, Luan V. Tran, Seon Ho Kim, Gabriel Kahn, Cyrus Shahabi
SIGMOD Conference3
2020 DeepTRANS: A Deep Learning System for Public Bus Travel Time Estimation using Traffic Forecasting
abstract
In the public transportation domain, accurate estimation of travel times helps to manage rider expectations as well as to provide a powerful tool for transportation agencies to coordinate the public transport vehicles. Although many statistical and machine learning methods have been proposed to estimate travel times, none of the methods consider utilizing predicted traffic information. Forecasting how congestion is going to evolve is critical for accurate travel time estimations. In this paper, we present DeepTRANS, which incorporates traffic forecasting information to our prior Deep Learning-based Bus Estimated Time of Arrival (ETA) model, increasing its accuracy by 21% in estimating bus travel time.
Luan V. Tran, Minyoung Mun, Matthew Lim, Jonah Yamato, Nathan Huh, Cyrus Shahabi
Proc. VLDB Endow.1
2020 Real-Time Distance-Based Outlier Detection in Data Streams
abstract
Real-time outlier detection in data streams has drawn much attention recently as many applications need to be able to detect abnormal behaviors as soon as they occur. The arrival and departure of streaming data on edge devices impose new challenges to process the data quickly in real-time due to memory and CPU limitations of these devices. Existing methods are slow and not memory efficient as they mostly focus on quick detection of inliers and pay less attention to expediting neighbor searches for outlier candidates. In this study, we propose a new algorithm, CPOD, to improve the efficiency of outlier detections while reducing its memory requirements. CPOD uses a unique data structure called "core point" with multi-distance indexing to both quickly identify inliers and reduce neighbor search spaces for outlier candidates. We show that with six real-world and one synthetic dataset, CPOD is, on average, 10, 19, and 73 times faster than M_MCOD, NETS, and MCOD, respectively, while consuming low memory.
Luan V. Tran, Minyoung Mun, Cyrus Shahabi
Proc. VLDB Endow.1
2019 Outlier Detection in Non-stationary Data Streams
abstract
Continuous outlier detection in data streams is an important topic in data mining and has applications in various domains such as fraud detection, weather analysis, and intrusion detection. The non-stationary characteristic of real-world data streams brings the challenge of updating the outlier detection model in a timely and accurate manner. In this paper, we propose a framework for outlier detection in non-stationary data streams (O-NSD) which detects changes in the underlying data distribution to trigger a model update. We propose an improved distance function between sliding windows which offers a monotonicity property; we develop two accurate change detection algorithms, one of which is parameter-free; and we further propose new evaluation measures that quantify the timeliness of the detected changes. Our extensive experiments with real-world and synthetic datasets show that our change detection algorithms outperform the state-of-the-art solution. In addition, we demonstrate our O-NSD framework with two popular unsupervised outlier classifiers. Empirical results show that our framework offers higher accuracy and requires a much lower running time, compared to retrain-based and incremental update approaches.
Luan V. Tran, Liyue Fan, Cyrus Shahabi
SSDBM1
2018 A Real-Time Framework for Task Assignment in Hyperlocal Spatial Crowdsourcing
abstract
Spatial Crowdsourcing (SC) is a novel platform that engages individuals in the act of collecting various types of spatial data. This method of data collection can significantly reduce cost and turnover time and is particularly useful in urban environmental sensing, where traditional means fail to provide fine-grained field data. In this study, we introduce hyperlocal spatial crowdsourcing, where all workers who are located within the spatiotemporal vicinity of a task are eligible to perform the task (e.g., reporting the precipitation level at their area and time). In this setting, there is often a budget constraint, either for every time period or for the entire campaign, on the number of workers to activate to perform tasks. The challenge is thus to maximize the number of assigned tasks under the budget constraint despite the dynamic arrivals of workers and tasks. We introduce a taxonomy of several problem variants, such as budget-per-time-period vs. budget-per-campaign and binary-utility vs. distance-based-utility . We study the hardness of the task assignment problem in the offline setting and propose online heuristics which exploit the spatial and temporal knowledge acquired over time. Our experiments are conducted with spatial crowdsourcing workloads generated by the SCAWG tool, and extensive results show the effectiveness and efficiency of our proposed solutions.
Luan V. Tran, Hien To, Liyue Fan, Cyrus Shahabi
ACM Trans. Intell. Syst. Technol.1
2016 Real-time task assignment in hyperlocal spatial crowdsourcing under budget constraints
abstract
Spatial Crowdsourcing (SC) is a novel platform that engages individuals in the act of collecting various types of spatial data. This method of data collection can significantly reduce cost and turnover time, and is particularly useful in environmental sensing, where traditional means fail to provide fine-grained field data. In this study, we introduce hyperlocal spatial crowdsourcing, where all workers who are located within the spatiotemporal vicinity of a task are eligible to perform the task, e.g., reporting the precipitation level at their area and time. In this setting, there is often a budget constraint, either for every time period or for the entire campaign, on the number of workers to activate to perform tasks. The challenge is thus to maximize the number of assigned tasks under the budget constraint, despite the dynamic arrivals of workers and tasks as well as their co-location relationship. We study two problem variants in this paper: budget is constrained for every timestamp, i.e. fixed, and budget is constrained for the entire campaign, i.e. dynamic. For each variant, we study the complexity of its offline version and then propose several heuristics for the online version which exploit the spatial and temporal knowledge acquired over time. Extensive experiments with real-world and synthetic datasets show the effectiveness and efficiency of our proposed solutions.
Hien To, Liyue Fan, Luan V. Tran, Cyrus Shahabi
PerCom3
2016 Distance-based Outlier Detection in Data Streams
abstract
Continuous outlier detection in data streams has important applications in fraud detection, network security, and public health. The arrival and departure of data objects in a streaming manner impose new challenges for outlier detection algorithms, especially in time and space efficiency. In the past decade, several studies have been performed to address the problem of distance-based outlier detection in data streams (DODDS), which adopts an unsupervised definition and does not have any distributional assumptions on data values. Our work is motivated by the lack of comparative evaluation among the state-of-the-art algorithms using the same datasets on the same platform. We systematically evaluate the most recent algorithms for DODDS under various stream settings and outlier rates. Our extensive results show that in most settings, the MCOD algorithm offers the superior performance among all the algorithms, including the most recent algorithm Thresh_LEAP.
Luan V. Tran, Liyue Fan, Cyrus Shahabi
Proc. VLDB Endow.1