Tin Vu

dblp:229/8682 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
5since 2021 · last 2024
0000-0003-4088-7091ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021
YearPublicationVenuePosition
2024 A learning-based framework for spatial join processing: estimation, optimization and tuning
abstract
Abstract The importance and complexity of spatial join operation resulted in the availability of many join algorithms, some of which are tailored for big-data platforms like Hadoop and Spark. The choice among them is not trivial and depends on different factors. This paper proposes the first machine-learning-based framework for spatial join query optimization which can accommodate both the characteristics of spatial datasets and the complexity of the different algorithms. The main challenge is how to develop portable cost models that once trained can be applied to any pair of input datasets, because they are able to extract the important input characteristics, such as data distribution and spatial partitioning, the logic of spatial join algorithms, and the relationship between the two input datasets. The proposed system defines a set of features that can be computed efficiently for the data to catch the intricate aspects of spatial join. Then, it uses these features to train five machine learning models that are used to identify the best spatial join algorithm. The first two are regression models that estimate two important measures of the spatial join performance and they act as the cost model. The third model chooses the best partitioning strategy to use with spatial join. The fourth and fifth models further tune two important parameters, number of partitions and plane-sweep direction, to get the best performance. Experiments on large-scale synthetic and real data show the efficiency of the proposed models over baseline methods.
Tin Vu, Alberto Belussi, Sara Migliorini 0001, Ahmed Eldawy
VLDB J.1
2022 Towards a Learned Cost Model for Distributed Spatial Join: Data, Code & Models
abstract
Geospatial data comprise around 60% of all the publicly available data. One of the essential and most complex operations that brings together multiple geospatial datasets is the spatial join operation. Due to its complexity, there is a lot of partitioning techniques and parallel algorithms for the spatial join problem. This leads to a complex query optimization problem: which algorithm to use for a given pair of input datasets that we want to join? With the rise of machine learning, there is a promise in addressing this problem with the use of various learned models. However, one of the concerns is the lack of a standard and publicly available data to train and test on, as well as the lack of accessible baseline models. This resource paper helps the research community to solve this problem by providing synthetic and real datasets for spatial join, source code for constructing more datasets, and several baseline solutions that researchers can further extend and compare to.
Tin Vu, Alberto Belussi, Sara Migliorini 0001, Ahmed Eldawy
CIKM1
2021 Beast: Scalable Exploratory Analytics on Spatio-temporal Data
abstract
This paper introduces the open-source Beast system for scalable exploratory data science on big spatio-temporal data. Beast is based on well-established research and has been released to assist the research community with analyzing big spatio-temporal data. Beast provides a set of extensible components that naturally integrate with Spark to build exploratory data science pipelines. Beast can install in less than a minute on an existing Spark cluster and provides a wide array of features including loading vector and raster data represented in standard file formats, synthetic data generation for benchmarking, load-balanced spatial partitioning, data summarization, interactive visualization, and more. Beast builds on several research projects; its goal is to make all this research widely available to researchers in one integrative and coherent system.
Ahmed Eldawy, Vagelis Hristidis, Saheli Ghosh, Majid Saeedan, Akil Sevim, A. B. Siddique 0001, Samriddhi Singla, Ganesh Sivaram, Tin Vu, Yaming Zhang
CIKM9
2021 A Learned Query Optimizer for Spatial Join
abstract
The importance and complexity of spatial join resulted in many join algorithms, some of which run on big-data platforms such as Hadoop and Spark. This paper proposes the first machine-learning-based query optimizer for spatial join operation which can accommodate the skewness of the spatial datasets and the complexity of the different algorithms. The main challenge is how to develop portable cost models that take into account the important input characteristics such as data distribution, spatial partitioning, logic of spatial join algorithms, and the relationship between the two datasets. The proposed system defines a set of features that can all be computed efficiently for the data to catch the intricate aspects of spatial join. Then, it uses these features to train three machine learning models that capture several metrics to estimate the cost of four spatial join algorithms according to user requirements. The first model can estimate the cardinality of spatial join algorithm. The second model can predict the number of rough comparisons for a specific join algorithm. Finally, the third model is a classification model that can choose the best join algorithm to run. Experiments on large scale synthetic and real data show the efficiency of the proposed models over baseline methods.
Tin Vu, Alberto Belussi, Sara Migliorini 0001, Ahmed Eldawy
SIGSPATIAL/GIS1
2021 Incremental Partitioning for Efficient Spatial Data Analytics
abstract
Big spatial data has become ubiquitous, from mobile applications to satellite data. In most of these applications, data is continuously growing to huge volumes. Existing systems for big spatial data organize records at either the record-level or block-level. Systems that use record-level structures include key-value stores and LSM-Tree stores, which support insert and delete operations and they are optimized for highly-selective queries. On the other hand, systems like GeoSpark that use block-level structures (e.g. 128 MB each) are more efficient for analytical queries, but they cannot incrementally maintain the partitioned data and do not support delete operations. This paper proposes a general framework that enables block-level systems to incrementally maintain spatial partitions, in the presence of bulk insertions and deletions, in distributed file system (DFS) blocks. We first formally study the incremental spatial partitioning problem for big data and demonstrate its NP-hardness. Then, we propose a cost model to estimate the performance of queries on the partitioned data and the effect of modifying it as the data grows. After that, we provide three different implementations of the incremental partitioning framework. Comprehensive experiments on large real datasets show that our proposed partitioning algorithms outperforms state-of-the-art spatial partitioning methods.
Tin Vu, Ahmed Eldawy, Vagelis Hristidis, Vassilis J. Tsotras
Proc. VLDB Endow.1
2020 SpiderWeb: A Spatial Data Generator on the Web
abstract
This demonstration presents a web-based generator for spatial data. This generator allows users to choose from a wide range of spatial data distributions and configure the cardinality of the data and the distribution parameters. It then provides three functionalities. First, it provides a visualization of how the data will look like. Second, it allows users to download this data in several standard formats including CSV and GeoJSON. Third, it provides a permalink that users can bookmark or share with their team members to reproduce the same dataset later. This service is a step towards standardized benchmarking for spatial data systems.
Puloma Katiyar, Tin Vu, Ahmed Eldawy, Sara Migliorini 0001, Alberto Belussi
SIGSPATIAL/GIS2
2020 Noise Prediction for Geocoding Queries using Word Geospatial Embedding and Bidirectional LSTM
abstract
User geocoding queries in map applications often contain noisy tokens such as typos in street, city name, wrong postal code, redundant words due to copy-paste action, etc. This issue becomes worse with the rapid growth of mobile devices, where errors from user input are inevitable. Such noisy tokens may fail the searching process if they are passed as-is to the downstream query processing components. In particular, there might be nothing or irrelevant results returned to the user. Therefore, noisy tokens in geocoding queries should be recognized and handled properly prior to the searching process. In this paper, a deep learning based noise prediction model for geocoding queries is proposed. It combines a novel Word Geospatial Embedding (WGE) and a Bidirectional LSTM based sequence tagging model. The proposed WGE is the first language model that allows geospatial semantics to be encoded into the vector representations. It allows geo-related machine learning/deep learning models making spatial-aware prediction.
Tin Vu, Solluna Liu, Renzhong Wang, Kumarswamy Valegerepura
SIGSPATIAL/GIS1
2019 Deep Query Optimization
abstract
In recent decades, we observed the rapid growth of several big data platforms. Each of them is designed for specific demands. For instance, Spark can efficiently process iterative queries, while Storm is designed for in-memory processing. In this context, the complexity of these distributed systems make it much harder to develop rigorous cost models for query optimization problems. This paper aims to address two problems of the query optimization process: cost estimation and index selection. The cost estimation problem predicts the best execution plan by measuring the cost of alternative query plans. The index selection problem determines the most suitable indexing method with a given dataset. Both problems require the development of a complex function that measures the cost or suitability of alternatives to a specific dataset. Therefore, we employ deep learning to solve those problems due to its capability of learning complicated models. We first address a simple form of cost estimation problem: selectivity estimation. Our preliminary results show that our deep learning models work efficiently with the accuracy of selectivity estimation up to 97%.
Tin Vu
SIGMOD Conference1
2018 R-Grove: growing a family of R-trees in the big-data forest
abstract
The rapid growth of big spatial data urged the research community to develop several big spatial data systems. Regardless of their architecture, one of the fundamental requirements of all these systems is to partition the data efficiently across machines. A widely-used technique for big spatial indexing is to reuse existing search trees asis, e.g., the R-tree family, by building a temporary tree for a sample of the input and use its leaf nodes as partition boundaries. However, we show in this paper that this approach has major limitations that make it unsuitable for the big data environment. This paper studies the use of three popular trees from the R-tree family to index big spatial data, namely, the original R-tree by Guttman, R*-tree, and RR*-tree. We show that the entire family of R-trees is not ready to grow in the big data forest due to fundamental limitations in their design. To overcome these limitations, we propose three new indexes, namely, R-Grove, R*-Grove, and RR*-Grove, which are fundamentally modified to work with big data while inheriting the main characteristics of their traditional index counterparts. With all the proposed indexes publicly available as open source, we hope that these new indexes will be adopted by the community to better serve big spatial data research.
Tin Vu, Ahmed Eldawy
SIGSPATIAL/GIS1