Hoang Tam Vo

dblp:32/8510 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 4 first-authorArtificial intelligence and machine learning · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Distributed and cloud data management · 29% Query processing and optimization · 29% Indexing and storage engines · 14%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Cloud and datacenter computing · 74% Distributed systems · 21% Parallel and multicore computing · 5%

Topics — the 15 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed and cloud data management
peer-to-peer data management
0.322014
BestPeer++: A Peer-to-Peer BasedLarge-Scale Data Processing Platform · IEEE Trans. Knowl. Data Eng. 2014
BestPeer++: A Peer-to-Peer Based Large-Scale Data Processing Platform · ICDE 2012
Data stream processing
complex event processing
0.312017
Multi-Query Optimization for Complex Event Processing in SAP ESP · ICDE 2017
Query processing and optimization
multi-query optimization
0.312017
Multi-Query Optimization for Complex Event Processing in SAP ESP · ICDE 2017
Query processing and optimization › query optimization
cost-based optimization
0.212014
ScalaGiST: Scalable Generalized Search Trees for MapReduce Systems [Innovative Systems Paper] · Proc. VLDB Endow. 2014
Query processing and optimization
database access optimization
0.212014
ScalaGiST: Scalable Generalized Search Trees for MapReduce Systems [Innovative Systems Paper] · Proc. VLDB Endow. 2014
Indexing and storage engines › tree index
generalized search tree
0.212014
ScalaGiST: Scalable Generalized Search Trees for MapReduce Systems [Innovative Systems Paper] · Proc. VLDB Endow. 2014
Distributed and cloud data management
cloud database
0.222012
A Framework for Supporting DBMS-like Indexes in the Cloud · Proc. VLDB Endow. 2011
LogBase: A Scalable Log-structured Database System in the Cloud · Proc. VLDB Endow. 2012
Indexing and storage engines
index management
0.112011
A Framework for Supporting DBMS-like Indexes in the Cloud · Proc. VLDB Endow. 2011
Cloud and datacenter computing
cloud storage
0.112011
ES2: A cloud data storage system for supporting both OLTP and OLAP · ICDE 2011
Distributed and cloud data management › distributed data store
cloud storage
0.112010
Towards Elastic Transactional Cloud Storage with Range Query Support · Proc. VLDB Endow. 2010
Distributed systems
peer-to-peer systems
0.112014
BestPeer++: A Peer-to-Peer BasedLarge-Scale Data Processing Platform · IEEE Trans. Knowl. Data Eng. 2014
Transaction processing and concurrency control › concurrency control
multiversion concurrency control
0.012012
LogBase: A Scalable Log-structured Database System in the Cloud · Proc. VLDB Endow. 2012
Transaction processing and concurrency control › concurrency control
optimistic concurrency control
0.012010
Towards Elastic Transactional Cloud Storage with Range Query Support · Proc. VLDB Endow. 2010
Distributed systems › replication
data replication
0.012010
Towards Elastic Transactional Cloud Storage with Range Query Support · Proc. VLDB Endow. 2010
Parallel and multicore computing
load balancing
0.012010
Towards Elastic Transactional Cloud Storage with Range Query Support · Proc. VLDB Endow. 2010

Methods — techniques the papers use, named apart from their topics

cloud computing · 0.7peer-to-peer · 0.4peer-to-peer networking · 0.3benchmarking · 0.3operator transformation sharing · 0.3merge sharing · 0.3decomposition sharing · 0.3elastic power-aware platform · 0.2cost modeling · 0.2
YearPublicationVenuePosition
2018 Research Directions in Blockchain Data Management and Analytics
Hoang Tam Vo, Ashish Kundu, Mukesh K. Mohania
EDBT1
2017 Blockchain-based Data Management and Analytics for Micro-insurance Applications
abstract
In this paper, we demonstrate a blockchain-based solution for transparently managing and analyzing data in a pay-as-you-go car insurance application. This application allows drivers who rarely use cars to only pay insurance premium for particular trips they would like to travel. One of the key challenges from database perspective is how to ensure all the data pertaining to the actual trip and premium payment made by the users are transparently recorded so that every party in the insurance contract including the driver, the insurance company, and the financial institution is confident that the data are tamper-proof and traceable.
Hoang Tam Vo, Lenin Mehedy, Mukesh K. Mohania, Ermyas Abebe
CIKM1
2017 Multi-Query Optimization for Complex Event Processing in SAP ESP
abstract
SAP Event Stream Processor (ESP) platform aims at delivering real-time stream processing and analytics in many time-critical areas such as Capital Markets, Internet of Things (IoT) and Data Center Intelligence. SAP ESP allows users to realize complex event processing (CEP) in the form of pattern queries. In this paper, we present MOTTO - a multi-query optimizer in SAP ESP in order to improve the performance of many concurrent pattern queries. This is motivated by the observations that many real-world applications usually have concurrent pattern queries working on the same data streams, leading to tremendous sharing opportunities among queries. In MOTTO, we leverage three major sharing techniques, namely merge, decomposition and operator transformation sharing, to reduce redundant computation among pattern queries. In addition, MOTTO supports nested pattern queries as well as pattern queries with different window sizes. The experiments demonstrate the efficiency of the MOTTO with real-world application scenarios and sensitivity studies.
Shuhao Zhang 0001, Hoang Tam Vo, Daniel Dahlmeier, Bingsheng He
ICDE2
2015 Cost-Model Oblivious Database Tuning with Reinforcement Learning
Debabrota Basu, Qian Lin 0002, Weidong Chen 0004, Hoang Tam Vo, Zihong Yuan, Pierre Senellart, Stéphane Bressan
DEXA (1)4
2014 ScalaGiST: Scalable Generalized Search Trees for MapReduce Systems [Innovative Systems Paper]
abstract
MapReduce has become the state-of-the-art for data parallel processing. Nevertheless, Hadoop, an open-source equivalent of MapReduce, has been noted to have sub-optimal performance in the database context since it is initially designed to operate on raw data without utilizing any type of indexes. To alleviate the problem, we present ScalaGiST - scalable generalized search tree that can be seamlessly integrated with Hadoop, together with a cost-based data access optimizer for efficient query processing at run-time. ScalaGiST provides extensibility in terms of data and query types, hence is able to support unconventional queries (e.g., multi-dimensional range and k -NN queries) in MapReduce systems, and can be dynamically deployed in large cluster environments for handling big users and data. We have built ScalaGiST and demonstrated that it can be easily instantiated to common B + -tree and R-tree indexes yet for dynamic distributed environments. Our extensive performance study shows that ScalaGiST can provide efficient write and read performance, elastic scaling property, as well as effective support for MapReduce execution of ad-hoc analytic queries. Performance comparisions with recent proposals of specialized distributed index structures, such as SpatialHadoop, Data Mapping, and RT-CAN further confirm its efficiency.
Peng Lu 0013, Gang Chen 0001, Beng Chin Ooi, Hoang Tam Vo, Sai Wu
Proc. VLDB Endow.4
2014 BestPeer++: A Peer-to-Peer BasedLarge-Scale Data Processing Platform
abstract
The corporate network is often used for sharing information among the participating companies and facilitating collaboration in a certain industry sector where companies share a common interest. It can effectively help the companies to reduce their operational costs and increase the revenues. However, the inter-company data sharing and processing poses unique challenges to such a data management system including scalability, performance, throughput, and security. In this paper, we present BestPeer++, a system which delivers elastic data sharing services for corporate network applications in the cloud based on BestPeer - a peer-to-peer (P2P) based data management platform. By integrating cloud computing, database, and P2P technologies into one system, BestPeer++ provides an economical, flexible and scalable platform for corporate network applications and delivers data sharing services to participants based on the widely accepted pay-as-you-go business model. We evaluate BestPeer++ on Amazon EC2 Cloud platform. The benchmarking results show that BestPeer++ outperforms HadoopDB, a recently proposed large-scale data processing system, in performance when both systems are employed to handle typical corporate network workloads. The benchmarking results also demonstrate that BestPeer++ achieves near linear scalability for throughput with respect to the number of peer nodes.
Gang Chen 0001, Tianlei Hu, Dawei Jiang, Peng Lu 0013, Kian-Lee Tan, Hoang Tam Vo, Sai Wu
IEEE Trans. Knowl. Data Eng.6
2012 BestPeer++: A Peer-to-Peer Based Large-Scale Data Processing Platform
abstract
The corporate network is often used for sharing information among the participating companies and facilitating collaboration in a certain industry sector where companies share a common interest. It can effectively help the companies to reduce their operational costs and increase the revenues. However, the inter-company data sharing and processing poses unique challenges to such a data management system including scalability, performance, throughput, and security. In this paper, we present Best Peer++, a system which delivers elastic data sharing services for corporate network applications in the cloud based on Best Peer -- a peer-to-peer (P2P) based data management platform. By integrating cloud computing, database, and P2P technologies into one system, Best Peer++ provides an economical, flexible and scalable platform for corporate network applications and delivers data sharing services to participants based on the widely accepted pay-as-you-go business model. We evaluate Best Peer++ on Amazon EC2 Cloud platform. The benchmarking results show that Best Peer++ outperforms Hadoop DB, a recently proposed large-scale data processing system, in performance when both systems are employed to handle typical corporate network workloads. The benchmarking results also demonstrate that Best Peer++ achieves near linear scalability for throughput with respect to the number of peer nodes.
Gang Chen 0001, Tianlei Hu, Dawei Jiang, Peng Lu 0013, Kian-Lee Tan, Hoang Tam Vo, Sai Wu
ICDE6
2012 LogBase: A Scalable Log-structured Database System in the Cloud
abstract
Numerous applications such as financial transactions (e.g., stock trading) are write-heavy in nature. The shift from reads to writes in web applications has also been accelerating in recent years. Write-ahead-logging is a common approach for providing recovery capability while improving performance in most storage systems. However, the separation of log and application data incurs write overheads observed in write-heavy environments and hence adversely affects the write throughput and recovery time in the system. In this paper, we introduce LogBase -- a scalable log-structured database system that adopts log-only storage for removing the write bottleneck and supporting fast system recovery. It is designed to be dynamically deployed on commodity clusters to take advantage of elastic scaling property of cloud environments. LogBase provides in-memory multiversion indexes for supporting efficient access to data maintained in the log. LogBase also supports transactions that bundle read and write operations spanning across multiple records. We implemented the proposed system and compared it with HBase and a disk-based log-structured record-oriented system modeled after RAMCloud. The experimental results show that LogBase is able to provide sustained write throughput, efficient data access out of the cache, and effective system recovery.
Hoang Tam Vo, Sheng Wang 0011, Divyakant Agrawal, Gang Chen 0001, Beng Chin Ooi
Proc. VLDB Endow.1
2011 ES2: A cloud data storage system for supporting both OLTP and OLAP
abstract
Cloud computing represents a paradigm shift driven by the increasing demand of Web based applications for elastic, scalable and efficient system architectures that can efficiently support their ever-growing data volume and large-scale data analysis. A typical data management system has to deal with real-time updates by individual users, and as well as periodical large scale analytical processing, indexing, and data extraction. While such operations may take place in the same domain, the design and development of the systems have somehow evolved independently for transactional and periodical analytical processing. Such a system-level separation has resulted in problems such as data freshness as well as serious data storage redundancy. Ideally, it would be more efficient to apply ad-hoc analytical processing on the same data directly. However, to the best of our knowledge, such an approach has not been adopted in real implementation. Intrigued by such an observation, we have designed and implemented epiC, an elastic power-aware data-itensive Cloud platform for supporting both data intensive analytical operations (ref. as OLAP) and online transactions (ref. as OLTP). In this paper, we present ES2- the elastic data storage system of epiC, which is designed to support both functionalities within the same storage. We present the system architecture and the functions of each system component, and experimental results which demonstrate the efficiency of the system.
Chun Chen 0001, Dawei Jiang, Beng Chin Ooi, Hoang Tam Vo, Sai Wu, Quanqing Xu
ICDE7
2011 A Framework for Supporting DBMS-like Indexes in the Cloud
Gang Chen 0001, Hoang Tam Vo, Sai Wu, Beng Chin Ooi, M. Tamer Özsu
Proc. VLDB Endow.2
2010 Providing Scalable Database Services on the Cloud
Chun Chen 0001, Gang Chen 0001, Dawei Jiang, Beng Chin Ooi, Hoang Tam Vo, Sai Wu, Quanqing Xu
WISE5
2010 Towards Elastic Transactional Cloud Storage with Range Query Support
abstract
Cloud storage is an emerging infrastructure that offers Platforms as a Service (PaaS). On such platforms, storage and compute power are adjusted dynamically, and therefore it is important to build a highly scalable and reliable storage that can elastically scale on-demand with minimal startup cost. In this paper, we propose ecStore -- an elastic cloud storage system that supports automated data partitioning and replication, load balancing, efficient range query, and transactional access. In ecStore, data objects are distributed and replicated in a cluster of commodity computer nodes located in the cloud. Users can access data via transactions which bundle read and write operations on multiple data items stored on possibly different cluster nodes. The architecture of ecStore follows a stratum design that leverages an underlying distributed index with a replication layer in the middle and a transaction management layer on top. ecStore provides adaptive read consistency on replicated data. We also enhance the system with an effective load balancing scheme using a self-tuning replication technique that is specially designed for large-scale data. Furthermore, a multi-version optimistic concurrency control scheme matches well with the characteristics of data in cloud storages. To validate the performance of the system, we have conducted extensive experiments on various platforms including a commercial cloud (Amazon's EC2), an in-house cluster, and PlanetLab.
Hoang Tam Vo, Chun Chen 0001, Beng Chin Ooi
Proc. VLDB Endow.1