Dai Hai Ton That

dblp:168/0902 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
4since 2021 · last 2022
0000-0002-3935-2471ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2022 Selecting Approaches for Enabling Enterprise Data Search: NASA's Science Mission Directorate (SMD) Catalog
abstract
NASA's Science Mission Directorate (SMD) is working to build an open-source science infrastructure to accelerate open, collaborative and interdisciplinary science. One key component in the open-source science infrastructure is the SMD data catalog. In this paper, we present our process for selecting a technical approach to building a NASA SMD enterprise-wide integrated search capability for science users across multiple science disciplines to support discovery and access to complex scientific data.
Kaylin M. Bugbee, Rahul Ramachandran, Ashish Acharya, Dai Hai Ton That, John Hedman, Ahmed Eleish, Charles Driessnack, Wesley Adams, Emily Foshee
IGARSS4
2021 LDI: Learned Distribution Index for Column Stores
abstract
In column stores, which ingest large amounts of data into multiple column groups, query performance deteriorates. Commercial column stores use log-structured merge (LSM) tree on projections to ingest data rapidly. LSM improves ingestion performance, but in column stores the sort-merge phase is I/O-intensive, which slows concurrent queries and reduces overall throughput. In this paper, we aim to reduce the sorting and merging cost that arise when data is ingested in column stores. We present LDI, a learned distribution index for column stores. LDI learns a frequency-based data distribution and constructs a bucket worth of data based on the learned distribution. Filled buckets that conform to the distribution are written out to disk; unfilled buckets are retained to achieve the desired level of sortedness, thus avoiding the expensive sort-merge phase. We present an algorithm to learn and adapt to distributions, and a robust implementation that takes advantage of disk parallelism. We compare LDI with LSM and production columnar stores using real and synthetic datasets.
Dai Hai Ton That, Mohammadsaleh Gharehdaghi, Alexander Rasin, Tanu Malik
IEEE BigData1
2021 On Lowering Merge Costs of an LSM Tree
abstract
In column stores, which ingest large amounts of data into multiple column groups, query performance deteriorates. Commercial column stores use log-structured merge (LSM) tree on projections to ingest data rapidly. LSM tree improves ingestion performance, but for column stores the sort-merge maintenance phase in an LSM tree is I/O-intensive, which slows concurrent queries and reduces overall throughput. In this paper, we present a simple heuristic approach to reduce the sorting and merging cost that arise when data is ingested in column stores. We demonstrate how a Min-Max heuristic can construct buckets and identify the level of sortedness in each range of data. Filled and relatively-sorted buckets are written out to disk; unfilled buckets are retained to achieve a better level of sortedness, thus avoiding the expensive sort-merge phase. We compare our Min-Max approach with LSM tree and production columnar stores using real and synthetic datasets.
Dai Hai Ton That, Mohammadsaleh Gharehdaghi, Alexander Rasin, Tanu Malik
SSDBM1
2021 Mobile participatory sensing with strong privacy guarantees using secure probes
Iulian Sandu Popa, Dai Hai Ton That, Karine Zeitouni, Cristian Borcea
GeoInformatica2
2020 ODSA: Open Database Storage Access
James Wagner, Alexander Rasin, Dai Hai Ton That, Tanu Malik, Jonathan Grier
EDBT3
2020 Advancing Open Science Through Innovative Data System Solutions: The Joint ESA-NASA Multi-Mission Algorithm and Analysis Platform (MAAP)'s Data Ecosystem
abstract
Collaborative open science practices are changing the way research is conducted. These changes affect how scientists work together on data, code and information. Data systems enhance open science by offering forward thinking technological solutions, such as providing data and computation on the cloud, to enable collaboration, sharing and analysis. In this paper, we present our vision for a conceptual data system on the cloud that enables open science. We also present our work on the Multi-Mission Algorithm and Analysis Platform (MAAP) which has served as a pathfinder data system for this conceptual approach.
Kaylin M. Bugbee, Rahul Ramachandran, Manil Maskey, Aimee Barciauskas, Aaron Kaulfus, Dai Hai Ton That, Katrina Virts, Kel N. Markert, Christopher Lynnes
IGARSS6
2019 SciInc: A Container Runtime for Incremental Recomputation
abstract
The conduct of reproducible science improves when computations are portable and verifiable. A container runtime provides an isolated environment for running computations and thus is useful for porting applications on new machines. Current container engines, such as LXC and Docker, however, do not track provenance, which is essential for verifying computations. In this paper, we present SciInc, a container runtime that tracks the provenance of computations during container creation. We show how container engines can use audited provenance data for efficient container replay. SciInc observes inputs to computations, and, if they change, propagates the changes, re-using partially memoized computations and data that are identical across replay and original run. We chose light-weight data structures for storing the provenance trace to maintain the invariant of shareable and portable container runtime. To determine the effectiveness of change propagation and memoization, we compared popular container technology and incremental recomputation methods using published data analysis experiments.
Andrew Youngdahl, Dai Hai Ton That, Tanu Malik
eScience2
2019 PLI $$^+$$ + : efficient clustering of cloud databases
Dai Hai Ton That, James Wagner, Alexander Rasin, Tanu Malik
Distributed Parallel Databases1
2017 Sciunits: Reusable Research Objects
abstract
Science is conducted collaboratively, often requiring knowledge sharing about computational experiments. When experiments include only datasets, they can be shared using Uniform Resource Identifiers (URIs) or Digital Object Identifiers (DOIs). An experiment, however, seldom includes only datasets, but more often includes software, its past execution, provenance, and associated documentation. The Research Object has recently emerged as a comprehensive and systematic method for aggregation and identification of diverse elements of computational experiments. While a necessary method, mere aggregation is not sufficient for the sharing of computational experiments. Other users must be able to easily recompute on these shared research objects. In this paper, we present the sciunit, a reusable research object in which aggregated content is recomputable. We describe a Git-like client that efficiently creates, stores, and repeats sciunits. We show through analysis that sciunits repeat computational experiments with minimal storage and processing overhead. Finally, we provide an overview of sharing and reproducible cyberinfrastructure based on sciunits gaining adoption in the domain of geosciences.
Dai Hai Ton That, Gabriel Fils, Zhihao Yuan, Tanu Malik
eScience1
2017 PLI: Augmenting Live Databases with Custom Clustered Indexes
abstract
RDBMSes only support one clustered index per database table that can speed up query processing. Database applications, that continually ingest large amounts of data, perceive slow query response times to long downtimes, as the clustered index ordering must be strictly maintained. In this paper, we show that application slowdown or downtime, however, can often be avoided if database systems expose the physical location of attributes that are completely or approximately clustered.
James Wagner, Alexander Rasin, Dai Hai Ton That, Tanu Malik
SSDBM3
2016 PAMPAS: Privacy-Aware Mobile Participatory Sensing Using Secure Probes
abstract
Mobile participatory sensing could be used in many applications such as vehicular traffic monitoring, pollution tracking, or even health surveying. However, its success depends on finding a solution for querying large numbers of users which protects user location privacy and works in real-time. This paper presents PAMPAS, a privacy-aware mobile distributed system for efficient data aggregation in mobile participatory sensing. In PAMPAS, mobile devices enhanced with secure hardware, called secure probes (SPs), perform distributed query processing, while preventing users from accessing other users' data. A supporting server infrastructure (SSI) coordinates the inter-SP communication and the computation tasks executed on SPs. PAMPAS ensures that SSI cannot link the location reported by SPs to the user identities even if SSI has additional background information. In addition to its novel system architecture, PAMPAS also proposes two new protocols for privacy-aware location-based aggregation and adaptive spatial partitioning of SPs that work efficiently on resource-constrained SPs. Our experimental results and security analysis demonstrate that these protocols are able to collect the data, aggregate them, and share statistics or derived models in real-time, without any location privacy leakage.
Dai Hai Ton That, Iulian Sandu Popa, Karine Zeitouni, Cristian Borcea
SSDBM1
2015 PPTM: Privacy-Aware Participatory Traffic Monitoring Using Mobile Secure Probes
abstract
Privacy became one of the main concerns in location-based services in general and in community-based traffic monitoring in particular. This demonstration presents a new approach for privacy preserving online traffic monitoring using mobile probes. It combines hardware and software solutions, and a secure protocol to collect, aggregate and share the traffic information.
Dai Hai Ton That, Iulian Sandu Popa, Karine Zeitouni
MDM (1)1