Botong Huang

dblp:131/4058 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
2since 2021 · last 2025
0000-0001-7870-4997ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 5 first-author · 2 since 2021Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Query processing and optimization · 84% Machine learning and data management · 9% Distributed and cloud data management · 5%
Computer architecture, parallel and distributed computing, and storage systems
7 papers
Cloud and datacenter computing · 90% Parallel and multicore computing · 10%

Topics — the 20 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › incremental computation
incremental query processing
1.732025
Agamotto: Scheduling of Deadline-Oriented Incremental Query Execution under Uncertain Resource Price · Proc. VLDB Endow. 2025
Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing · Proc. VLDB Endow. 2020
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Query processing and optimization
incremental computation
1.122023
Tempura: a general cost-based optimizer framework for incremental data processing (Journal Version) · VLDB J. 2023
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.932020
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Hydra: a federated resource manager for data-center scale analytics · NSDI 2019
Resource Elasticity for Large-Scale Machine Learning · SIGMOD Conference 2015
Cloud and datacenter computing › resource management
resource management and scheduling
0.912025
Agamotto: Scheduling of Deadline-Oriented Incremental Query Execution under Uncertain Resource Price · Proc. VLDB Endow. 2025
Query processing and optimization › query optimization
cost-based optimization
0.822023
Tempura: a general cost-based optimizer framework for incremental data processing (Journal Version) · VLDB J. 2023
Cumulon: optimizing statistical data analysis in the cloud · SIGMOD Conference 2013
Query processing and optimization
query optimization
0.412020
Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing · Proc. VLDB Endow. 2020
Parallel and multicore computing › parallel scheduling
resource-aware scheduling
0.412020
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Cloud and datacenter computing › cluster resource management and scheduling
resource scheduling
0.412019
Hydra: a federated resource manager for data-center scale analytics · NSDI 2019
Cloud and datacenter computing › resource management
cloud resource management
0.312017
Cümülön-D: Data Analytics in a Dynamic Spot Market · Proc. VLDB Endow. 2017
Cloud and datacenter computing › utility computing › cloud pricing
spot market
0.312017
Cümülön-D: Data Analytics in a Dynamic Spot Market · Proc. VLDB Endow. 2017
Cloud and datacenter computing
workflow scheduling
0.312017
Cümülön-D: Data Analytics in a Dynamic Spot Market · Proc. VLDB Endow. 2017
Machine learning and data management › scalable machine learning
distributed learning
0.212015
Resource Elasticity for Large-Scale Machine Learning · SIGMOD Conference 2015
Machine learning and data management
scalable machine learning
0.212015
Resource Elasticity for Large-Scale Machine Learning · SIGMOD Conference 2015
Cloud and datacenter computing
cluster resource management and scheduling
0.212015
Cumulon: Matrix-Based Data Analytics in the Cloud with Spot Instances · Proc. VLDB Endow. 2015
Cloud and datacenter computing › job scheduling › economic scheduling
cost-aware scheduling
0.212015
Cumulon: Matrix-Based Data Analytics in the Cloud with Spot Instances · Proc. VLDB Endow. 2015
Cloud and datacenter computing › utility computing › cloud pricing
spot instance
0.212015
Cumulon: Matrix-Based Data Analytics in the Cloud with Spot Instances · Proc. VLDB Endow. 2015
Data integration and cleaning
data warehouse
0.112020
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Query processing and optimization › view maintenance
incremental view maintenance
0.112020
Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing · Proc. VLDB Endow. 2020
Mathematical optimization › sequential decision making
markov decision processes
0.112017
Cümülön-D: Data Analytics in a Dynamic Spot Market · Proc. VLDB Endow. 2017
Cloud and datacenter computing › resource management
resource negotiation
0.112015
Resource Elasticity for Large-Scale Machine Learning · SIGMOD Conference 2015

Methods — techniques the papers use, named apart from their topics

markov decision process · 2.3prophet scheduler · 1.7incremental batch processing · 0.9cost-based optimizer framework · 0.7risk quantification · 0.6resource optimization · 0.4cost-based optimization · 0.4time-varying relations · 0.4rewrite rules · 0.4plan space exploration · 0.4benchmarking · 0.3simulation · 0.2search · 0.2modeling · 0.2
YearPublicationVenuePosition
2025 Agamotto: Scheduling of Deadline-Oriented Incremental Query Execution under Uncertain Resource Price
abstract
Incremental query processing is widely used in data warehouses and streaming systems. While many optimization techniques are developed to generate incremental query plans, the scheduling support for incremental processing remains preliminary. Typically, execution is triggered with fixed frequencies specified by the user. In this paper, we propose a novel scheduling problem for incremental query execution under a deadline, assuming the resource has a fluctuating and unforeseen price. We propose two naive solutions as well as a prophet scheduler that foresees the future. We present an end-to-end system Agamotto that models future probabilities offline with a Markov Decision Process (MDP) and makes cost-based and dynamic scheduling decisions online. We show how Agamotto can be extended to handle a workflow of dependent queries, so that they can all incrementally execute in an asynchronous fashion. Experiments show that Agamotto consistently outperforms the naive solutions, and the achieved cost is on average 10x closer to the theoretical lower bound provided by the prophet scheduler.
Botong Huang, Lianggui Weng, Wei Chen 0133, Zuozhi Wang, Kai Zeng 0002, Chen Li 0001, Yihui Feng, Bolin Ding, Jingren Zhou 0001
Proc. VLDB Endow.1
2023 Tempura: a general cost-based optimizer framework for incremental data processing (Journal Version)
Zuozhi Wang, Kai Zeng 0002, Botong Huang, Wei Chen 0133, Xiaozong Cui, Liya Fan, Dachuan Qu, Chen Li 0001, Jingren Zhou 0001
VLDB J.3
2020 Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing
abstract
As the primary approach to deriving decision-support insights, automated recurring routine analytic jobs account for a major part of cluster resource usages in modern enterprise data warehouses. These recurring routine jobs usually have stringent schedule and deadline determined by external business logic, and thus cause dreadful resource skew and severe resource over-provision in the cluster. In this paper, we present Grosbeak, a novel data warehouse that supports resource-aware incremental computing to process recurring routine jobs, smooths the resource skew, and optimizes the resource usage. Unlike batch processing in traditional data warehouses, Grosbeak leverages the fact that data is continuously ingested. It breaks an analysis job into small batches that incrementally process the progressively available data, and schedules these small-batch jobs intelligently when the cluster has free resources. In this demonstration, we showcase Grosbeak using real-world analysis pipelines. Users can interact with the data warehouse by registering recurring queries and observing the incremental scheduling behavior and smoothed resource usage pattern.
Zuozhi Wang, Kai Zeng 0002, Botong Huang, Wei Chen 0133, Xiaozong Cui, Liya Fan, Dachuan Qu, Chen Li 0001, Jingren Zhou 0001
SIGMOD Conference3
2020 Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing
abstract
Incremental processing is widely-adopted in many applications, ranging from incremental view maintenance, stream computing, to recently emerging progressive data warehouse and intermittent query processing. Despite many algorithms developed on this topic, none of them can produce an incremental plan that always achieves the best performance, since the optimal plan is data dependent. In this paper, we develop a novel cost-based optimizer framework, called Tempura, for optimizing incremental data processing. We propose an incremental query planning model called TIP based on the concept of time-varying relations, which can formally model incremental processing in its most general form. We give a full specification of Tempura, which can not only unify various existing techniques to generate an optimal incremental plan, but also allow the developer to add their rewrite rules. We study how to explore the plan space and search for an optimal incremental plan. We evaluate Tempura in various incremental processing scenarios to show its effectiveness and efficiency.
Zuozhi Wang, Kai Zeng 0002, Botong Huang, Wei Chen 0133, Xiaozong Cui, Liya Fan, Dachuan Qu, Chen Li 0001, Jingren Zhou 0001
Proc. VLDB Endow.3
2019 Hydra: a federated resource manager for data-center scale analytics
Carlo Curino, Subru Krishnan, Konstantinos Karanasos, Sriram Rao, Giovanni Matteo Fumarola, Botong Huang, Kishore Chaliparambil, Arun Suresh, Young Chen, Solom Heddaya, Roni Burd, Sarvesh Sakalanaga, Chris Douglas, Bill Ramsey, Raghu Ramakrishnan 0001
NSDI6
2017 Cümülön-D: Data Analytics in a Dynamic Spot Market
abstract
We present a system called Cümülön-D for matrix-based data analysis in a spot market of a public cloud. Prices in such markets fluctuate over time: while users can acquire machines usually at a very low bid price, the cloud can terminate these machines as soon as the market price exceeds their bid price. The distinguishing features of Cümülön-D include its continuous, proactive adaptation to the changing market, and its ability to quantify and control the monetary risk involved in paying for a workflow execution. We solve the dynamic optimization problem in a principled manner with a Markov decision process, and account for practical details that are often ignored previously but nonetheless important to performance. We evaluate Cümülön-D's effectiveness and advantages over previous approaches with experiments on Amazon EC2.
Botong Huang, Jun Yang 0001
Proc. VLDB Endow.1
2015 Resource Elasticity for Large-Scale Machine Learning
abstract
Declarative large-scale machine learning (ML) aims at flexible specification of ML algorithms and automatic generation of hybrid runtime plans ranging from single node, in-memory computations to distributed computations on MapReduce (MR) or similar frameworks. State-of-the-art compilers in this context are very sensitive to memory constraints of the master process and MR cluster configuration. Different memory configurations can lead to significant performance differences. Interestingly, resource negotiation frameworks like YARN allow us to explicitly request preferred resources including memory. This capability enables automatic resource elasticity, which is not just important for performance but also removes the need for a static cluster configuration, which is always a compromise in multi-tenancy environments. In this paper, we introduce a simple and robust approach to automatic resource elasticity for large-scale ML. This includes (1) a resource optimizer to find near-optimal memory configurations for a given ML program, and (2) dynamic plan migration to adapt memory configurations during runtime. These techniques adapt resources according to data, program, and cluster characteristics. Our experiments demonstrate significant improvements up to 21x without unnecessary over-provisioning and low optimization overhead.
Botong Huang, Matthias Boehm 0001, Yuanyuan Tian 0001, Berthold Reinwald, Shirish Tatikonda, Frederick Reiss 0001
SIGMOD Conference1
2015 Cumulon: Matrix-Based Data Analytics in the Cloud with Spot Instances
abstract
We describe Cümülön, a system aimed at helping users develop and deploy matrix-based data analysis programs in a public cloud. A key feature of Cümülön is its end-to-end support for the so-called spot instances ---machines whose market price fluctuates over time but is usually much lower than the regular fixed price. A user sets a bid price when acquiring spot instances, and loses them as soon as the market price exceeds the bid price. While spot instances can potentially save cost, they are difficult to use effectively, and run the risk of not finishing work while costing more. Cümülön provides a highly elastic computation and storage engine on top of spot instances, and offers automatic cost-based optimization of execution, deployment, and bidding strategies. Cümülön further quantifies how the uncertainty in the market price translates into the cost uncertainty of its recommendations, and allows users to specify their risk tolerance as an optimization constraint.
Botong Huang, Nicholas W. D. Jarrett, Shivnath Babu, Sayan Mukherjee 0001, Jun Yang 0001
Proc. VLDB Endow.1
2013 Cumulon: optimizing statistical data analysis in the cloud
abstract
We present Cumulon, a system designed to help users rapidly develop and intelligently deploy matrix-based big-data analysis programs in the cloud. Cumulon features a flexible execution model and new operators especially suited for such workloads. We show how to implement Cumulon on top of Hadoop/HDFS while avoiding limitations of MapReduce, and demonstrate Cumulon's performance advantages over existing Hadoop-based systems for statistical data analysis. To support intelligent deployment in the cloud according to time/budget constraints, Cumulon goes beyond database-style optimization to make choices automatically on not only physical operators and their parameters, but also hardware provisioning and configuration settings. We apply a suite of benchmarking, simulation, modeling, and search techniques to support effective cost-based optimization over this rich space of deployment plans.
Botong Huang, Shivnath Babu, Jun Yang 0001
SIGMOD Conference1