Harold Lim

dblp:18/7542 · also Harold C. Lim · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-authorSystems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Energy-efficient computing · 48% Parallel and multicore computing · 18% High-performance computing · 18%
Databases, data mining, and information retrieval
2 papers
Data stream processing · 53% Query processing and optimization · 35% Data mining · 12%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data stream processing
continuous query processing
0.212013
Execution and optimization of continuous queries with cyclops · SIGMOD Conference 2013
Parallel and multicore computing › data-parallel programming
mapreduce
0.112012
Stubby: A Transformation-based Optimizer for MapReduce Workflows · Proc. VLDB Endow. 2012
High-performance computing › scientific workflow
workflow optimization
0.112012
Stubby: A Transformation-based Optimizer for MapReduce Workflows · Proc. VLDB Endow. 2012
Energy-efficient computing
datacenter power management
0.112011
Power Budgeting for Virtualized Data Centers · USENIX ATC 2011
Cloud and datacenter computing › resource management
datacenter resource management
0.112011
Power Budgeting for Virtualized Data Centers · USENIX ATC 2011
Energy-efficient computing › power management
power budgeting
0.112011
Power Budgeting for Virtualized Data Centers · USENIX ATC 2011
Energy-efficient computing
power management
0.112011
Power Budgeting for Virtualized Data Centers · USENIX ATC 2011
Data stream processing
batching
0.012013
Execution and optimization of continuous queries with cyclops · SIGMOD Conference 2013
Data mining
large-scale data analytics
0.012013
Execution and optimization of continuous queries with cyclops · SIGMOD Conference 2013

Methods — techniques the papers use, named apart from their topics

plan-to-plan transformation · 0.3cost-based search · 0.3
YearPublicationVenuePosition
2013 How to Fit when No One Size Fits
Harold Lim, Yuzhang Han, Shivnath Babu
CIDR1
2013 Execution and optimization of continuous queries with cyclops
abstract
As the data collected by enterprises grows in scale, there is a growing trend of performing data analytics on large datasets. Batch processing systems that can handle petabyte scale of data, such as Hadoop, have flourished and gained traction in the industry. As the results of batch analytics have been used to continuously improve front-facing user experience, there is a growing interest in pushing the processing latency down. This trend has fueled a resurgence in the development and usage of execution engines that can process continuous queries.
Harold Lim, Shivnath Babu
SIGMOD Conference1
2012 Stubby: A Transformation-based Optimizer for MapReduce Workflows
abstract
There is a growing trend of performing analysis on large datasets using workflows composed of MapReduce jobs connected through producer-consumer relationships based on data. This trend has spurred the development of a number of interfaces---ranging from program-based to query-based interfaces---for generating MapReduce workflows. Studies have shown that the gap in performance can be quite large between optimized and unoptimized workflows. However, automatic cost-based optimization of MapReduce workflows remains a challenge due to the multitude of interfaces, large size of the execution plan space, and the frequent unavailability of all types of information needed for optimization. We introduce a comprehensive plan space for MapReduce workflows generated by popular workflow generators. We then propose Stubby , a cost-based optimizer that searches selectively through the subspace of the full plan space that can be enumerated correctly and costed based on the information available in any given setting. Stubby enumerates the plan space based on plan-to-plan transformations and an efficient search algorithm. Stubby is designed to be extensible to new interfaces and new types of optimizations, which is a desirable feature given how rapidly MapReduce systems are evolving. Stubby's efficiency and effectiveness have been evaluated using representative workflows from many domains.
Harold Lim, Herodotos Herodotou, Shivnath Babu
Proc. VLDB Endow.1
2011 Starfish: A Self-tuning System for Big Data Analytics
Herodotos Herodotou, Harold Lim, Nedyalko Borisov, Fatma Bilgen Cetin, Shivnath Babu
CIDR2
2011 Power Budgeting for Virtualized Data Centers
Harold Lim, Aman Kansal, Jie Liu 0001
USENIX ATC1