Jörg Schad

dblp:07/1890 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
0since 2021 · last 2013
0000-0002-1552-382XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Distributed systems · 42% Cloud and datacenter computing · 26% Performance modeling and evaluation · 12%
Databases, data mining, and information retrieval
2 papers
Indexing and storage engines · 70% Query processing and optimization · 30%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › cluster computing framework
mapreduce optimization
0.322012
Only Aggressive Elephants are Fast Elephants · Proc. VLDB Endow. 2012
Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) · Proc. VLDB Endow. 2010
Distributed systems › fault tolerance
failure recovery
0.222011
RAFT at work: speeding-up mapreduce applications under task and node failures · SIGMOD Conference 2011
RAFTing MapReduce: Fast recovery on the RAFT · ICDE 2011
Distributed systems
fault tolerance
0.222011
RAFT at work: speeding-up mapreduce applications under task and node failures · SIGMOD Conference 2011
RAFTing MapReduce: Fast recovery on the RAFT · ICDE 2011
Parallel and multicore computing › data-parallel programming
mapreduce
0.222011
RAFTing MapReduce: Fast recovery on the RAFT · ICDE 2011
RAFT at work: speeding-up mapreduce applications under task and node failures · SIGMOD Conference 2011
Storage systems › file systems
distributed file system
0.112012
Only Aggressive Elephants are Fast Elephants · Proc. VLDB Endow. 2012
Distributed systems › fault tolerance
checkpointing
0.112011
RAFT at work: speeding-up mapreduce applications under task and node failures · SIGMOD Conference 2011
Query processing and optimization › join processing › distributed join
mapreduce join
0.112010
Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) · Proc. VLDB Endow. 2010
Performance modeling and evaluation
benchmarking
0.112010
Runtime Measurements in the Cloud: Observing, Analyzing, and Reducing Variance · Proc. VLDB Endow. 2010
Cloud and datacenter computing › cluster computing framework
hadoop
0.122012
Only Aggressive Elephants are Fast Elephants · Proc. VLDB Endow. 2012
Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) · Proc. VLDB Endow. 2010
Cloud and datacenter computing › cluster resource management and scheduling
cluster scheduling
0.012011
RAFTing MapReduce: Fast recovery on the RAFT · ICDE 2011
Cloud and datacenter computing
cloud infrastructure
0.012010
Runtime Measurements in the Cloud: Observing, Analyzing, and Reducing Variance · Proc. VLDB Endow. 2010
Performance modeling and evaluation › workload characterization › parallel workload analysis
mapreduce workload characterization
0.012010
Runtime Measurements in the Cloud: Observing, Analyzing, and Reducing Variance · Proc. VLDB Endow. 2010
Performance modeling and evaluation
workload characterization
0.012010
Runtime Measurements in the Cloud: Observing, Analyzing, and Reducing Variance · Proc. VLDB Endow. 2010

Methods — techniques the papers use, named apart from their topics

upload pipeline modification · 0.3clustered indexes on block replicas · 0.3clustered indexing · 0.2UDF injection · 0.2query metadata checkpointing · 0.1checkpointing · 0.1microbenchmarking · 0.1
YearPublicationVenuePosition
2013 Elephant, Do Not Forget Everything! Efficient Processing of Growing Datasets
abstract
MapReduce has become quite popular to analyse very large datasets. Nevertheless, users typically have to run their MapReduce jobs over the whole dataset every time the dataset is appended by new records. Some researchers have proposed to reuse the intermediate data produced by previous MapReduce jobs. However, existing works still have to read the whole dataset in order to identify which parts of the dataset changed. Furthermore, storing intermediate results is not suitable in some cases, because it can lead to a very high storage overhead. In this paper, we propose Itchy, a MapReduce-based system that employes a set of different techniques to efficiently deal with growing datasets. Itchy uses an optimizer to automatically choose the right technique to process a MapReduce job. The beauty of Itchy is that it does not have to read the whole dataset again to deal with new records. In more detail, Itchy keeps track of the provenance of intermediate results in order to selectively recompute intermediate results as required. But, if intermediate results are small or the computational cost of map functions is high, Itchy can automatically start storing intermediate results rather than the provenance information. Additionally, Itchy also supports the option of directly merging outputs from several jobs in cases where MapReduce jobs allow for such kind of processing. We evaluate Itchy using two different benchmarks and compare it with Hadoop and Incoop. The results show the superiority of Itchy over both baseline systems for processing incremental jobs. In terms of job runtime, Itchy is more than one order of magnitude faster than Hadoop (up to ~41 times faster) and Incoop (up to ~11 times faster).
Jörg Schad, Jorge-Arnulfo Quiané-Ruiz, Jens Dittrich
IEEE CLOUD1
2012 Only Aggressive Elephants are Fast Elephants
abstract
Yellow elephants are slow. A major reason is that they consume their inputs entirely before responding to an elephant rider's orders. Some clever riders have trained their yellow elephants to only consume parts of the inputs before responding. However, the teaching time to make an elephant do that is high. So high that the teaching lessons often do not pay off. We take a different approach. We make elephants aggressive; only this will make them very fast. We propose HAIL (Hadoop Aggressive Indexing Library), an enhancement of HDFS and Hadoop MapReduce that dramatically improves runtimes of several classes of MapReduce jobs. HAIL changes the upload pipeline of HDFS in order to create different clustered indexes on each data block replica. An interesting feature of HAIL is that we typically create a win-win situation: we improve both data upload to HDFS and the runtime of the actual Hadoop MapReduce job. In terms of data upload, HAIL improves over HDFS by up to 60% with the default replication factor of three. In terms of query execution, we demonstrate that HAIL runs up to 68x faster than Hadoop. In our experiments, we use six clusters including physical and EC2 clusters of up to 100 nodes. A series of scalability experiments also demonstrates the superiority of HAIL.
Jens Dittrich, Jorge-Arnulfo Quiané-Ruiz, Stefan Richter 0007, Stefan Schuh, Alekh Jindal, Jörg Schad
Proc. VLDB Endow.6
2011 RAFTing MapReduce: Fast recovery on the RAFT
abstract
MapReduce is a computing paradigm that has gained a lot of popularity as it allows non-expert users to easily run complex analytical tasks at very large-scale. At such scale, task and node failures are no longer an exception but rather a characteristic of large-scale systems. This makes fault-tolerance a critical issue for the efficient operation of any application. MapReduce automatically reschedules failed tasks to available nodes, which in turn recompute such tasks from scratch. However, this policy can significantly decrease performance of applications. In this paper, we propose a family of Recovery Algorithms for Fast-Tracking (RAFT) MapReduce. As ease-of-use is a major feature of MapReduce, RAFT focuses on simplicity and also non-intrusiveness, in order to be implementation-independent. To efficiently recover from task failures, RAFT exploits the fact that MapReduce produces and persists intermediate results at several points in time. RAFT piggy-backs checkpoints on the task progress computation. To deal with multiple node failures, we propose query metadata checkpointing. We keep track of the mapping between input key-value pairs and intermediate data for all reduce tasks. Thereby, RAFT does not need to re-execute completed map tasks entirely. Instead RAFT only recomputes intermediate data that were processed for local reduce tasks and hence not shipped to another node for processing. We also introduce a scheduling strategy taking full advantage of these recovery algorithms. We implemented RAFT on top of Hadoop and evaluated it on a 45-node cluster using three common analytical tasks. Overall, our experimental results demonstrate that RAFT outperforms Hadoop runtimes by 23% on average under task and node failures. The results also show that RAFT has negligible runtime overhead.
Jorge-Arnulfo Quiané-Ruiz, Christoph Pinkel, Jörg Schad, Jens Dittrich
ICDE3
2011 RAFT at work: speeding-up mapreduce applications under task and node failures
abstract
The MapReduce framework is typically deployed on very large computing clusters where task and node failures are no longer an exception but the rule. Thus, fault-tolerance is an important aspect for the efficient operation of MapReduce jobs. However, currently MapReduce implementations fully recompute failed tasks (subparts of a job) from the beginning. This can significantly decrease the runtime performance of MapReduce applications. We present an alternative system that implements RAFT ideas. RAFT is a family of powerful and inexpensive Recovery Algorithms for Fast-Tracking MapReduce jobs under task and node failures. To recover from task failures, RAFT exploits the intermediate results persisted by MapReduce at several points in time. RAFT piggybacks checkpoints on the task progress computation. To recover from node failures, RAFT maintains a per-map task list of all input key-value pairs producing intermediate results and pushes intermediate results to reducers. In this demo, we demonstrate that RAFT recovers efficiently from both task and node failures. Further, the audience can compare RAFT with Hadoop via an easy-to-use web interface.
Jorge-Arnulfo Quiané-Ruiz, Christoph Pinkel, Jörg Schad, Jens Dittrich
SIGMOD Conference3
2010 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing)
abstract
MapReduce is a computing paradigm that has gained a lot of attention in recent years from industry and research. Unlike parallel DBMSs, MapReduce allows non-expert users to run complex analytical tasks over very large data sets on very large clusters and clouds. However, this comes at a price: MapReduce processes tasks in a scan-oriented fashion. Hence, the performance of Hadoop --- an open-source implementation of MapReduce --- often does not match the one of a well-configured parallel DBMS. In this paper we propose a new type of system named Hadoop++: it boosts task performance without changing the Hadoop framework at all (Hadoop does not even 'notice it'). To reach this goal, rather than changing a working system (Hadoop), we inject our technology at the right places through UDFs only and affect Hadoop from inside . This has three important consequences: First, Hadoop++ significantly outperforms Hadoop. Second, any future changes of Hadoop may directly be used with Hadoop++ without rewriting any glue code. Third, Hadoop++ does not need to change the Hadoop interface. Our experiments show the superiority of Hadoop++ over both Hadoop and HadoopDB for tasks related to indexing and join processing.
Jens Dittrich, Jorge-Arnulfo Quiané-Ruiz, Alekh Jindal, Yagiz Kargin, Vinay Setty, Jörg Schad
Proc. VLDB Endow.6
2010 Runtime Measurements in the Cloud: Observing, Analyzing, and Reducing Variance
abstract
One of the main reasons why cloud computing has gained so much popularity is due to its ease of use and its ability to scale computing resources on demand. As a result, users can now rent computing nodes on large commercial clusters through several vendors, such as Amazon and rackspace. However, despite the attention paid by Cloud providers, performance unpredictability is a major issue in Cloud computing for (1) database researchers performing wall clock experiments, and (2) database applications providing service-level agreements. In this paper, we carry out a study of the performance variance of the most widely used Cloud infrastructure (Amazon EC2) from different perspectives. We use established microbenchmarks to measure performance variance in CPU, I/O, and network. And, we use a multi-node MapReduce application to quantify the impact on real dataintensive applications. We collected data for an entire month and compare it with the results obtained on a local cluster. Our results show that EC2 performance varies a lot and often falls into two bands having a large performance gap in-between --- which is somewhat surprising. We observe in our experiments that these two bands correspond to the different virtual system types provided by Amazon. Moreover, we analyze results considering different availability zones, points in time, and locations. This analysis indicates that, among others, the choice of availability zone also influences the performance variability. A major conclusion of our work is that the variance on EC2 is currently so high that wall clock experiments may only be performed with considerable care. To this end, we provide some hints to users.
Jörg Schad, Jens Dittrich, Jorge-Arnulfo Quiané-Ruiz
Proc. VLDB Endow.1
2009 On Analyzing the Database Performance for Different Classes of XML Documents based on the used Storage Approach
Hagen Höpfner, Jörg Schad, Essam Mansour 0001
ICSOFT (2)2