Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Stefan Bora

dblp:271/9943 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Performance modeling and evaluation · 56% Parallel and multicore computing · 37% Cloud and datacenter computing · 6%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation › queueing models › queueing network model
fork-join systems
1.122023
The Tiny-Tasks Granularity Trade-Off: Balancing Overhead Versus Performance in Parallel Systems · IEEE Trans. Parallel Distributed Syst. 2023
Tiny Tasks - A Remedy for Synchronization Constraints in Multi-Server Systems · INFOCOM 2020
Performance modeling and evaluation
queueing models
1.122023
The Tiny-Tasks Granularity Trade-Off: Balancing Overhead Versus Performance in Parallel Systems · IEEE Trans. Parallel Distributed Syst. 2023
Tiny Tasks - A Remedy for Synchronization Constraints in Multi-Server Systems · INFOCOM 2020
Parallel and multicore computing
task granularity
1.122023
The Tiny-Tasks Granularity Trade-Off: Balancing Overhead Versus Performance in Parallel Systems · IEEE Trans. Parallel Distributed Syst. 2023
Tiny Tasks - A Remedy for Synchronization Constraints in Multi-Server Systems · INFOCOM 2020
Parallel and multicore computing
task scheduling
0.712023
The Tiny-Tasks Granularity Trade-Off: Balancing Overhead Versus Performance in Parallel Systems · IEEE Trans. Parallel Distributed Syst. 2023
Parallel and multicore computing
parallel scheduling
0.612022
Performance and Scaling of Parallel Systems with Blocking Start and/or Departure Barriers · INFOCOM 2022
Performance modeling and evaluation
queueing analysis
0.612022
Performance and Scaling of Parallel Systems with Blocking Start and/or Departure Barriers · INFOCOM 2022
Performance modeling and evaluation › stability analysis
stability region
0.612022
Performance and Scaling of Parallel Systems with Blocking Start and/or Departure Barriers · INFOCOM 2022
Performance modeling and evaluation › network performance analysis
stochastic network calculus
0.412020
Tiny Tasks - A Remedy for Synchronization Constraints in Multi-Server Systems · INFOCOM 2020
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.212022
Performance and Scaling of Parallel Systems with Blocking Start and/or Departure Barriers · INFOCOM 2022
Parallel and multicore computing › parallel programming models
degree of parallelism
0.212022
Performance and Scaling of Parallel Systems with Blocking Start and/or Departure Barriers · INFOCOM 2022
Cloud and datacenter computing
cluster resource management and scheduling
0.112020
Tiny Tasks - A Remedy for Synchronization Constraints in Multi-Server Systems · INFOCOM 2020
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
mapreduce scheduling
0.112020
Tiny Tasks - A Remedy for Synchronization Constraints in Multi-Server Systems · INFOCOM 2020

Methods — techniques the papers use, named apart from their topics

queueing theory · 1.2analytical modeling · 1.2simulation · 1.1stochastic network calculus · 0.4
YearPublicationVenuePosition
2023 The Tiny-Tasks Granularity Trade-Off: Balancing Overhead Versus Performance in Parallel Systems
abstract
Models of parallel processing systems typically assume that one has$l$workers and jobs are split into an equal number of$k=l$tasks. Splitting jobs into$k > l$smaller tasks, i.e. using “tiny tasks”, can yield performance and stability improvements because it reduces the variance in the amount of work assigned to each worker, but as$k$increases, the overhead involved in scheduling and managing the tasks begins to overtake the performance benefit. We perform extensive experiments on the effects of task granularity on an Apache Spark cluster, and based on these, develop a four-parameter model for task and job overhead that, in simulation, produces sojourn time distributions that match those of the real system. We also present analytical results which illustrate how using tiny tasks improves the stability region of split-merge systems, and analytical bounds on the sojourn and waiting time distributions of both split-merge and single-queue fork-join systems with tiny tasks. Finally we combine the overhead model with the analytical models to produce an analytical approximation to the sojourn and waiting time distributions of systems with tiny tasks which include overhead. We also perform analogous tiny-tasks experiments on a hybrid multi-processor shared memory system based on MPI and OpenMP which has no load-balancing between nodes. Though no longer strict analytical bounds, our analytical approximations with overhead match both the Spark and MPI/OpenMP experimental results very well.
Stefan Bora, Brenton D. Walker, Markus Fidler
IEEE Trans. Parallel Distributed Syst.1
2022 Performance and Scaling of Parallel Systems with Blocking Start and/or Departure Barriers
abstract
Parallel systems divide jobs into smaller tasks that can be serviced by many workers at the same time. Some parallel systems have blocking barriers that require all of their tasks to start and/or depart in unison. This is true of many parallelized machine learning workloads, and the popular Apache Spark processing engine has recently added support for Barrier Execution Mode, which allows users to add such barriers to their jobs. The drawback of these barriers is reduced performance and stability compared to equivalent non-blocking systems.We derive analytical expressions for the stability regions for parallel systems with blocking start and/or departure barriers. We extend results from queueing theory to derive waiting and sojourn time bounds for systems with blocking start barriers. Our results show that for a given system utilization and number of servers, there is an optimal degree of parallelism that balances waiting time and job execution time. This observation leads us to propose and implement a class of self-adaptive schedulers, we call "Take-Half", that modulate the allowed degree of parallelism based on the instantaneous system load, improving mean performance and eliminating stability issues.
Brenton D. Walker, Stefan Bora, Markus Fidler
INFOCOM2
2020 Tiny Tasks - A Remedy for Synchronization Constraints in Multi-Server Systems
abstract
Models of parallel processing systems typically assume that one has l servers and jobs are split into an equal number of k = l tasks. This seemingly simple approximation has surprisingly large consequences for the resulting stability and performance bounds. In reality, best practices for modern mapreduce systems indicate that a job's partitioning factor should be much larger than the number of servers available, with some researchers going to far as to advocate for a "tiny tasks" regime, where jobs are split into over 10,000 tasks. In this paper we use recent advances in stochastic network calculus to fundamentally understand the effects of task granularity on parallel systems' scaling, stability, and performance. For the split-merge model, we show that when one allows for tiny tasks, the stability region is actually much better than had previously been concluded. For the single-queue fork-join model, we show that sojourn times quickly approach the optimal case when l "big tasks" are subdivided into k≫ l "tiny tasks". Our results are validated using extensive simulations, and the applicability of the models used is validated by experPiments on an Apache Spark cluster.
Markus Fidler, Brenton D. Walker, Stefan Bora
INFOCOM3