EDBT 2026 Demo / reviewers in the wild / expert
Aristeidis Mastoras
dblp:116/7075
· DBLP profile ↗
7ranked-venue papers
7as first author
1since 2021 · last 2023
0000-0002-5235-8499ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Parallel and multicore computing · 80% Electronic design automation · 14% Processor architecture and microarchitecture · 6% | |
| Databases, data mining, and information retrieval
1 paper |
Query processing and optimization · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › pipeline parallelism
dynamic linear pipeline |
0.8 | 2 | 2020 | Chunking for Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2020 Efficient and Scalable Execution of Fine-Grained Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2019 |
Parallel and multicore computing › task scheduling
pipeline scheduling |
0.8 | 2 | 2020 | Chunking for Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2020 Efficient and Scalable Execution of Fine-Grained Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2019 |
Query processing and optimization
lazy evaluation |
0.7 | 1 | 2023 | Design and Implementation for Nonblocking Execution in GraphBLAS: Tradeoffs and Performance · ACM Trans. Archit. Code Optim. 2023 |
Electronic design automation › high-level synthesis › pipeline synthesis
loop pipelining |
0.6 | 2 | 2018 | Unifying Fixed Code Mapping, Communication, Synchronization and Scheduling Algorithms for Efficient and Scalable Loop Pipelining · IEEE Trans. Parallel Distributed Syst. 2018 Unifying fixed code and fixed data mapping of load-imbalanced pipelined loops · PPoPP 2016 |
Parallel and multicore computing › task scheduling
dynamic scheduling |
0.3 | 1 | 2018 | Unifying Fixed Code Mapping, Communication, Synchronization and Scheduling Algorithms for Efficient and Scalable Loop Pipelining · IEEE Trans. Parallel Distributed Syst. 2018 |
Parallel and multicore computing
parallel programming models and runtimes |
0.3 | 1 | 2018 | Unifying Fixed Code Mapping, Communication, Synchronization and Scheduling Algorithms for Efficient and Scalable Loop Pipelining · IEEE Trans. Parallel Distributed Syst. 2018 |
Parallel and multicore computing
load balancing |
0.2 | 1 | 2016 | Unifying fixed code and fixed data mapping of load-imbalanced pipelined loops · PPoPP 2016 |
Processor architecture and microarchitecture › instruction scheduling
software pipelining |
0.2 | 1 | 2016 | Unifying fixed code and fixed data mapping of load-imbalanced pipelined loops · PPoPP 2016 |
Parallel and multicore computing › parallel scheduling
runtime scheduling |
0.2 | 2 | 2020 | Chunking for Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2020 Efficient and Scalable Execution of Fine-Grained Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2019 |
Parallel and multicore computing › locality optimization
data locality optimization |
0.2 | 1 | 2023 | Design and Implementation for Nonblocking Execution in GraphBLAS: Tradeoffs and Performance · ACM Trans. Archit. Code Optim. 2023 |
Parallel and multicore computing › synchronization
synchronization overhead |
0.1 | 1 | 2020 | Chunking for Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2020 |
Parallel and multicore computing
synchronization |
0.1 | 1 | 2019 | Efficient and Scalable Execution of Fine-Grained Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2019 |
Methods — techniques the papers use, named apart from their topics
dynamic data dependence analysis · 1.3analytic modeling · 1.3ticket mechanism · 0.7static scheduling · 0.7work-stealing · 0.3work stealing · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Design and Implementation for Nonblocking Execution in GraphBLAS: Tradeoffs and PerformanceabstractGraphBLASis a recent standard that allows the expression of graph algorithms in the language of linear algebra and enables automatic code parallelization and optimization. GraphBLAS operations are memory bound and may benefit from data locality optimizations enabled by nonblocking execution. However, nonblocking execution remains under-evaluated. In this article, we present a novel design and implementation that investigates nonblocking execution in GraphBLAS. Lazy evaluation enables runtime optimizations that improve data locality, and dynamic data dependence analysis identifies operations that may reuse data in cache. The nonblocking execution of an arbitrary number of operations results in dynamic parallelism, and the performance of the nonblocking execution depends on two parameters, which are automatically determined, at run-time, based on a proposed analytic model. The evaluation confirms the importance of nonblocking execution for various matrices of three algorithms, by showing up to 4.11× speedup over blocking execution as a result of better cache utilization. The proposed analytic model makes the nonblocking execution reach up to 5.13× speedup over the blocking execution. The fully automatic performance is very close to that obtained by using the best manual configuration for both small and large matrices. Finally, the evaluation includes a comparison with other state-of-the-art frameworks for numerical linear algebra programming that employ parallel execution and similar optimizations to those discussed in this work, and the presented nonblocking execution reaches up to 16.1× speedup over the state of the art. Aristeidis Mastoras, Sotiris Anagnostidis, Albert-Jan Nicholas Yzelman |
ACM Trans. Archit. Code Optim. | 1 |
| 2020 | Chunking for Dynamic Linear PipelinesabstractDynamic scheduling and dynamic creation of the pipeline structure are crucial for efficient execution of pipelined programs. Nevertheless, dynamic systems imply higher overhead than static systems. Therefore, chunking is the key to decrease the synchronization and scheduling overhead by grouping activities. We present a chunking algorithm for dynamic systems that handles dynamic linear pipelines, which allow the number and duration of stages to be determined at run-time. The evaluation on 44 cores shows that chunking brings the overhead of dynamic scheduling down to that of a static scheduler, and it enables efficient and scalable execution of fine-grained dynamic linear pipelines. Aristeidis Mastoras, Thomas R. Gross |
ACM Trans. Archit. Code Optim. | 1 |
| 2019 | Load-balancing for load-imbalanced fine-grained linear pipelinesabstractPipelining is a well-known technique to overlap loop iterations by partitioning the loop body into a sequence of stages. A large class of programs can be expressed as linear pipelines if data dependences only flow from earlier to later stages. Various pipelining techniques have been explored but reconciling load-balancing and efficient execution is still a challenge for two main reasons. First, partitioning of the loop body into stages that lead to load-balancing may depend on the data set as well as system properties, e.g., number of cores. Second, the configuration of the runtime system is far from obvious. In this article, we present Pipelight , a technique that achieves load-balancing for linear pipelines. Pipelight relies on a way of mapping stages onto threads that simplifies partitioning and enables the design of a lightweight algorithm for dynamic scheduling. Furthermore, Pipelight introduces a concurrent data structure that exploits the properties of data dependences presented in linear pipelines to provide efficient communication and synchronization. This data structure simplifies the configuration of the runtime system and makes Pipelight a practical solution. The evaluation on a 44-core system shows the efficiency of Pipelight for a set of programs selected from widely-used collections. Although Pipelight simplifies parallelization of linear pipelines, it performs similarly to the most efficient properly configured state-of-the-art technique. The price paid for the benefits of Pipelight is additional overhead for fine-grained loops. However, this overhead can be amortized successfully with chunking. To make Pipelight a promising solution, we propose a directive-based transformation for Pipelight , which is developed in a prototype source-to-source compiler. Consequently, Pipelight is an efficient and practical solution to achieve load-balancing for fine-grained linear pipelines. Aristeidis Mastoras, Thomas R. Gross |
Parallel Comput. | 1 |
| 2019 | Efficient and Scalable Execution of Fine-Grained Dynamic Linear PipelinesabstractWe present Pipelite , a dynamic scheduler that exploits the properties of dynamic linear pipelines to achieve high performance for fine-grained workloads. The flexibility of Pipelite allows the stages and their data dependences to be determined at runtime. Pipelite unifies communication, scheduling, and synchronization algorithms with suitable data structures. This unified design introduces the local suspension mechanism and a wait-free enqueue operation, which allow efficient dynamic scheduling. The evaluation on a 44-core machine, using programs from three widely used benchmark suites, shows that Pipelite implies low overhead and significantly outperforms the state of the art in terms of speedup, scalability, and memory usage. Aristeidis Mastoras, Thomas R. Gross |
ACM Trans. Archit. Code Optim. | 1 |
| 2018 | Unifying Fixed Code Mapping, Communication, Synchronization and Scheduling Algorithms for Efficient and Scalable Loop PipeliningabstractPipelining allows the execution of loop iterations with cross-iteration dependences to overlap in time, provided that the loop body is partitioned into stages such that the data dependences are not violated. Then, the stages are mapped onto threads and communication and synchronization between stages is typically achieved using queues. Pipelining techniques that rely on static scheduling perform poorly for load-imbalanced loops. Moreover, previous research efforts that achieve load-balancing are restricted to work-stealing and imply high overhead for fine-grained loops. In this article, we present URTS, a unified runtime system with compiler support that provides a lightweight dynamic scheduler by combining mapping, communication and synchronization algorithms with a suitable data structure and an efficient ticket mechanism. Particularly, URTS shows that it is possible to combine the efficiency of static scheduling with the load-imbalance tolerance of work-stealing by using a unified design that exploits the properties of a novel data structure. The evaluation on 8- and 32-core machines shows that URTS implies low overhead, of the same order as a static scheduler, for a set of benchmarks chosen from widely-used collections. URTS is a scalable solution that performs efficient dynamic scheduling for fine-grained loops, i.e., a class of interesting loops that is poorly handled by the state-of-the-art due to high overhead. Aristeidis Mastoras, Thomas R. Gross |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Unifying fixed code and fixed data mapping of load-imbalanced pipelined loopsabstractSome loops with cross-iteration dependences can execute in parallel by pipelining. The loop body is partitioned into stages such that the data dependences are not violated and then the stages are mapped onto threads. Two well-known mapping techniques are fixed code and fixed data; they achieve high performance for load-balanced loops, but they fail to perform well for load-imbalanced loops. In this article, we present a novel hybrid mapping that eliminates drawbacks of both prior mapping techniques and enables dynamic scheduling of stages. Aristeidis Mastoras, Thomas R. Gross |
PPoPP | 1 |
| 2015 | Ariadne - Directive-based parallelism extraction from recursive functions
Aristeidis Mastoras, George Manis |
J. Parallel Distributed Comput. | 1 |