Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Aristeidis Mastoras

dblp:116/7075 · DBLP profile ↗
← Back
7ranked-venue papers
7as first author
1since 2021 · last 2023
0000-0002-5235-8499ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 7 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 80% Electronic design automation · 14% Processor architecture and microarchitecture · 6%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › pipeline parallelism
dynamic linear pipeline
0.822020
Chunking for Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2020
Efficient and Scalable Execution of Fine-Grained Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2019
Parallel and multicore computing › task scheduling
pipeline scheduling
0.822020
Chunking for Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2020
Efficient and Scalable Execution of Fine-Grained Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2019
Query processing and optimization
lazy evaluation
0.712023
Design and Implementation for Nonblocking Execution in GraphBLAS: Tradeoffs and Performance · ACM Trans. Archit. Code Optim. 2023
Electronic design automation › high-level synthesis › pipeline synthesis
loop pipelining
0.622018
Unifying Fixed Code Mapping, Communication, Synchronization and Scheduling Algorithms for Efficient and Scalable Loop Pipelining · IEEE Trans. Parallel Distributed Syst. 2018
Unifying fixed code and fixed data mapping of load-imbalanced pipelined loops · PPoPP 2016
Parallel and multicore computing › task scheduling
dynamic scheduling
0.312018
Unifying Fixed Code Mapping, Communication, Synchronization and Scheduling Algorithms for Efficient and Scalable Loop Pipelining · IEEE Trans. Parallel Distributed Syst. 2018
Parallel and multicore computing
parallel programming models and runtimes
0.312018
Unifying Fixed Code Mapping, Communication, Synchronization and Scheduling Algorithms for Efficient and Scalable Loop Pipelining · IEEE Trans. Parallel Distributed Syst. 2018
Parallel and multicore computing
load balancing
0.212016
Unifying fixed code and fixed data mapping of load-imbalanced pipelined loops · PPoPP 2016
Processor architecture and microarchitecture › instruction scheduling
software pipelining
0.212016
Unifying fixed code and fixed data mapping of load-imbalanced pipelined loops · PPoPP 2016
Parallel and multicore computing › parallel scheduling
runtime scheduling
0.222020
Chunking for Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2020
Efficient and Scalable Execution of Fine-Grained Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2019
Parallel and multicore computing › locality optimization
data locality optimization
0.212023
Design and Implementation for Nonblocking Execution in GraphBLAS: Tradeoffs and Performance · ACM Trans. Archit. Code Optim. 2023
Parallel and multicore computing › synchronization
synchronization overhead
0.112020
Chunking for Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2020
Parallel and multicore computing
synchronization
0.112019
Efficient and Scalable Execution of Fine-Grained Dynamic Linear Pipelines · ACM Trans. Archit. Code Optim. 2019

Methods — techniques the papers use, named apart from their topics

dynamic data dependence analysis · 1.3analytic modeling · 1.3ticket mechanism · 0.7static scheduling · 0.7work-stealing · 0.3work stealing · 0.3
YearPublicationVenuePosition
2023 Design and Implementation for Nonblocking Execution in GraphBLAS: Tradeoffs and Performance
abstract
GraphBLASis a recent standard that allows the expression of graph algorithms in the language of linear algebra and enables automatic code parallelization and optimization. GraphBLAS operations are memory bound and may benefit from data locality optimizations enabled by nonblocking execution. However, nonblocking execution remains under-evaluated. In this article, we present a novel design and implementation that investigates nonblocking execution in GraphBLAS. Lazy evaluation enables runtime optimizations that improve data locality, and dynamic data dependence analysis identifies operations that may reuse data in cache. The nonblocking execution of an arbitrary number of operations results in dynamic parallelism, and the performance of the nonblocking execution depends on two parameters, which are automatically determined, at run-time, based on a proposed analytic model. The evaluation confirms the importance of nonblocking execution for various matrices of three algorithms, by showing up to 4.11× speedup over blocking execution as a result of better cache utilization. The proposed analytic model makes the nonblocking execution reach up to 5.13× speedup over the blocking execution. The fully automatic performance is very close to that obtained by using the best manual configuration for both small and large matrices. Finally, the evaluation includes a comparison with other state-of-the-art frameworks for numerical linear algebra programming that employ parallel execution and similar optimizations to those discussed in this work, and the presented nonblocking execution reaches up to 16.1× speedup over the state of the art.
Aristeidis Mastoras, Sotiris Anagnostidis, Albert-Jan Nicholas Yzelman
ACM Trans. Archit. Code Optim.1
2020 Chunking for Dynamic Linear Pipelines
abstract
Dynamic scheduling and dynamic creation of the pipeline structure are crucial for efficient execution of pipelined programs. Nevertheless, dynamic systems imply higher overhead than static systems. Therefore, chunking is the key to decrease the synchronization and scheduling overhead by grouping activities. We present a chunking algorithm for dynamic systems that handles dynamic linear pipelines, which allow the number and duration of stages to be determined at run-time. The evaluation on 44 cores shows that chunking brings the overhead of dynamic scheduling down to that of a static scheduler, and it enables efficient and scalable execution of fine-grained dynamic linear pipelines.
Aristeidis Mastoras, Thomas R. Gross
ACM Trans. Archit. Code Optim.1
2019 Load-balancing for load-imbalanced fine-grained linear pipelines
abstract
Pipelining is a well-known technique to overlap loop iterations by partitioning the loop body into a sequence of stages. A large class of programs can be expressed as linear pipelines if data dependences only flow from earlier to later stages. Various pipelining techniques have been explored but reconciling load-balancing and efficient execution is still a challenge for two main reasons. First, partitioning of the loop body into stages that lead to load-balancing may depend on the data set as well as system properties, e.g., number of cores. Second, the configuration of the runtime system is far from obvious. In this article, we present Pipelight , a technique that achieves load-balancing for linear pipelines. Pipelight relies on a way of mapping stages onto threads that simplifies partitioning and enables the design of a lightweight algorithm for dynamic scheduling. Furthermore, Pipelight introduces a concurrent data structure that exploits the properties of data dependences presented in linear pipelines to provide efficient communication and synchronization. This data structure simplifies the configuration of the runtime system and makes Pipelight a practical solution. The evaluation on a 44-core system shows the efficiency of Pipelight for a set of programs selected from widely-used collections. Although Pipelight simplifies parallelization of linear pipelines, it performs similarly to the most efficient properly configured state-of-the-art technique. The price paid for the benefits of Pipelight is additional overhead for fine-grained loops. However, this overhead can be amortized successfully with chunking. To make Pipelight a promising solution, we propose a directive-based transformation for Pipelight , which is developed in a prototype source-to-source compiler. Consequently, Pipelight is an efficient and practical solution to achieve load-balancing for fine-grained linear pipelines.
Aristeidis Mastoras, Thomas R. Gross
Parallel Comput.1
2019 Efficient and Scalable Execution of Fine-Grained Dynamic Linear Pipelines
abstract
We present Pipelite , a dynamic scheduler that exploits the properties of dynamic linear pipelines to achieve high performance for fine-grained workloads. The flexibility of Pipelite allows the stages and their data dependences to be determined at runtime. Pipelite unifies communication, scheduling, and synchronization algorithms with suitable data structures. This unified design introduces the local suspension mechanism and a wait-free enqueue operation, which allow efficient dynamic scheduling. The evaluation on a 44-core machine, using programs from three widely used benchmark suites, shows that Pipelite implies low overhead and significantly outperforms the state of the art in terms of speedup, scalability, and memory usage.
Aristeidis Mastoras, Thomas R. Gross
ACM Trans. Archit. Code Optim.1
2018 Unifying Fixed Code Mapping, Communication, Synchronization and Scheduling Algorithms for Efficient and Scalable Loop Pipelining
abstract
Pipelining allows the execution of loop iterations with cross-iteration dependences to overlap in time, provided that the loop body is partitioned into stages such that the data dependences are not violated. Then, the stages are mapped onto threads and communication and synchronization between stages is typically achieved using queues. Pipelining techniques that rely on static scheduling perform poorly for load-imbalanced loops. Moreover, previous research efforts that achieve load-balancing are restricted to work-stealing and imply high overhead for fine-grained loops. In this article, we present URTS, a unified runtime system with compiler support that provides a lightweight dynamic scheduler by combining mapping, communication and synchronization algorithms with a suitable data structure and an efficient ticket mechanism. Particularly, URTS shows that it is possible to combine the efficiency of static scheduling with the load-imbalance tolerance of work-stealing by using a unified design that exploits the properties of a novel data structure. The evaluation on 8- and 32-core machines shows that URTS implies low overhead, of the same order as a static scheduler, for a set of benchmarks chosen from widely-used collections. URTS is a scalable solution that performs efficient dynamic scheduling for fine-grained loops, i.e., a class of interesting loops that is poorly handled by the state-of-the-art due to high overhead.
Aristeidis Mastoras, Thomas R. Gross
IEEE Trans. Parallel Distributed Syst.1
2016 Unifying fixed code and fixed data mapping of load-imbalanced pipelined loops
abstract
Some loops with cross-iteration dependences can execute in parallel by pipelining. The loop body is partitioned into stages such that the data dependences are not violated and then the stages are mapped onto threads. Two well-known mapping techniques are fixed code and fixed data; they achieve high performance for load-balanced loops, but they fail to perform well for load-imbalanced loops. In this article, we present a novel hybrid mapping that eliminates drawbacks of both prior mapping techniques and enables dynamic scheduling of stages.
Aristeidis Mastoras, Thomas R. Gross
PPoPP1
2015 Ariadne - Directive-based parallelism extraction from recursive functions
Aristeidis Mastoras, George Manis
J. Parallel Distributed Comput.1