Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Alexandros Tzannes

dblp:64/7750 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
0since 2021 · last 2019
0000-0002-4984-2036ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 100%
Software engineering, system software, and programming languages
1 paper
Concurrent programming · 40% Programming languages and type systems · 20% Program analysis · 20%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel scheduling
adaptive scheduling
0.322014
Lazy Scheduling: A Runtime Adaptive Scheduler for Declarative Parallelism · ACM Trans. Program. Lang. Syst. 2014
Lazy binary-splitting: a run-time adaptive work-stealing scheduler · PPoPP 2010
Parallel and multicore computing › load balancing › dynamic load balancing
work stealing
0.322014
Lazy Scheduling: A Runtime Adaptive Scheduler for Declarative Parallelism · ACM Trans. Program. Lang. Syst. 2014
Lazy binary-splitting: a run-time adaptive work-stealing scheduler · PPoPP 2010
Program verification
annotation inference
0.212015
Region and Effect Inference for Safe Parallelism (T) · ASE 2015
Concurrent programming › deterministic execution
deterministic parallelism
0.212015
Region and Effect Inference for Safe Parallelism (T) · ASE 2015
Concurrent programming › concurrency correctness
safe parallelism
0.212015
Region and Effect Inference for Safe Parallelism (T) · ASE 2015
Program analysis
static analysis
0.212015
Region and Effect Inference for Safe Parallelism (T) · ASE 2015
Programming languages and type systems
type inference
0.212015
Region and Effect Inference for Safe Parallelism (T) · ASE 2015
Parallel and multicore computing › parallel programming models
declarative parallelism
0.212014
Lazy Scheduling: A Runtime Adaptive Scheduler for Declarative Parallelism · ACM Trans. Program. Lang. Syst. 2014
Parallel and multicore computing
parallel programming models and runtimes
0.212014
Lazy Scheduling: A Runtime Adaptive Scheduler for Declarative Parallelism · ACM Trans. Program. Lang. Syst. 2014
Parallel and multicore computing
task scheduling
0.212014
Lazy Scheduling: A Runtime Adaptive Scheduler for Declarative Parallelism · ACM Trans. Program. Lang. Syst. 2014
Parallel and multicore computing
parallel programming models
0.112010
Lazy binary-splitting: a run-time adaptive work-stealing scheduler · PPoPP 2010
Parallel and multicore computing › parallel scheduling
runtime scheduling
0.112010
Lazy binary-splitting: a run-time adaptive work-stealing scheduler · PPoPP 2010
Parallel and multicore computing › parallelization strategies
nested parallelism
0.012010
Lazy binary-splitting: a run-time adaptive work-stealing scheduler · PPoPP 2010

Methods — techniques the papers use, named apart from their topics

constraint satisfaction · 0.2load balancing · 0.2dynamic scheduling · 0.2lazy binary splitting · 0.1eager binary splitting · 0.1
YearPublicationVenuePosition
2019 Region and effect inference for safe parallelism
Alexandros Tzannes, Stephen Heumann, Lamyaa Eloussi, Mohsen Vakilian, Vikram S. Adve, Michael Han 0003
Autom. Softw. Eng.1
2015 Scalable Task Scheduling and Synchronization Using Hierarchical Effects
abstract
Several concurrent programming models that give strong safety guarantees employ effect specifications that indicate what effects on shared state a piece of code may perform. These specifications can be much more expressive than traditional synchronization mechanisms like locks, and they are amenable to static and/or dynamic checking approaches for ensuring safety properties. The Tasks With Effects (TWE) programming model uses dynamic checking to give nearly the strongest safety guarantees of any existing shared memory language while providing the flexibility to express both structured and unstructured concurrency. Like several other systems, TWE's effect specifications use hierarchical memory regions, which can naturally model nested and modular data structures and allow effects to be expressed at different levels of granularity in different parts of a program. To implement a programming model like TWE with high performance, particularly for programs with many fine-grain tasks, the run-time task scheduler must employ an algorithm that can enforce task isolation (mutual exclusion of tasks with conflicting effects) with low overhead and high scalability. This paper describes such an algorithm for TWE. It uses a scheduling tree designed to take advantage of the hierarchical structure of TWE effects, obtaining two key properties that lead to high scalability: (a) effects need to be compared only for ancestor and descendant nodes in the tree, and not any other nodes, and (b) the scheduler can use fine-grain locking of tree nodes to enable highly concurrent scheduling operations. We prove formally that the algorithm guarantees task isolation. Experimental results with a range of programs show that the algorithm provides very good scalability, even with fine-grain tasks.
Stephen Heumann, Alexandros Tzannes, Vikram S. Adve
PACT2
2015 Region and Effect Inference for Safe Parallelism (T)
abstract
In this paper, we present the first full regions-and-effects inference algorithm for explicitly parallel fork-join programs. We infer annotations inspired by Deterministic Parallel Java (DPJ) for a type-safe subset of C++. We chose the DPJ annotations because they give the strongest safety guarantees of any existing concurrency-checking approach we know of, static or dynamic, and it is also the most expressive static checking system we know of that gives strong safety guarantees. This expressiveness, however, makes manual annotation difficult and tedious, which motivates the need for automatic inference, but it also makes the inference problem very challenging: the code may use region polymorphism, imperative updates with complex aliasing, arbitrary recursion, hierarchical region specifications, and wildcard elements to describe potentially infinite sets of regions. We express the inference as a constraint satisfaction problem and develop, implement, and evaluate an algorithm for solving it. The region and effect annotations inferred by the algorithm constitute a checkable proof of safe parallelism, and it can be recorded both for documentation and for fast and modular safety checking.
Alexandros Tzannes, Stephen Heumann, Lamyaa Eloussi, Mohsen Vakilian, Vikram S. Adve, Michael Han 0003
ASE1
2014 Lazy Scheduling: A Runtime Adaptive Scheduler for Declarative Parallelism
abstract
Lazy scheduling is a runtime scheduler for task-parallel codes that effectively coarsens parallelism on load conditions in order to significantly reduce its overheads compared to existing approaches, thus enabling the efficient execution of more fine-grained tasks. Unlike other adaptive dynamic schedulers, lazy scheduling does not maintain any additional state to infer system load and does not make irrevocable serialization decisions. These two features allow it to scale well and to provide excellent load balancing in practice but at a much lower overhead cost compared to work stealing, the golden standard of dynamic schedulers. We evaluate three variants of lazy scheduling on a set of benchmarks on three different platforms and find it to substantially outperform popular work stealing implementations on fine-grained codes. Furthermore, we show that the vast performance gap between manually coarsened and fully parallel code is greatly reduced by lazy scheduling, and that, with minimal static coarsening, lazy scheduling delivers performance very close to that of fully tuned code. The tedious manual coarsening required by the best existing work stealing schedulers and its damaging effect on performance portability have kept novice and general-purpose programmers from parallelizing their codes. Lazy scheduling offers the foundation for a declarative parallel programming methodology that should attract those programmers by minimizing the need for manual coarsening and by greatly enhancing the performance portability of parallel code.
Alexandros Tzannes, George C. Caragea, Uzi Vishkin, Rajeev Barua
ACM Trans. Program. Lang. Syst.1
2011 Improving Run-Time Scheduling for General-Purpose Parallel Code
abstract
Summary form only given. Today, almost all desktop and laptop computers are shared-memory multicores, but the code they run is over whelmingly serial. High level language extensions and libraries (e.g., OpenMP, Cilk++, TBB) make it much easier for programmers to write parallel code than previous approaches (e.g., MPI), in large part thanks to the efficient work-stealing scheduler that allows the programmer to expose more parallelism than the actual hardware parallelism. But when the parallel tasks are too short or too many, the scheduling overheads become significant and hurt performance. Because this happens frequently (e.g, data-parallelism, PRAM algorithms), programmers need to manually coarsen tasks for performance by combining many of them into longer tasks.
Alexandros Tzannes, Rajeev Barua, Uzi Vishkin
PACT1
2010 Resource-Aware Compiler Prefetching for Many-Cores
abstract
Super-scalar, out-of-order processors that can have tens of read and write requests in the execution window place significant demands on Memory Level Parallelism (MLP). Multi-and many-cores with shared parallel caches further increase MLP demand. Current cache hierarchies however have been unable to keep up with this trend, with modern designs allowing only 4-16 concurrent cache misses. This disconnect is exacerbated by recent highly parallel architectures (e.g. GPUs) where power and area per-core budget favor lighter cores with less resources. Support for hardware and software prefetch increase MLP pressure since these techniques overlap multiple memory requests with existing computation. In this paper, we propose and evaluate a novel Resource-Aware Prefetching (RAP) compiler algorithm that is aware of the number of simultaneous prefetches supported, and optimized for the same. We show that in situations where not enough resources are available to issue prefetch instructions for all references in a loop, it is more beneficial to decrease the prefetch distance and prefetch for as many references as possible, rather than use a fixed prefetched distance and skip prefetching for some references, as in current approaches. We implemented our algorithm in a GCC-derived compiler and evaluated its performance using an emerging fine-grained many-core architecture. Our results show that the RAP algorithm outperforms a well-known loop prefetching algorithm by up to 40.15% and the state-of-the art GCC implementation by up to 34.79%. Moreover, we compare the RAP algorithm with a simple hardware prefetching mechanism, and show improvements of up to 24.61%.
George C. Caragea, Alexandros Tzannes, Fuat Keceli, Rajeev Barua, Uzi Vishkin
ISPDC2
2010 Lazy binary-splitting: a run-time adaptive work-stealing scheduler
abstract
We present Lazy Binary Splitting (LBS), a user-level scheduler of nested parallelism for shared-memory multiprocessors that builds on existing Eager Binary Splitting work-stealing (EBS) implemented in Intel's Threading Building Blocks (TBB), but improves performance and ease-of-programming. In its simplest form (SP), EBS requires manual tuning by repeatedly running the application under carefully controlled conditions to determine a stop-splitting-threshold (sst)for every do-all loop in the code. This threshold limits the parallelism and prevents excessive overheads for fine-grain parallelism. Besides being tedious, this tuning also over-fits the code to some particular dataset, platform and calling context of the do-all loop, resulting in poor performance portability for the code. LBS overcomes both the performance portability and ease-of-programming pitfalls of a manually fixed threshold by adapting dynamically to run-time conditions without requiring tuning.
Alexandros Tzannes, George C. Caragea, Rajeev Barua, Uzi Vishkin
PPoPP1