EDBT 2026 Demo / reviewers in the wild / expert
Abdelhalim Amer
dblp:119/9063
· DBLP profile ↗
12ranked-venue papers
5as first author
0since 2021 · last 2020
0000-0001-5856-0172ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-authorComputer networks · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Parallel and multicore computing · 71% High-performance computing · 15% Memory systems · 11% | |
| Software engineering, system software, and programming languages
3 papers |
Operating systems · 82% Runtime systems and virtual machines · 18% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Operating systems › resource management › process management › CPU scheduling
thread scheduling |
0.8 | 2 | 2020 | Analyzing the Performance Trade-Off in Implementing User-Level Threads · IEEE Trans. Parallel Distributed Syst. 2020 Lessons learned from analyzing dynamic promotion for user-level threading · SC 2018 |
Operating systems › resource management › process management
user-level threads |
0.8 | 2 | 2020 | Analyzing the Performance Trade-Off in Implementing User-Level Threads · IEEE Trans. Parallel Distributed Syst. 2020 Lessons learned from analyzing dynamic promotion for user-level threading · SC 2018 |
Parallel and multicore computing
parallel programming runtimes |
0.3 | 1 | 2018 | Argobots: A Lightweight Low-Level Threading and Tasking Framework · IEEE Trans. Parallel Distributed Syst. 2018 |
Parallel and multicore computing › parallel programming models
message passing |
0.3 | 1 | 2017 | Why is MPI so slow?: analyzing the fundamental limits in implementing MPI-3.1 · SC 2017 |
Parallel and multicore computing › parallel programming models › message passing
MPI implementation |
0.3 | 1 | 2017 | Why is MPI so slow?: analyzing the fundamental limits in implementing MPI-3.1 · SC 2017 |
Memory systems
non-uniform memory access |
0.3 | 1 | 2017 | An Efficient Abortable-locking Protocol for Multi-level NUMA Systems · PPoPP 2017 |
Parallel and multicore computing
parallel programming models |
0.3 | 1 | 2017 | Why is MPI so slow?: analyzing the fundamental limits in implementing MPI-3.1 · SC 2017 |
High-performance computing
performance optimization at scale |
0.3 | 1 | 2017 | Why is MPI so slow?: analyzing the fundamental limits in implementing MPI-3.1 · SC 2017 |
Parallel and multicore computing
synchronization |
0.3 | 1 | 2017 | An Efficient Abortable-locking Protocol for Multi-level NUMA Systems · PPoPP 2017 |
Parallel and multicore computing
MPI |
0.2 | 1 | 2015 | MPI+Threads: runtime contention and remedies · PPoPP 2015 |
Parallel and multicore computing › parallel programming models › message passing
MPI runtime |
0.1 | 1 | 2017 | An Efficient Abortable-locking Protocol for Multi-level NUMA Systems · PPoPP 2017 |
Parallel and multicore computing › parallel computing › parallel communication
multithreaded communication |
0.1 | 1 | 2015 | MPI+Threads: runtime contention and remedies · PPoPP 2015 |
Methods — techniques the papers use, named apart from their topics
instruction-level analysis · 1.2cache-level analysis · 0.9user-level threading · 0.7tasking model · 0.7argobots · 0.7dynamic promotion · 0.3timeout · 0.3communication stack optimization · 0.3HMCS-T · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Analyzing the Performance Trade-Off in Implementing User-Level ThreadsabstractUser-level threads have been widely adopted as a means of achieving lightweight concurrent execution without the costs of OS-level threads. Nevertheless, the costs of managing user-level threads represent a performance barrier that dictates how fine grained the concurrency exposed by an application can be without incurring significant overheads; this in turn may translate into insufficient parallelism to exploit highly parallel systems. This article is a deep dive into the fundamental costs in implementing user-level threads. We first identify that one of the highest sources of fork-join overheads stems from deviations, events that incur context switching during the execution of a thread and disrupt a run-to-completion execution. We then conduct an in-depth investigation of a wide spectrum of methods with respect to how they handle deviations while covering both parent- and child-first scheduling policies. Our methodology involves a comprehensive instruction- and cache-level analysis of all methods on several modern CPU architectures. The primary finding of our evaluation is that dynamic promotion methods that assume the absence of deviation and dynamically provide context-switching support offer the best trade-off between performance and capability when the likelihood of deviation is low. Shintaro Iwasaki, Abdelhalim Amer, Kenjiro Taura, Pavan Balaji |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | BOLT: Optimizing OpenMP Parallel Regions with User-Level ThreadsabstractOpenMP is widely used by a number of applications, computational libraries, and runtime systems. As a result, multiple levels of the software stack use OpenMP independently of one another, often leading to nested parallel regions. Although exploiting such nested parallelism is a potential opportunity for performance improvement, it often causes destructive performance with leading OpenMP runtimes because of their reliance on heavyweight OS-level threads. User-level threads (ULTs) are more lightweight alternatives but existing ULT-based runtimes suffer from several shortcomings: 1) thread management costs remain significant and outweigh the benefits from additional parallelism; 2) the shift to ULTs often hurts the more common flat parallelism case; and 3) absence of user control over thread-to-CPU binding, a critical feature on modern systems. This paper presents BOLT, a practical ULT-based OpenMP runtime system that efficiently supports both flat and nested parallelism. This is accomplished on three fronts: 1) advanced data reuse and thread synchronization strategies; 2) thread coordination that adapts to the level of oversubscription; and 3) an implementation of the modern OpenMP thread-to-CPU binding interface tailored to ULT-based runtimes. The result is a highly optimized runtime that transparently achieves similar performance compared with leading state-of-the-art widely used OpenMP runtimes under flat parallelism, while outperforming all existing runtimes under nested parallelism. Shintaro Iwasaki, Abdelhalim Amer, Kenjiro Taura, Pavan Balaji |
PACT | 2 |
| 2019 | Software combining to mitigate multithreaded MPI contentionabstractEfforts to mitigate lock contention from concurrent threaded accesses to MPI have reduced contention through fine-grained locking, avoided locking altogether by offloading communication to dedicated threads, or alleviated negative side effects from contention by using better lock management protocols. The blocking nature of lock-based methods, however, wastes the asynchrony benefits of nonblocking MPI operations, and the offloading model sacrifices CPU resources and incurs unnecessary software offloading overheads under low contention. Abdelhalim Amer, Charles Archer, Michael Blocksome, Chongxiao Cao, Michael Chuvelev, Hajime Fujita 0002, María Jesús Garzarán, Yanfei Guo, Jeff R. Hammond, Shintaro Iwasaki, Kenneth Raffenetti, Mikhail Shiryaev, Min Si, Kenjiro Taura, Sagar Thapaliya, Pavan Balaji |
ICS | 1 |
| 2018 | Lessons learned from analyzing dynamic promotion for user-level threading
Shintaro Iwasaki, Abdelhalim Amer, Kenjiro Taura, Pavan Balaji |
SC | 2 |
| 2018 | 8th International Workshop on Programming Models and Applications for Multicores and Manycores (PMAM'17)
Abdelhalim Amer |
Parallel Comput. | 1 |
| 2018 | Argobots: A Lightweight Low-Level Threading and Tasking FrameworkabstractIn the past few decades, a number of user-level threading and tasking models have been proposed in the literature to address the shortcomings of OS-level threads, primarily with respect to cost and flexibility. Current state-of-the-art user-level threading and tasking models, however, either are too specific to applications or architectures or are not as powerful or flexible. In this paper, we present Argobots, a lightweight, low-level threading and tasking framework that is designed as a portable and performant substrate for high-level programming models or runtime systems. Argobots offers a carefully designed execution model that balances generality of functionality with providing a rich set of controls to allow specialization by end users or high-level programming models. We describe the design, implementation, and performance characterization of Argobots and present integrations with three high-level models: OpenMP, MPI, and colocated I/O services. Evaluations show that (1) Argobots, while providing richer capabilities, is competitive with existing simpler generic threading runtimes; (2) our OpenMP runtime offers more efficient interoperability capabilities than production OpenMP runtimes do; (3) when MPI interoperates with Argobots instead of Pthreads, it enjoys reduced synchronization costs and better latency-hiding capabilities; and (4) I/O services with Argobots reduce interference with colocated applications while achieving performance competitive with that of a Pthreads approach. Abdelhalim Amer, Pavan Balaji, Cyril Bordage, George Bosilca, Alex Brooks, Philip H. Carns, Adrián Castelló 0001, Damien Genet, Thomas Hérault, Shintaro Iwasaki, Prateek Jindal, Laxmikant V. Kalé, Sriram Krishnamoorthy, Jonathan Lifflander, Huiwei Lu, Esteban Meneses, Marc Snir, Yanhua Sun, Kenjiro Taura, Pete Beckman |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Advanced Thread Synchronization for Multithreaded MPI ImplementationsabstractConcurrent multithreaded access to the Message Passing Interface (MPI) is gaining importance to support emerging hybrid MPI applications. The interoperability between threads and MPI, however, is complex and renders efficient implementations nontrivial. Prior studies showed that threads waiting for communication progress (waiting threads) often interfere with others (active threads) and degrade their progress. This situation occurs when both classes of threads compete for the same MPI resource and ownership passing to waiting threads does not guarantee communication to advance. The best-known practical solution prioritizes active threads and adapts first-in-first-out arbitration within each class. This approach, however, suffers from residual wasted resource acquisitions (waste) and ignores data locality, thus resulting in poor scalability. In this work, we propose thread synchronization improvements to eliminate waste while preserving data locality in a production MPI implementation. First, we leverage MPI knowledge and a fast synchronization method to eliminate waste and accelerate progress. Second, we rely on a cooperative progress model that dynamically elects and restricts a single waiting thread to drive a communication context for improved data locality. Third, we prioritize active threads and synchronize them with a locality-preserving lock that is hierarchical and exploits unbounded bias for high throughput. Results show significant improvement in synthetic microbenchmarks and two MPI+OpenMP applications. Hoang-Vu Dang, Abdelhalim Amer, Pavan Balaji |
CCGrid | 3 |
| 2017 | An Efficient Abortable-locking Protocol for Multi-level NUMA SystemsabstractThe popularity of Non-Uniform Memory Access (NUMA) architectures has led to numerous locality-preserving hierarchical lock designs, such as HCLH, HMCS, and cohort locks. Locality-preserving locks trade fairness for higher throughput. Hence, some instances of acquisitions can incur long latencies, which may be intolerable for certain applications. Few locks admit a waiting thread to abandon its protocol on a timeout. State-of-the-art abortable locks are not fully locality aware, introduce high overheads, and unsuitable for frequent aborts. Enhancing locality-aware locks with lightweight timeout capability is critical for their adoption. In this paper, we design and evaluate the HMCS-T lock, a Hierarchical MCS (HMCS) lock variant that admits a timeout. HMCS-T maintains the locality benefits of HMCS while ensuring aborts to be lightweight. HMCS-T offers the progress guarantee missing in most abortable queuing locks. Our evaluations show that HMCS-T offers the timeout feature at a moderate overhead over its HMCS analog. HMCS-T, used in an MPI runtime lock, mitigated the poor scalability of an MPI+OpenMP BFS code and resulted in 4.3x superior scaling. Milind Chabbi, Abdelhalim Amer, Shasha Wen, Xu Liu 0001 |
PPoPP | 2 |
| 2017 | Why is MPI so slow?: analyzing the fundamental limits in implementing MPI-3.1abstractThis paper provides an in-depth analysis of the software overheads in the MPI performance-critical path and exposes mandatory performance overheads that are unavoidable based on the MPI-3.1 specification. We first present a highly optimized implementation of the MPI-3.1 standard in which the communication stack---all the way from the application to the low-level network communication API---takes only a few tens of instructions. We carefully study these instructions and analyze the root cause of the overheads based on specific requirements from the MPI standard that are unavoidable under the current MPI standard. We recommend potential changes to the MPI standard that can minimize these overheads. Our experimental results on a variety of network architectures and applications demonstrate significant benefits from our proposed changes. Kenneth Raffenetti, Abdelhalim Amer, Lena Oden, Charles Archer, Wesley Bland, Hajime Fujita 0002, Yanfei Guo, Tomislav Janjusic, Dmitry Durnov, Michael Blocksome, Min Si, Akhil Langer, Gengbin Zheng, Masamichi Takagi, Paul K. Coffman, Sayantan Sur, Alexander Sannikov, Sergey Oblomov, Michael Chuvelev, Masayuki Hatanaka, Paul F. Fischer, Thilina Ratnayaka, Matthew Otten, Misun Min, Pavan Balaji |
SC | 2 |
| 2015 | Characterizing MPI and Hybrid MPI+Threads Applications at Scale: Case Study with BFSabstractWith the increasing prominence of many-core architectures and decreasing per-core resources on large supercomputers, a number of applications developers are investigating the use of hybrid MPI+threads programming to utilize computational units while sharing memory. An MPI-only model that uses one MPI process per system core is capable of effectively utilizing the processing units, but it fails to fully utilize the memory hierarchy and relies on fine-grained internodes communication. Hybrid MPI+threads models, on the other hand, can handle internodes parallelism more effectively and alleviate some of the overheads associated with internodes communication by allowing more coarse-grained data movement between address spaces. The hybrid model, however, can suffer from locking and memory consistency overheads associated with data sharing. In this paper, we use a distributed implementation of the breadth-first search algorithm in order to understand the performance characteristics of MPI-only and MPI+threads models at scale. We start with a baseline MPI-only implementation and propose MPI+threads extensions where threads independently communicate with remote processes while cooperating for local computation. We demonstrate how the coarse-grained communication of MPI+threads considerably reduces time and space overheads that grow with the number of processes. At large scale, however, these overheads constitute performance barriers for both models and require fixing the root causes, such as the excessive polling for communication progress and inefficient global synchronizations. To this end, we demonstrate various techniques to reduce such overheads and show performance improvements on up to 512K cores of a Blue Gene/Q system. Abdelhalim Amer, Huiwei Lu, Pavan Balaji, Satoshi Matsuoka |
CCGRID | 1 |
| 2015 | MPI+Threads: runtime contention and remediesabstractHybrid MPI+Threads programming has emerged as an alternative model to the “MPI everywhere” model to better handle the increasing core density in cluster nodes. While the MPI standard allows multithreaded concurrent communication, such flexibility comes with the cost of maintaining thread safety within the MPI implementation, typically implemented using critical sections. In contrast to previous works that studied the importance of critical-section granularity in MPI implementations, in this paper we investigate the implication of critical-section arbitration on communication performance. We first analyze the MPI runtime when multithreaded concurrent communication takes place on hierarchical memory systems. Our results indicate that the mutex-based approach that most MPI implementations use today can incur performance penalties due to unfair arbitration. We then present methods to mitigate these penalties with a first-come, first-served arbitration and a priority locking scheme that favors threads doing useful work. Through evaluations using several benchmarks and applications, we demonstrate up to 5-fold improvement in performance. Abdelhalim Amer, Huiwei Lu, Yanjie Wei, Pavan Balaji, Satoshi Matsuoka |
PPoPP | 1 |
| 2012 | Using Bittorrent and SVC for efficient video sharing and streamingabstractMassive and large scale content distribution over Internet is attracting a lot of research efforts as many challenges remain to be solved. Recent studies show that Internet video including video-to-TV and video calling is dominating the Internet traffic. As Internet becomes widely accessible to wired, mobile and wireless users, it is important to design a system that can ensure video streaming across variable network conditions while simultaneously handling devices and end-user heterogeneities. Most of the proposed solutions, such as CDN and peer-to-peer (P2P), solve the scalability problem but fail to handle receiver's heterogeneity. In this paper, we combine P2P network and SVC (Scalable Video Coding) to provide an efficient video sharing and streaming system. Our solution consists of an SVC layered extension of the widely used Bittorrent protocol to support real-time content delivery with different video qualities given the receivers capabilities. Thus, we propose different optimization techniques to organize peers in an overlay. The results, obtained by means of simulation, show that our system outperforms solutions that relay on single layer streams such as AVC (Advanced Video Coding) and this in terms of receivers perceived QoS. Abdelhalim Amer, Ahmed Touflk, Walid-Khaled Hidouci, Satoshi Matsuoka |
ISCC | 1 |