EDBT 2026 Demo / reviewers in the wild / expert
Ahmad Afsahi
dblp:11/1112
· DBLP profile ↗
44ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-2924-6851ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable, Topology- and Multi-HCA-Aware Hierarchical GPU Allgather Using Parallel Rings
Amirreza Barati Sedeh, Ryan Grant, Ahmad Afsahi |
CCGrid | 3 |
| 2025 | Cascade: a Collaborative Algorithm for Scalable and Efficient Neighborhood AllgatherabstractNeighborhood collectives are a critical feature of MPI, enabling efficient communication in applications with sparse communication patterns. This research proposes Cascade, a new algorithm for neighborhood allgather collective that organizes computing nodes along multiple paths based on their distance to the current node. In this approach, messages are forwarded along these paths and propagated until all outgoing neighbors receive them, reducing the communication time. Three performance models are developed to analyze the efficiency of the Cascade algorithm, the default Open MPI algorithm, and the recently proposed Distance-halving neighborhood algorithm in the literature, offering insight into communication cost, scalability, and expected behavior of the algorithms across different system configurations. Experimental results demonstrate that the Cascade algorithm achieves up to 9.54x and 7.05x speedup over Open MPI for random sparse graphs and Moore neighborhoods, respectively. Additionally, the algorithm improves performance by up to$5.25 x$for a sparse matrix-matrix multiplication kernel. The Cascade algorithm outperforms the Distance-halving neighborhood algorithm by up to 2.57 x and 4.81 x speedup for random sparse graphs and Moore neighborhoods, respectively. Moreover, Cascade achieves up to 1.61x performance gain over the Distance-halving neighborhood for the sparse matrix-matrix multiplication kernel. The predictions of our performance models closely match the experimental results. Hamed Sharifian, Amir Hossein Sojoodi, Ahmad Afsahi |
CLUSTER | 3 |
| 2025 | Utilizing Network Hardware Parallelism for MPI Partitioned Collective CommunicationabstractParallel distributed applications running on large-scale high-performance computing systems depend on effective point-to-point and collective communication to meet performance goals. Beginning with version 4.0, the Message Passing Interface (MPI) introduced the partitioned communication API, providing tools for addressing communication bottlenecks raised by hybrid communication models. This API allows individual actors (CPU threads, GPU threads, etc.) to initiate communication on portions of complete buffers, enabling additional communication/computation overlap. Intuitively, the utility of partitioned communication could benefit from network-level support: If there are multiple paths between endpoints, an MPI-aware network could disperse partitions across these paths, avoiding the data serialization entailed by a dependency on a single path. The Cerio Rockport Ethernet Fabric has the ability to expose this capability to communication middleware. In this work we develop this capability to allow for user-level path selection for MPI partitioned communication and explore how this capability impacts point-to-point performance, collective design, and Allreduce efficiency in a Large Language Model task Yiltan Hassan Temuçin, Amirreza Barati Sedeh, Whit Schonbein, Ryan E. Grant, Ahmad Afsahi |
PDP | 5 |
| 2024 | A Topology- and Load-Aware Design for Neighborhood AllgatherabstractNeighborhood collective communications were introduced in MPI 3.0 to enable application developers to define new communication patterns and take advantage of the sparsity in the communication patterns of applications. In this research, we propose a novel topology- and load-aware distance-halving design for neighborhood allgather. In this algorithm, each rank recursively halves the communicator and finds an agent on the opposite half to offload its outgoing neighbors. This approach limits communication with distant ranks, thereby decreasing the latency of neighborhood allgather. Our experimental study demonstrates that our proposed algorithm can outperform the default implementation of Open MPI by up to 30x and 14x speedup for Random Sparse Graph and Moore neighborhood micro-benchmarks, respectively. Furthermore, our design exhibits up to 4.92x performance gain for an SpMM Kernel. Hamed Sharifian, Amir Hossein Sojoodi, Ahmad Afsahi |
CLUSTER | 3 |
| 2023 | A Dynamic Network-Native MPI Partitioned Aggregation Over InfiniBand VerbsabstractModern HPC systems require efficient hybrid programming model to utilize their hardware resources effectively. The Message Passing Interface (MPI) has accommodated next-generation hardware by providing new APIs such as the MPI Partitioned interface. This API provides a user with fine-grain communication without the overhead of traditional MPI point-to-point communication in multi-threaded workloads.To the best of our knowledge, we present the first work on detailed low-level design for an MPI Partitioned implementation. We guide readers through a method to map the MPI Partitioned interface to the InfiniBand Verbs API. Alongside implementation details, we also study the aggregation of user partitions and how we can efficiently send them over the network. We study a brute force approach and using the Partitioned LogGP (PLogGP) model to predict ideal aggregation. We observe that using the PLogGP model provides comparable performance without exhausting computing resources to search the entire solution space. The PLogGP design was further optimized by considering how the partition arrival pattern can be used to dynamically modify our aggregation scheme. We profiled our micro-benchmarks to provide analysis on how and why this additional optimization is beneficial to our results and how we can fine-tune this mechanism. Finally, we evaluated our PLogGP and Timer-based PLogGP designs with a commonly used communication pattern in HPC (communication sweep) to observe the impact when communicating with multiple processes in an application-like scenario at 1024 cores. Yiltan Hassan Temuçin, Scott Levy, Whit Schonbein, Ryan E. Grant, Ahmad Afsahi |
CLUSTER | 5 |
| 2022 | Micro-Benchmarking MPI Partitioned Point-to-Point CommunicationabstractModern High-Performance Computing (HPC) architectures have developed the need for scalable hybrid programming models. The latest Message Passing Interface (MPI) 4.0 standard has introduced a new communication model: MPI Partitioned Point-to-Point communication. This new model allows for the contribution of data from multiple threads with lower overheads than with traditional MPI point-to-point communication. In this paper, we design the first publicly available micro-benchmark suite for MPI Partitioned to measure various metrics that can give insight into the benefits of using this new model and scenarios where MPI point-to-point is better suited. Suggestions are provided to application developers on how to choose partition size for their application based on compute and message size. We evaluate MPI Partitioned communication with both a hot and cold CPU cache, system noise with different probability distributions, point-to-point communication directly, and with commonly used MPI communication patterns such as a halo exchange and Sweep3D. Yiltan Hassan Temuçin, Ryan E. Grant, Ahmad Afsahi |
ICPP | 3 |
| 2022 | Efficient Process Arrival Pattern Aware Collective Communication for Deep LearningabstractMPI collective communication operations are used extensively in parallel applications. As such, researchers have been investigating how to improve their performance and scalability to directly impact application performance. Unfortunately, most of these studies are based on the premise that all processes arrive at the collective call simultaneously. A few studies though have shown that imbalanced Process Arrival Pattern (PAP) is ubiquitous in real environments, significantly affecting the collective performance. Therefore, devising PAP-aware collective algorithms that could improve performance, while challenging, is highly desirable. This paper is along those lines but in the context of Deep Learning (DL) workloads that have become maintstream. Pedram Alizadeh, Amir Hossein Sojoodi, Yiltan Hassan Temuçin, Ahmad Afsahi |
EuroMPI | 4 |
| 2021 | Efficient Multi-Path NVLink/PCIe-Aware UCX based Collective Communication for Deep LearningabstractHigh-performance communication for very large messages on modern multi-GPU nodes has become increasingly important for Deep Learning workloads. These computing nodes are equipped with state-of-the-art interconnects, such as Nvidia's NVLink and PCIe, to facilitate communications between GPUs, and GPUs with the host processors. In this paper, we take on the challenge to design efficient intra-socket GPU-to-GPU communication using multiple NVLink channels at the UCX and MPI levels, and then utilise it to design an intra-node hierarchical NVLink/PCIe-aware GPU based MPI_Allreduce to enhance Horovod + TensorFlow with different models. UCX only utilises a small portion of the available NVLink bandwidth for intra-socket GPU-to-GPU communication. We propose a novel data transfer mechanism that stripes the message across multiple intra-socket communication channels and multiple memory regions using multiple GPU streams to utilise all available NVLink paths. Our approach achieves 1.69x and 1.84x higher bandwidth for UCX and Open MPI + UCX, respectively. We observe similar bandwidth improvements for large messages for MPI point-to-point communication when compared to other MPI implementations as they are also limited by data transfers by a single path. We then propose a 3-stage hierarchical, pipelined MPI_Allreduce design that incorporates the new multi-path NVLink data transfer mechanism for intra-socket communications in the first and third stages of the collective, and PCIe and X-bus channels for inter-socket GPU communication in the second stage with minimal interference. For large messages, our proposed algorithm achieves a high speedup when compared to Spectrum MPI, Open MPI + UCX, Open MPI + HPC-X, MVAPICH2-GDR, and NCCL. We also observe significant speedup for the proposed MPI_Allreduce for Horovod with TensorFlow with a variety of Deep Learning models. Yiltan Hassan Temuçin, Amir Hossein Sojoodi, Pedram Alizadeh, Ahmad Afsahi |
HOTI | 4 |
| 2020 | Communication-aware message matching in MPIabstractSummary The Message Passing Interface (MPI) is the de facto standard for parallel programming in High Performance Computing (HPC). Asynchronous communications in MPI involve message matching semantics that must be satisfied by the conforming libraries. The matching performance is in the critical path of communications in MPI. However, the current message matching approaches suffer from scalability issues and/or do not consider the message queue characteristics of the applications. In this paper, we propose a new message matching mechanism for MPI that can speed up the operation by allocating dedicated queues for certain communications of an application. More specifically, we propose a design that categorizes communications into a set of partners and non‐partners based on the communication frequency in the corresponding queues. We propose a static and a dynamic approach for our message matching design. While the static approach works based on the information from a profiling stage, the dynamic approach utilizes the message queue characteristics at runtime. Our experimental evaluations show that the proposed design can provide up to 28x speedup in queue search time for long list traversals without degrading the performance for short list traversals. We can also gain up to 5x speedup for the FDS application, which is highly affected by the message matching performance. S. Mahdieh Ghazimirsaeed, Seyed Hessam Mirsadeghi, Ahmad Afsahi |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Fuzzy Matching: Hardware Accelerated MPI Communication MiddlewareabstractContemporary parallel scientific codes often rely on message passing for inter-process communication. However, inefficient coding practices or multithreading (e.g., via MPI_THREAD_MULTIPLE) can severely stress the underlying message processing infrastructure, resulting in potentially un-acceptable impacts on application performance. In this article, we propose and evaluate a novel method for addressing this issue: 'Fuzzy Matching'. This approach has two components. First, it exploits the fact most server-class CPUs include vector operations to parallelize message matching. Second, based on a survey of point-to-point communication patterns in representative scientific applications, the method further increases parallelization by allowing matches based on 'partial truth', i.e., by identifying probable rather than exact matches. We evaluate the impact of this approach on memory usage and performance on Knight's Landing and Skylake processors. At scale (262,144 Intel Xeon Phi cores), the method shows up to 1.13 GiB of memory savings per node in the MPI library, and improvement in matching time of 95.9%; smaller-scale runs show run-time improvements of up to 31.0% for full applications, and up to 6.1% for optimized proxy applications. Matthew G. F. Dosanjh, Whit Schonbein, Ryan E. Grant, Patrick G. Bridges, S. Mahdieh Ghazimirsaeed, Ahmad Afsahi |
CCGRID | 6 |
| 2019 | An Efficient Collaborative Communication Mechanism for MPI Neighborhood CollectivesabstractNeighborhood collectives are introduced in MPI3.0 standard to provide users with the opportunitv to define their own communication patterns through the process topologv interface of MPI. In this paper, we propose a collaborative communication mechanism based on common neighborhoods that might exist among groups of k processes. Such common neighborhoods are used to decrease the number of communication stages through message combining. We show how designing our desired communication pattern can be modeled as a maximum weighted matching problem in distributed hvpergraphs, and propose a distributed algorithm to solve it. Moreover, we consider two design alternatives: topologvagnostic and topologv-aware. The former ignores the phvsical topologv o7 the svstem and the mapping o7 processes, whereas the latter takes them into account to further optimize the communication pattern. Our experimental results show that we can gain up to 8x and 5.2x improvement for various process topologies and a SpMM kernel, respectivelv. S. Mahdieh Ghazimirsaeed, Seyed Hessam Mirsadeghi, Ahmad Afsahi |
IPDPS | 3 |
| 2019 | A dynamic, unified design for dedicated message matching engines for collective and point-to-point communications
S. Mahdieh Ghazimirsaeed, Ryan E. Grant, Ahmad Afsahi |
Parallel Comput. | 3 |
| 2018 | The Case for Semi-Permanent Cache Occupancy: Understanding the Impact of Data Locality on Network ProcessingabstractThe performance critical path for MPI implementations relies on fast receive side operation, which in turn requires fast list traversal. The performance of list traversal is dependent on data-locality; whether the data is currently contained in a close-to-core cache due to its temporal locality or if its spacial locality allows for predictable pre-fetching. In this paper, we explore the effects of data locality on the MPI matching problem by examining both forms of locality. First, we explore spacial locality, by combining multiple entries into a single linked list element, we can control and modify this form of locality. Secondly, we explore temporal locality by utilizing a new technique called "hot caching", a process that creates a thread to periodically access certain data, increasing its temporal locality. In this paper, we show that by increasing data locality, we can improve MPI performance on a variety of architectures up to 4x for micro-benchmarks and up to 2x for an application. Matthew G. F. Dosanjh, S. Mahdieh Ghazimirsaeed, Ryan E. Grant, Whit Schonbein, Michael J. Levenhagen, Patrick G. Bridges, Ahmad Afsahi |
ICPP | 7 |
| 2018 | Design considerations for GPU-aware collective communications in MPIabstractSummary GPU accelerators have established themselves in the state‐of‐the‐art clusters by offering high performance and energy efficiency. In such systems, efficient inter‐process GPU communication is of paramount importance to application performance. This paper investigates various algorithms in conjunction with the latest GPU features to improve GPU collective operations. First, we propose a GPU Shared Buffer‐aware (GSB) algorithm and a Binomial Tree Based (BTB) algorithm for GPU collectives on single‐GPU nodes. We then propose a hierarchical framework for clusters with multi‐GPU nodes. By studying various combinations of algorithms, we highlight the importance of choosing the right algorithm within each level. The evaluation of our framework on MPI_Allreduce shows promising performance results for large message sizes. To address the shortcoming for small and medium messages, we present the benefit of using the Hyper‐Q feature and the MPS service in jointly using CUDA IPC and host‐staged copy types to perform multiple inter‐process communications. However, we argue that efficient designs are still required to further harness this potential. Accordingly, we propose a static and a dynamic algorithm for MPI_Allgather and MPI_Allreduce and present their effectiveness on various message sizes. Our profiling results indicate that the achieved performance is mainly rooted in overlapping different copy types. Iman Faraji, Ahmad Afsahi |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Exploiting Common Neighborhoods to Optimize MPI Neighborhood CollectivesabstractNeighborhood collectives were added to the Message Passing Interface (MPI) to better support sparse communication patterns found in many applications. These new collectives encourage more scalable programming styles, and greatly extend the scope of MPI collectives by allowing users to define their own collective communication patterns. In this paper, we describe a new, distributed algorithm for computing improved communication schedules for neighborhood collectives. We show how to discover common process neighborhoods in fully general MPI distributed graph topologies, and how to exploit this information to build message-combining communication schedules for the MPI neighborhood collectives. Our experimental results show considerable performance improvements for application communication topologies of various shapes and sizes. On average, the performance gain is around 50%, but it can also be as much as 71% for topologies with larger numbers of neighbors. Seyed Hessam Mirsadeghi, Jesper Larsson Träff, Pavan Balaji, Ahmad Afsahi |
HiPC | 4 |
| 2017 | Exploiting heterogeneity of communication channels for efficient GPU selection on multi-GPU nodes
Iman Faraji, Seyed Hessam Mirsadeghi, Ahmad Afsahi |
Parallel Comput. | 3 |
| 2016 | MAGC: A Mapping Approach for GPU ClustersabstractGPU accelerators have been increasingly used in modern heterogeneous HPC clusters by offering high performance and energy efficiency. Such heterogeneous GPU clusters consisting of multiple CPU cores and GPU devices have become the platform of choice for many HPC applications. The communication channels among these processing elements expose different latency and bandwidth characteristics. Thus, efficient utilization of communication channels becomes an important factor for achieving higher inter-process communication performance. In this paper, we exploit topology awareness for a better utilization of communication channels in GPU clusters. We first discuss the challenges associated with topology-aware mapping in GPU clusters, and then propose MAGC, a Mapping Approach for GPU Clusters. MAGC seeks to improve the total communication performance by a joint consideration of both CPU-to-CPU and GPU-to-GPU communications of the application, and CPU and GPU physical topologies of the underlying GPU cluster. It provides a unified framework for topology-aware process-to-core mapping and GPU-to-process assignment across a GPU cluster. We study the potential benefits of MAGC with two different mapping algorithms: a) the Scotch graph mapping library, and b) a heuristic designed to explicitly consider maximum congestion. We evaluate our design through extensive experiments at micro-benchmark and application levels on two GPU clusters with different GPU types and topologies. We have developed a micro-benchmark suite to model various communication patterns among CPU cores and among GPU devices. For application results, we use the molecular dynamics simulator, HOOMD-blue. Micro-benchmark results show that we can achieve up to 91.4% improvement in communication time. At the application level, we can achieve up to 8% performance improvement. Seyed Hessam Mirsadeghi, Iman Faraji, Ahmad Afsahi |
SBAC-PAD | 3 |
| 2015 | Scalable connectionless RDMA over unreliable datagrams
Ryan E. Grant, Mohammad J. Rashti, Pavan Balaji, Ahmad Afsahi |
Parallel Comput. | 4 |
| 2014 | Nonblocking Epochs in MPI One-Sided CommunicationabstractThe synchronization model of the MPI one-sided communication paradigm can lead to serialization and latency propagation. For instance, a process can propagate non-RMA communication-related latencies to remote peers waiting in their respective epoch-closing routines in matching epochs. In this work, we discuss six latency issues that were documented for MPI-2.0 and show how they evolved in MPI-3.0. Then, we propose entirely nonblocking RMA synchronizations that allow processes to avoid waiting even in epoch-closing routines. The proposal provides contention avoidance in communication patterns that require back to back RMA epochs. It also fixes the latency propagation issues. Moreover, it allows the MPI progress engine to orchestrate aggressive schedulings to cut down the overall completion time of sets of epochs without introducing memory consistency hazards. Our test results show noticeable performance improvements for a lower-upper matrix decomposition as well as an application pattern that performs massive atomic updates. Judicael A. Zounmevo, Pavan Balaji, William Gropp, Ahmad Afsahi |
SC | 5 |
| 2014 | A fast and resource-conscious MPI message queue mechanism for large-scale jobs
Judicael A. Zounmevo, Ahmad Afsahi |
Future Gener. Comput. Syst. | 2 |
| 2013 | Toward Asynchronous and MPI-Interoperable Active MessagesabstractMany new large-scale applications have emerged recently and become important in areas such as bioinformatics and social networks. These applications are often data-intensive and involve irregular communication patterns and complex operations on remote processes. Active messages have proven effective for parallelizing such nontraditional applications. However, most current active messages frameworks are low-level and system specific, do not efficiently support asynchronous progress, and are not interoperable with two-sided and collective communications. In this paper, we present the design and implementation of an active messages framework inside MPI to provide portability and programmability, and we explore challenges when asynchronously handling active messages and other messages from the network as well as from shared memory. We test our implementation with a set of comprehensive benchmarks. Evaluation results show that our framework has the advantages of overlapping and interoperability, while introducing only a modest overhead. Darius Buntinas, Judicael A. Zounmevo, James Dinan, David Goodell, Pavan Balaji, Rajeev Thakur, Ahmad Afsahi, William Gropp |
CCGRID | 8 |
| 2013 | Mercury: Enabling remote procedure call for high-performance computingabstractRemote procedure call (RPC) is a technique that has been largely adopted by distributed services. This technique, now more and more used in the context of high-performance computing (HPC), allows the execution of routines to be delegated to remote nodes, which can be set aside and dedicated to specific tasks. However, existing RPC frameworks assume a socket-based network interface (usually on top of TCP/IP), which is not appropriate for HPC systems, because this API does not typically map well to the native network transport used on those systems, resulting in lower network performance. In addition, existing RPC frameworks often do not support handling large data arguments, such as those found in read or write calls. We present in this paper an asynchronous RPC interface, called Mercury, specifically designed for use in HPC systems. The interface allows asynchronous transfer of parameters and execution requests and provides direct support of large data arguments. Mercury is generic in order to allow any function call to be shipped. Additionally, the network implementation is abstracted, allowing easy porting to future systems and efficient use of existing native transport mechanisms. Jérome Soumagne, Dries Kimpe, Judicael A. Zounmevo, Mohamad Chaarawi, Quincey Koziol, Ahmad Afsahi, Robert B. Ross |
CLUSTER | 6 |
| 2013 | Using MPI in high-performance computing servicesabstractThe Message Passing Interface (MPI) is one of the most portable high-performance computing (HPC) programming models, with platform-optimized implementations typically delivered with new HPC systems. Therefore, for distributed services requiring portable, high-performance, user-level network access, MPI promises to be an attractive alternative to custom network portability layers, platform-specific methods, or portable but less performant interfaces such as BSD sockets. In this paper, we present our experiences in using MPI as a network transport for a large-scale, distributed storage system. We discuss the features of MPI that facilitate adoption as well as challenges and recommendations. Judicael A. Zounmevo, Dries Kimpe, Robert B. Ross, Ahmad Afsahi |
EuroMPI | 4 |
| 2012 | Designing an Offloaded Nonblocking MPI_Allgather Collective Using CORE-DirectabstractCollective communication operations in the Message Passing Interface (MPI) consume a significant amount of time at scale, degrading the performance of scientific applications. Optimizing collectives is key to application performance and scalability. This paper focuses on hiding the latency of the allgather collective by efficiently offloading it to the networking hardware. We have investigated the use of Mellanox CORE-Direct offloading technology for independent progression of communication within the collective in order to achieve high communication/computation overlap. This study evaluates several design options for the nonblocking allgather collective and discusses implementations of offloaded Standard Exchange, Ring and Bruck algorithms in flat and hierarchical communicators under single-port and k-port modelling. We have applied our findings to improving the performance of the redesigned Radix Sort application kernel. Performance results suggest that our offloaded nonblocking all gather compares favourably to the blocking variant (with improvements of up to 68% for medium messages in a hierarchical collective) while providing high overlap capability. Multiport modelling is shown to be beneficial, especially in a flat communicator. Radix Sort enjoys up to 40% improvement in its runtime. Grigori Inozemtsev, Ahmad Afsahi |
CLUSTER | 2 |
| 2012 | An Efficient MPI Message Queue Mechanism for Large-scale JobsabstractThe Message Passing Interface (MPI) message queues have been shown to grow proportionately to the job size for many applications. With such a behaviour and knowing that message queues are used very frequently, ensuring fast queue operations at large scales is of paramount importance in the current and the upcoming exascale computing eras. Scalability, however, is two-fold. With the growing processor core density per node, and the expected smaller memory density per core at larger scales, a queue mechanism that is blind on memory requirements poses another scalability issue even if it solves the speed of operation problem. In this work we propose a multidimensional queue traversal mechanism whose operation time and memory overhead grow sub-linearly with the job size. We compare our proposal with a linked list-based approach which is not scalable in terms of speed of operation, and with an array-based method which is not scalable in terms of memory consumption. Our proposed multidimensional approach yields queue operation time speedups that translate to up to 4-fold execution time improvement over the linked list design for the applications studied in this work. It also shows a consistent lower memory footprint compared to the array-based design. Judicael A. Zounmevo, Ahmad Afsahi |
ICPADS | 2 |
| 2011 | Investigating Scenario-Conscious Asynchronous Rendezvous over RDMAabstractIn this paper, we propose a light-weight asynchronous message progression mechanism for large message transfers in Message Passing Interface (MPI) Rendezvous protocol that is scenario-conscious and consequently overhead-free in cases where independent message progression naturally happens. Without requiring a dedicated thread, we take advantage of small bursts of CPU to poll for message transfer conditions. The existing application thread is parasitized for the purpose of getting those small bursts of CPU. Our proposed approach is only triggered when the message transfer would otherwise be deferred to the MPI wait call, and it allows for full message progression, achieving 100% overlap. It does not add to the memory footprint of the applications, and is effective in improving the communication performance of most of the applications studied in this paper. Judicael A. Zounmevo, Ahmad Afsahi |
CLUSTER | 2 |
| 2011 | RDMA Capable iWARP over DatagramsabstractiWARP is a state of the art high-speed connection-based RDMA networking technology for Ethernet networks to provide InfiniBand-like zero-copy and one-sided communication capabilities over Ethernet. Despite the benefits offered by iWARP, many data center and web-based applications, such as stock-market trading and media-streaming applications, that rely on data gram-based semantics (mostly through UDP/IP) cannot take advantage of it because the iWARP standard is only defined over reliable, connection-oriented transports. This paper presents an RDMA model that functions over reliable and unreliable data grams. The ability to use data grams significantly expands the application space serviced by iWARP and can bring the scalability advantages of a connectionless transport to iWARP. In our previous work, we had developed an iWARP data gram solution using send/receive semantics showing excellent memory scalability and performance benefits over the current TCP-based iWARP. In this paper, we demonstrate an improved iWARP design that provides true RDMA semantics over data grams. Specifically, because traditional RDMA semantics do not map well to unreliable communication, we propose RDMA Write-Record, the first and the only method capable of supporting RDMA Write over both unreliable and reliable data grams. We demonstrate through a proof-of-concept software implementation that data gram-iWARP is feasible for real-world applications. Our proposed RDMA Write-Record method has been designed with data loss in mind and can provide superior performance under conditions of packet loss. It is shown through micro-benchmarks that by using RDMA capable data gram-iWARP a maximum of 256% increase in large message bandwidth and a maximum of 24.4\% improvement in small message latency can be achieved over traditional iWARP. For application results we focus on streaming applications, showing a 24% improvement in memory usage and up to a 74% improvement in performance, although the proposed approach is also applicable to the HPC domain. Ryan E. Grant, Mohammad J. Rashti, Ahmad Afsahi, Pavan Balaji |
IPDPS | 3 |
| 2011 | Multi-core and Network Aware MPI Topology Functions
Mohammad J. Rashti, Jonathan Green, Pavan Balaji, Ahmad Afsahi, William Gropp |
EuroMPI | 4 |
| 2010 | iWARP redefined: Scalable connectionless communication over high-speed EthernetabstractiWARP represents the leading edge of high performance Ethernet technologies. By utilizing an asynchronous communication model, iWARP brings the advantages of OS bypass and RDMA technology to Ethernet. The current specification of iWARP is only defined over connection-oriented transports such as TCP. The memory requirements of many connections along with TCP's flow and reliability controls lead to scalability and performance issues for large-scale HPC and datacenter applications. In this research, we propose guidelines to extend iWARP over datagrams to provide better scalability and performance. While the proposed extension is designed for use in both HPC and datacenters, the emphasis of this paper is on HPC applications. We present our software implementation of datagram-iWARP over UDP and MPI over datagram-iWARP. Our microbenchmark and MPI application results show performance and memory usage benefits for MPI applications, promoting the use of datagram-iWARP for large-scale HPC applications. Mohammad J. Rashti, Ryan E. Grant, Ahmad Afsahi, Pavan Balaji |
HiPC | 3 |
| 2010 | A study of hardware assisted IP over InfiniBand and its impact on enterprise data center performanceabstractHigh-performance sockets implementations such as the Sockets Direct Protocol (SDP) have traditionally showed major performance advantages compared to the TCP/IP stack over InfiniBand (IPoIB). These stacks bypass the kernel-based TCP/IP and take advantage of network hardware features, providing enhanced performance. SDP has excellent performance but limited utility as only applications relying on the TCP/IP sockets API can use it and other IP stack uses (IPSec, UDP, SCTP) or TCP layer modifications (iSCSI) cannot benefit from it. Recently, newer generations of InfiniBand adapters, such as ConnectX from Mellanox, have provided hardware support for the IP stack itself, such as Large Send Offload and Large Receive Offload. As such high performance socket networks are likely to be deployed or converged with existing Ethernet networking solutions, the performance of such technologies is important to assess. In this paper we take a first look at the performance advantages provided by these offload techniques and compare them to SDP. Our micro-benchmarks and enterprise data-center experiments show that hardware assisted IPoIB can provide competitive performance with SDP and even outperform it in some cases. Ryan E. Grant, Pavan Balaji, Ahmad Afsahi |
ISPASS | 3 |
| 2009 | Evaluation of ConnectX Virtual Protocol Interconnect for Data CentersabstractWith the emergence of new technologies such as Virtual Protocol Interconnect (VPI) for the modern data center, the separation between commodity networking technology and high-performance interconnects is shrinking. With VPI, a single network adapter on a data center server can easily be configured to use one port to interface with Ethernet traffic and another port to interface with high-bandwidth, low-latency InfiniBand technology. In this paper, we evaluate ConnectX VPI using microbenchmarks as well as real traces from a three-tier data center architecture. We find that with VPI each network segment in the data center can use the most optimal configuration (whether InfiniBand or Ethernet) without having to fall back to the lowest common denominator, as is currently the case. Our results show a maximum 26.7% increase in bandwidth, a 54.5% reduction in latency, and a 5% increase in real data center throughput. Ryan E. Grant, Ahmad Afsahi, Pavan Balaji |
ICPADS | 2 |
| 2009 | Improving RDMA-based MPI eager protocol for frequently-used buffersabstractMPI is the main standard for communication in high-performance clusters. MPI implementations use the eager protocol to transfer small messages. To avoid the cost of memory registration and pre-negotiation, the eager protocol involves a data copy to intermediate buffers at both sender and receiver sides. In this paper, however, we propose that when a user buffer is used frequently in an application, it is more efficient to register the sender buffer and avoid the sender-side data copy. The performance results of our proposed eager protocol on MVAPICH2 over InfiniBand indicate that up to 14% improvement can be achieved in a single medium-size message latency, comparable to a maximum 15% theoretical improvement on our platform. We also show that collective communications such as broadcast can benefit from the new protocol by up to 19%. In addition, the communication time in MPI applications with high buffer reuse is improved using this technique. Mohammad J. Rashti, Ahmad Afsahi |
IPDPS | 2 |
| 2009 | Improving energy efficiency of asymmetric chip multithreaded multiprocessors through reduced OS noise schedulingabstractAbstract The performance of the emerging chip multithreaded symmetric multiprocessors (SMPs) is of great importance to the high performance computing community. However, the growing power consumption of such systems is of increasing concern, and techniques that can be used to increase the overall system power efficiency while sustaining the performance are very desirable. Operating system (OS) noise can have a dramatic effect on the system performance. Effectively handling the smaller OS tasks while simultaneously preserving application thread synchronicity leads to gains in the overall system efficiency. Recently, under a fixed power budget, asymmetric multiprocessors (AMP) have been proposed to improve the performance of multithreaded applications. An AMP in this context is a multiprocessor system in which its processors are not operating at the same frequency. This paper proposes two simple scheduling methods that reduce the impact of OS noise, while simultaneously taking advantage of an opportunity to increase the overall machine energy efficiency on AMP servers. Prototyping AMPs on a commercial 2‐way dual‐core Hyper‐Threaded (HT) Intel Xeon SMP server, using real power measurements across six SPEC OpenMP applications, indicates that the first proposed scheduler performs better on average for HT‐enabled systems, whereas the second scheduler is superior on average for HT‐disabled systems. Copyright © 2009 John Wiley & Sons, Ltd. Ryan E. Grant, Ahmad Afsahi |
Concurr. Comput. Pract. Exp. | 2 |
| 2007 | Improving system efficiency through scheduling and power managementabstractThe performance of the emerging commercial chip multithreaded multiprocessors is of great importance to the high performance computing community. However, the growing power consumption of such systems is of increasing concern, and techniques that could be effectively used to increase overall system power efficiency while sustaining performance are very desirable. Ryan E. Grant, Ahmad Afsahi |
CLUSTER | 2 |
| 2007 | A feasibility analysis of power-awareness and energy minimization in modern interconnects for high-performance computingabstractHigh-performance computing (HPC) systems consume a significant amount of power, resulting in high operational costs, reduced reliability, and wasting of natural resources. Therefore, power consumption has become an increasingly important design constraint in high-performance clusters. In this regard, research on power-aware HPC has emerged. While most research has focused at understanding and utilizing applicationspsila behavior to scale down the CPU for energy savings, this paper demonstrates the positive impact of modern interconnects in delivering energy-efficiency in high-performance clusters. In this work, we first present the power-performance profiles of the Myrinet-2000 and Quadrics QsNetIIat the user-level and MPI-level in comparison to a traditional, non-offloaded Gigabit Ethernet. Such information enables us to devise a power-aware MPI runtime library that automatically and transparently performs message segmentation and re-assembly in order to increase energy savings. Secondly, by designing and evaluating a number of all-gather collectives, we argue that it is possible to increase the energy-efficiency of a cluster by optimizing its messaging layers. Reza Zamani, Ahmad Afsahi, V. Carl Hamacher |
CLUSTER | 2 |
| 2007 | RDMA-based and SMP-aware Multi-port All-Gather on Multi-rail QsNet^II SMP ClustersabstractClusters of symmetric multiprocessors (SMP) are more commonplace than ever in achieving high- performance. Scientific applications running on clusters employ collective communications extensively. Using shared memory communication among co- located processes on SMP nodes as well as remote direct memory access (RDMA) operations for inter- node communication and trying to overlap them is a proven technique in boosting the performance of collective operations. The effect is much more pronounced when efficient multi-port collectives on multi-rail networks are devised and implemented. In this work, we design and implement multi-port RDMA-based and SMP-aware all-gather algorithms with message striping over multi-rail QsNeIIdirectly at the Elan level. We compare our algorithms against RDMA-only traditional algorithms and the native elan_gather(). Our performance results indicate that the proposed SMP-aware Brack all-gather gains an improvement of up to 1.96 for 4KB messages over the native elanjgather(). Meanwhile, the direct algorithm achieves up to 1.49 improvement for 32 KB messages. Ahmad Afsahi |
ICPP | 2 |
| 2007 | A Comprehensive Analysis of OpenMP Applications on Dual-Core Intel Xeon SMPsabstractHybrid chip multithreaded SMPs present new challenges as well as new opportunities to maximize performance. Our intention is to discover the optimal operating configuration of such systems for scientific applications and to identify the shared resources that might become a bottleneck to performance under the different hardware configurations. This knowledge will be useful to the research community in developing software techniques to improve the performance of shared memory programs on modern multi-core multiprocessors. In this paper, we study a two-way dual-core Hyper-Threaded (HT) Intel Xeon SMP server under single program and multi-program multithreaded workloads using the NAS OpenMP benchmark suite. Our performance results indicate that in the single-program case, the CMP-based SMP and CMT-based SMP configurations have the highest average speedup across all of the applications. The most efficient architecture is a single HT-enabled dual-core processor that is almost comparable to the performance of a 2-way dual-core HT-disabled system. Ryan E. Grant, Ahmad Afsahi |
IPDPS | 2 |
| 2007 | 10-Gigabit iWARP Ethernet: Comparative Performance Analysis with InfiniBand and Myrinet-10GabstractiWARP is a set of standards enabling remote direct memory access (RDMA) over Ethernet. iWARP supporting RDMA and OS bypass, coupled with TCP/IP offload engines, can fully eliminate the host CPU involvement in an Ethernet environment. With the iWARP standard and the introduction of 10-Gigabit Ethernet, there is now an alternative path to the proprietary interconnects for high-performance computing, while maintaining compatibility with existing Ethernet infrastructure and protocols. Recently, NetEffect Inc. has introduced an iWARP-enabled 10-Gigabit Ethernet channel adapter. In this paper we assess the potential of such an interconnect for high-performance computing by comparing its performance with two leading cluster interconnects, infiniband and myrinet-10G. The results show that the NetEffect iWARP implementation achieves an unprecedented latency for Ethernet, and saturates 87% of the available bandwidth. It also scales better with multiple connections. At the MPI level, iWARP performs better than infiniband in queue usage and buffer re-use. Mohammad J. Rashti, Ahmad Afsahi |
IPDPS | 2 |
| 2006 | Power-performance efficiency of asymmetric multiprocessors for multi-threaded scientific applicationsabstractRecently, under a fixed power budget, asymmetric multiprocessors (AMP) have been proposed to improve the performance of multi-threaded applications compared to symmetric multiprocessors. An AMP is a multiprocessor system in which its processors are not operating at the same frequency. Power consumption has become an important design constraint in servers and high-performance server clusters. This paper explores the power-performance efficiency of hyper-threaded (HT) AMP servers, and proposes a new scheduling algorithm that can be used to reduce the overall power consumption of a server while maintaining a high level of performance. Prototyping AMPs on a commercial 4-way SMP server, we show that on average 15.6% energy savings and 6.1% slowdown for the HT-disabled case, and 7.1% energy savings and 4.8% slowdown for the HT-enabled case can be achieved across NAS and SPEC OpenMP applications. Ryan E. Grant, Ahmad Afsahi |
IPDPS | 2 |
| 2006 | Efficient RDMA-based multi-port collectives on multi-rail QsNetII clustersabstractMany scientific applications use MPI collective communications intensively. Therefore, efficient and scalable implementation of collective operations is critical to the performance of such applications running on clusters. Quadrics QsNetIIis a high-performance interconnect for clusters that implements some collectives at the Elan level. These collectives are directly used by their corresponding MPI collectives. Quadrics software supports point-to-point striping over multi-rail QsNetIInetworks. However, multi-rail collectives have not been supported. In this work, we propose a number of RDMA-based multi-port collectives over multi-rail QsNetIIclusters directly at the Elan level. Our performance results indicate that the proposed multi-port gather gains an improvement of up to 6.35 for 1MB message over the native elan_gather. The proposed multi-port all-to-all performs better than the native elan_alltoall by a factor of 2.19 for 16KB message. Moreover, we have also proposed two algorithms for the scatter operation Ahmad Afsahi |
IPDPS | 2 |
| 2004 | Myrinet Networks: A Performance StudyabstractAs network computing become commonplace, the interconnection networks and the communication system software become critical in achieving high performance. Thus, it is essential to systematically assess the features and performance of the new networks. Recently, Myricom has introduced a two-port "E-card" Myrinet/PCl-X interface. In this paper, we present the basic performance of its GM2.I messaging layer, as well as a set of microbenchmarks designed to assess the quality of MPI implementation on top of GM. These microbenchmarks measure the latency, bandwidth, intra-node performance, computation/communication overlap, parameters of the LogP model, buffer reuse impact, different traffic patterns, and collective communications. We have discovered that the MPI basic performance is close to those offered at the GM. We find that the host overhead is very small in our system. The Myrinet network is shown to be sensitive to the buffer reuse patterns. However, it provides opportunities for overlapping computation with communication. The Myrinet network is able to deliver up to 2000MB/s bandwidth for the permutation patterns. Ahmad Afsahi, Reza Zamani |
NCA | 2 |
| 2003 | Performance characteristics of openMP constructs, and application benchmarks on a large symmetric multiprocessorabstractWith the increasing popularity of small to large-scale symmetric multiprocessor (SMP) systems, there has been a dire need to have sophisticated, and flexible development and runtime environments for efficient and rapid development of parallel applications. To this end, OpenMP has emerged as the standard for parallel programming on shared-memory systems. It is very important to evaluate the performance of OpenMP constructs, kernels, and application benchmarks on large-scale SMP systems. We present the performance of the basic OpenMP constructs, class B of NAS OpenMP 3.0 benchmarks, and the SPEC OMPL2001 application benchmarks (large data set) on a contemporary 72-node Sun Fire 15K SMP node. We report the basic timings, scalability, and runtime profiles of different parallel regions within each benchmark in the NAS OpenMP 3.0, and the SPEC OMPL-2001 suites. We elaborate on the performance differences between the medium and large classes of the SPEC OMP2001 suites on our system, as well as a comparison among a number of large-scale symmetric multiprocessors for the SPEC OMPL2001. Nathan R. Fredrickson, Ahmad Afsahi |
ICS | 2 |
| 2002 | Efficient communication using message prediction for clusters of multiprocessorsabstractAbstract With the increasing uniprocessor and symmetric multiprocessor computational power available today, interprocessor communication has become an important factor that limits the performance of clusters of workstations/multiprocessors. Many factors including communication hardware overhead, communication software overhead, and the user environment overhead (multithreading, multiuser) affect the performance of the communication subsystems in such systems. A significant portion of the software communication overhead belongs to a number of message copying operations. Ideally, it is desirable to have a true zero‐copy protocol where the message is moved directly from the send buffer in its user space to the receive buffer in the destination without any intermediate buffering. However, due to the fact that message‐passing applications at the send side do not know the final receive buffer addresses, early arrival messages have to be buffered at a temporary area. In this paper, we show that there is a message reception communication locality in message‐passing applications. We have utilized this communication locality and devised different message predictors at the receiver sides of communications. In essence, these message predictors can be efficiently used to drain the network and cache the incoming messages even if the corresponding receive calls have not yet been posted. The performance of these predictors, in terms of hit ratio, on some parallel applications are quite promising and suggest that prediction has the potential to eliminate most of the remaining message copies. We also show that the proposed predictors do not have sensitivity to the starting message reception call, and that they perform better than (or at least equal to) our previously proposed predictors. Copyright © 2002 John Wiley & Sons, Ltd. Ahmad Afsahi, Nikitas J. Dimopoulos |
Concurr. Comput. Pract. Exp. | 1 |
| 1997 | Collective Communications on a Reconfigurable Optical Interconnect
Ahmad Afsahi, Nikitas J. Dimopoulos |
OPODIS | 1 |