EDBT 2026 Demo / reviewers in the wild / expert
Anthony Skjellum
dblp:40/6558
· DBLP profile ↗
65ranked-venue papers
5as first author
19since 2021 · last 2025
0000-0001-5252-6600ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 5 first-author · 8 since 2021Security and privacy · 9 · 5 since 2021Computer networks · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Authentic Learning Exercise for Kubernetes Misconfigurations: An Experience Report of Student PerceptionsabstractKubernetes has become a popular tool for automated container orchestration. Despite reported benefits, practitioners report that the secure configuration of Kubernetes is one of the primary challenges among practitioners. Moreover, there is a significant skill shortage of Kubernetes security experts. Understanding misconfigurations in Kubernetes can help practitioners prevent security incidents. We systematically investigate whether authentic learning can help students learn about misconfigurations in Kubernetes. We conduct an authentic learning exercise and collected responses from 295 students. Based on responses from the students, we find (i) students who have little to no experience in cybersecurity, software quality assurance, or static analysis perceived the authentic learning exercise as useful to learn misconfigurations in Kubernetes, and (ii) students perceptions of authentic learning exercise activities vary based on and educational background. We conclude our paper with recommendations for instructors and researchers. Md. Shazibul Islam Shamim, Fan Wu 0013, Hossain Shahriar, Anthony Skjellum, Akond Ashfaque Ur Rahman |
CSEE&T | 4 |
| 2025 | Concepts for Designing Modern C++ Interfaces for MPIabstractAbstract Since the C++ bindings were deleted in 2008, the Message Passing Interface (MPI) community has recently revived efforts in building high-level modern C++ interfaces. Such interfaces are either built to serve specific scientific application needs (with limited coverage to the underlying MPI functionality), or as an exercise in general-purpose programming model building, with the hope that bespoke interfaces can be broadly adopted to construct a variety of distributed-memory scientific applications. However, with the advent of modern C++-based heterogeneous programming models, GPUs and widespread Machine Learning (ML) usage in contemporary scientific computing, the role of prospective community-standardized high-level C++ interfaces to MPI is evolving. The success of such an interface clearly will depend on providing robust abstractions and features adhering to the generic programming principles that underpin the C++ programming language, without compromising on either performance or portability, the core principles upon which MPI was founded. However, there is a tension between idiomatic C++ handling of types and lifetimes and MPI’s loose interpretation of object lifetimes/ownership and insistence on maintaining global states. Instead of proposing “yet another” high-level C++ interface to MPI, overlooking or providing partial solutions to work around the key issues concerning the dissonance between MPI semantics and idiomatic C++, this paper focuses on the three fundamental aspects of a high-level interface: type system, object lifetimes, and communication buffers, while also identifying inconsistencies in the MPI specification. Presumptive solutions can be unrefined, and we hope the broader MPI and C++ communities will engage with us in productive exchange of ideas and concerns. C. Nicole Avans, Alfredo A. Correa, Matthias Schimek, Joseph Schuchart, Anthony Skjellum, Evan Drake Suggs, Tim Niklas Uhl |
EuroMPI | 6 |
| 2024 | Quantifying and Modeling Irregular MPI CommunicationabstractMany modern scientific applications have communication patterns where both the number of communication partners and amount of data transmitted between process pairs vary significantly and change over time. This work describes an approach to measure and model the behavior of these irregular, dynamic MPI communication patterns on modern high-performance computing systems. Specifically, this approach quantifies communication behavior using a small number of stochastic random variables that capture key features of irregular communication patterns, and estimates the distributions of these variables either parameterically or empirically. This work then demonstrates that the collected parameters and their distributions can be used to measure and model the communication performance of several MPI applications. It also presents a synthetic benchmark that uses these distributions to recreate statistically similar communication patterns. This approach provides a lightweight method to reproduce communication patterns of a variety of applications with minimal overhead while gaining additional insights into the performance and characteristics of various irregular communication patterns. Carson Woods, Derek Schafer, Patrick G. Bridges, Anthony Skjellum |
CCGrid | 4 |
| 2024 | SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC ApplicationsabstractCommunication switches have sometimes been augmented to process collectives, e.g., in the IBM BlueGene and Mellanox SHArP switches. In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third are higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). Pouya Haghi, Cheng Tan 0002, Anqi Guo, Chunshu Wu, Dongfang Liu, Ang Li 0006, Anthony Skjellum, Tong Geng, Martin C. Herbordt |
ICS | 7 |
| 2024 | Design and Development of XiveNet: A Hybrid CAN Research TestbedabstractWe have developed an affordable distributed Internet of Things (IoT) testbed, named XiveNet, to conduct in-vehicle security research. This testbed merges the adaptability of simulators with the real-time ECU characteristics of actual vehicles. The testbed is made up of ECU chips found in vehicles, Raspberry Pis, and is combined with a bus master simulator. Our experiments with CAN (controller area network) traffic from actual vehicles (Oak Ridge National Laboratories Road Data Set) demonstrate that our testbed closely replicates the attributes of a real vehicle. We have further authenticated our testbed by deploying SecCAN, a secure CAN algorithm, and evaluating its security by injecting invalid frames. Furthermore, we examined ORNL’s timing-based intrusion detection on our testbed and successfully produced alerts. Additionally, we incorporated Named Data Networking (NDN) capable nodes, providing researchers with an additional resource to develop future in-vehicle security solutions. Finally, we have proposed a bitrate hopping technique focused on preventing the denial of service attack and conducted a preliminary investigation using the testbed. Our evaluation and validation indicate that the testbed provides the real-world vehicle environment with the flexibility of a simulation environment that supports a wide range of hardware and software configurations. William Luke Lambert, Sheikh K. Ghafoor, Haley Burnell, Brennan Huber, Farah I. Kandah, Anthony Skjellum |
ISPDC | 6 |
| 2024 | Understanding GPU Triggering APIs for MPI+X Communication
Patrick G. Bridges, Anthony Skjellum, Evan Drake Suggs, Derek Schafer, Purushotham V. Bangalore |
EuroMPI | 2 |
| 2024 | MitM attacks on intellectual property and integrity of additive manufacturing systems: A security analysis
Hamza Alkofahi, Heba Alawneh, Anthony Skjellum |
Comput. Secur. | 3 |
| 2023 | FLASH: FPGA-Accelerated Smart Switches with GCN Case StudyabstractSome communication switches, e.g., the Mellanox SHArP and those in the IBM BlueGene clusters, are augmented to process packets at the application level with fixed-function collectives. This approach, however, lacks flexibility, which limits their applicability in diverse and dynamic workloads. Recently, a new type of programmable packet processor, which uses high-level languages, e.g., P4, has emerged as a possible candidate. P4-based switches, however, fall short in certain applications, including machine learning, where capabilities not currently supported by P4 are needed. These include more complex calculation, such as sparse computation and fused multiply-accumulate, data-intensive floating point operations, data reuse, and significant memory. The problem addressed here is that such a switch augmentation needs to support: a large amount of state, significant flexible compute capability, and ease of programming, all while maintaining full functionality, including ensuring high throughput, and demonstrating utility. Pouya Haghi, William Krska, Cheng Tan 0002, Tong Geng, Po-Hao Chen 0002, Connor Greenwood, Anqi Guo, Thomas M. Hines, Chunshu Wu, Ang Li 0006, Anthony Skjellum, Martin C. Herbordt |
ICS | 11 |
| 2023 | View-aware Message Passing Through the Integration of Kokkos and ExaMPIabstractKokkos provides in-memory advanced data structures, concurrency, and algorithms to support performance portable C++ parallel programming across CPUs and GPUs. The Message Passing Interface (MPI) provides the most widely used message passing model for inter-node communication. Many programmers use both Kokkos and MPI together. In this paper, Kokkos is integrated within an MPI implementation for ease of use in applications that use both Kokkos and MPI, without sacrificing performance. For instance, this model allows passing first-class Kokkos objects directly to extended C++-based MPI APIs. Evan Drake Suggs, Stephen Olivier, Jan Ciesko, Anthony Skjellum |
EuroMPI | 4 |
| 2023 | Design of a portable implementation of partitioned point-to-point communication primitivesabstractAbstract The Message Passing Interface (MPI) has been the dominant message passing solution for scientific computing for decades. MPI point‐to‐point communications are highly efficient mechanisms for process‐to‐process communication. However, MPI performance when processes utilize multiple threads is slowed by concurrency protections in the MPI library. MPI's current thread level interface imposes these overheads throughout the library when thread safety is needed. While much work has been done to reduce multithreading overheads in MPI, a solution is needed that reduces the number of messages exchanged in a threaded environment. Partitioned communication is included in the MPI 4.0 standard as an alternative that addresses the challenges of multithreaded communication in MPI today. Partitioned communication reduces overall message volume by creating a buffer‐sharing mechanism between threads such that they can indicate when portions of a communication buffer are available to be sent. Separation of the control and data planes in MPI is enabled by allowing persistent initialization and single occurrence message buffer matching from the indication that the data is ready to be sent. This enables the usage of underlying hardware primitives like triggered operations, where commands (destination, size, etc.) can be set up prior to data buffer readiness and readiness triggered with a simple doorbell/counter later. This approach is useful for future development of MPI operations in environments where traditional networking commands can have performance challenges, like accelerators (GPUs, FPGAs). In this paper, we detail the design and implementation of a layered library (built on top of MPI‐3.1) and an integrated Open MPI solution that supports the new, MPI‐4.0 partitioned communication feature set. The library will enable applications to use currently released MPI implementations and older legacy libraries to provide partitioned communication support while also enabling further exploration of this new communication model in new applications and use cases. We will compare the designs of the library and native Open MPI support, provide performance results and comparisons between the two approaches, and lessons learned from the implementation of partitioned communication in both library and native forms. We find that the native implementation and library have similar performance with a percentage difference under 0.94% in microbenchmarks and performance within 5% for a partitioned communication enabled proxy application. W. Pepper Marts, Andrew Worley, Prema Soundarajan, Derek Schafer, Matthew G. F. Dosanjh, Ryan E. Grant, Purushotham V. Bangalore, Anthony Skjellum, Sheikh K. Ghafoor |
Concurr. Comput. Pract. Exp. | 8 |
| 2023 | BEAST: Behavior as a Service for Trust management in IoT devices
Brennan Huber, Farah I. Kandah, Anthony Skjellum |
Future Gener. Comput. Syst. | 3 |
| 2022 | Reconfigurable switches for high performance and flexible MPI collectivesabstractAbstract There has been much effort in offloading MPI collective operations into hardware. But while NIC‐based collective acceleration is well‐studied, offloading their processing into the switching fabric, despite numerous advantages, has been much more limited. A major problem with fixed logic implementations is that either only a fraction of the possible collective communication is accelerated or that logic is wasted in the applications that do not need a particular capability. Using reconfigurable logic has numerous advantages: exactly the required operations can be implemented; the level of desired performance can be specified; and new, possibly complex, operations can be defined and implemented. We have designed an in‐switch collective accelerator,MPI‐FPGA, and demonstrated its use with seven MPI collectives and over a set of benchmarks and proxy applications (MiniApps). The accelerator uses a novel two‐level switch design containing fully pipelined vectorized aggregation logic units. Essential to this work is providing support for sub‐communicator collectives that enables communicators of arbitrary shape, and that is scalable to large systems. A streaming interface improves the performance for long messages. While this reconfigurable design is generally applicable, we prototype it with an FPGA‐centric cluster. A sampleMPI‐FPGAdesign in a direct network achieves considerable speedups over conventional clusters in the most likely scenarios. We also present results for indirect networks with reconfigurable high‐radix switches and show that this approach is competitive withSHArPtechnology for the subset of operations thatSHArPsupports.MPI‐FPGAis fully integrated into MPICH and is transparent to MPI applications. Pouya Haghi, Anqi Guo, Qingqing Xiong, Chen Yang 0010, Tong Geng, Justin T. Broaddus, Ryan J. Marshall, Derek Schafer, Anthony Skjellum, Martin C. Herbordt |
Concurr. Comput. Pract. Exp. | 9 |
| 2021 | An Overview of Cryptographic AccumulatorsabstractThis paper is a primer on cryptographic accumulators and how to apply them practically. A cryptographic accumulator is a space- and time-efficient data structure used for set-membership tests. Since it is possible to represent any computational problem where the answer is yes or no as a set-membership problem, cryptographic accumulators are invaluable data structures in computer science and engineering. But, to the best of our knowledge, there is neither a concise survey comparing and contrasting various types of accumulators nor a guide for how to apply the most appropriate one for a given application. Therefore, we address that gap by describing cryptographic accumulators while presenting their fundamental and so-called optional properties. We discuss the effects of each property on the given accumulator's performance in terms of space and time complexity, as well as communication overhead. Ilker Özçelik, Sai Medury, Justin T. Broaddus, Anthony Skjellum |
ICISSP | 4 |
| 2021 | Encryption is Futile: Reconstructing 3D-Printed Models Using the Power Side-ChannelabstractOutsourced Additive Manufacturing (AM) exposes sensitive design data to external malicious actors. Even with end-to-end encryption between the design owner and 3D-printer, side-channel attacks can be used to bypass cyber-security measures and obtain the underlying design. In this paper, we develop a method based on the power side-channel that enables accurate design reconstruction in the face of full encryption measures without any prior knowledge of the design. Our evaluation on a Fused Deposition Modeling (FDM) 3D Printer has shown 99 % accuracy in reconstruction, a significant improvement on the state of the art. This approach demonstrates the futility of pure cyber-security measures applied to Additive Manufacturing. Jacob Gatlin, Sofia Belikovetsky, Yuval Elovici, Anthony Skjellum, Joshua Lubell, Paul Witherell, Mark Yampolskiy |
RAID | 4 |
| 2021 | What Did You Add to My Additive Manufacturing Data?: Steganographic Attacks on 3D Printing FilesabstractAdditive Manufacturing (AM) adoption is increasing in home and industrial settings, but information security for this technology is still immature. Thus far, three security threat categories have been identified: technical data theft, sabotage, and illegal part manufacturing. In this paper, we expand to a new threat category: misuse of digital design files as a subliminal communication channel. We identify and explore attacks by which arbitrary information can be embedded steganographically in the most common digital design file format, the STL, without distorting the printed object. Because the technique will not change the manufactured object’s geometry, it is likely to remain unnoticed and can be exploited for data transfer. Further, even with knowledge of our methods, defenders cannot distinguish between actual data transfer and random manipulation of the files. This is the first info-hiding attack on this system, conducted despite the fact that random changes may spoil the physical artifact and result in detection. Mark Yampolskiy, Lynne Graves, Jacob Gatlin, Anthony Skjellum, Moti Yung |
RAID | 4 |
| 2021 | Availability analysis of a permissioned blockchain with a lightweight consensus protocol
Amani Altarawneh, Richard R. Brooks, Oluwakemi Hambolu, Lu Yu 0001, Anthony Skjellum |
Comput. Secur. | 6 |
| 2021 | Understanding the use of message passing interface in exascale proxy applicationsabstractSummary The Exascale Computing Project (ECP) focuses on the development of future exascale‐capable applications. Most ECP applications use the message passing interface (MPI) as their parallel programming model with mini‐apps serving as proxies. This paper explores the explicit usage of MPI in such ECP proxy applications. We empirically analyze 14 proxy applications from the ECP Proxy Apps Suite. We use the MPI profiling interface (PMPI) to collect MPI usage patterns in ECP proxy apps. Our analysis shows that a small subset of features from MPI is commonly used in the proxies of exascale‐capable applications, even when they reference third‐party libraries. This study is intended to provide a better understanding of the use of MPI in current exascale applications. The findings can help focus software investments made for exascale systems in the MPI middleware including optimization, fault‐tolerance, tuning, and hardware‐offload. Nawrin Sultana, Martin Ruefenacht, Anthony Skjellum, Purushotham V. Bangalore, Ignacio Laguna, Kathryn Mohror |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Radio Identity Verification-Based IoT Security Using RF-DNA Fingerprints and SVMabstractIt is estimated that the number of Internet-of-Things (IoT) devices will reach 75 billion in the next five years. Most of those currently and soon-to-be deployed devices lack sufficient security to protect themselves and their networks from attacks by malicious IoT devices masquerading as authorized devices in order to circumvent digital authentication approaches. This work presents a physical (PHY) layer IoT authentication approach capable of addressing this critical security need through the use of feature-reduced, radio frequency-distinct native attributes (RF-DNA) fingerprints and support vector machines (SVM). This work successfully demonstrates: 1) authorized identity (ID) verification across three trials of six randomly chosen radios at signal-to-noise ratios greater than or equal to 6 dB and 2) rejection of all rogue radio ID spoofing attacks at signal-to-noise ratios greater than or equal to 3 dB using RF-DNA fingerprints whose features are selected using the Relief-F algorithm. Donald R. Reising, Joseph Cancelleri, T. Daniel Loveless, Farah I. Kandah, Anthony Skjellum |
IEEE Internet Things J. | 5 |
| 2021 | Implementation and evaluation of MPI 4.0 partitioned communication libraries
Matthew G. F. Dosanjh, Andrew Worley, Derek Schafer, Prema Soundararajan, Sheikh K. Ghafoor, Anthony Skjellum, Purushotham V. Bangalore, Ryan E. Grant |
Parallel Comput. | 6 |
| 2020 | Accelerating MPI Collectives with FPGAs in the Network and Novel Communicator SupportabstractMPI collective operations can often be performance killers in HPC applications; we seek to solve this bottleneck by offloading them to reconfigurable hardware within the switch itself, rather than, e.g., the NIC. We have designed a hardware accelerator MPI-FPGA to implement six MPI collectives in the network. Preliminary results show that MPI-FPGA achieves $10 \times $ speedup in the most likely scenarios over conventional clusters. We introduce a novel mechanism that enables the hardware to support a large number of communicators of arbitrary shape, and that is scalable to very large systems. MPI-FPGA is fully integrated into MPICH and so transparent to MPI applications. Qingqing Xiong, Chen Yang 0010, Pouya Haghi, Anthony Skjellum, Martin C. Herbordt |
FCCM | 4 |
| 2020 | The Security Ingredients for Correct and Byzantine Fault-tolerant Blockchain Consensus AlgorithmsabstractThe blockchain technology revolution and the use of blockchains in various applications have resulted in many companies and programmers developing and customizing specific fit-for-purpose consensus algorithms. Security and performance are determined by the chosen consensus algorithm; hence, the reliability and security of these algorithms must be assured and tested, which requires an understanding of all the security assumptions that make such algorithms correct and byzantine fault-tolerant.This paper studies the "security ingredients" that enable a given consensus algorithm to achieve safety, liveness, and byzantine fault tolerance (BFT) in both permissioned and permissionless blockchain systems. The key contributions of this paper are the organization of these requirements and a new taxonomy that describes the requirements for security. The CAP Theorem is utilized to explain important tradeoffs between consistency and availability in consensus algorithm design, which are crucial depending on the specific application of a given algorithm. This topic has also been explored previously by De Angelis. However, this paper expands that prior explanation and dilemma of consistency vs. availability and then combines this with Buterin's Trilemma to complete the overall exposition of tradeoffs. Amani Altarawneh, Anthony Skjellum |
ISNCC | 2 |
| 2020 | Why is MPI (perceived to be) so complex?: Part 1 - Does strong progress simplify MPI?abstractStrong progress is optional in MPI. MPI allows implementations where progress (for example, updating the message-transport state machines or interaction with network devices) is only made during certain MPI procedure calls. Generally speaking, strong progress implies the ability to achieve progress (to transport data through the network from senders to receivers and exchange protocol messages) without explicit calls from user processes to MPI procedures. For instance, data given to a send procedure that matches a pre-posted receive on the receiving process is moved from source to destination in due course regardless of how often (including zero times) the sender or receiver processes call MPI in the meantime. Further, nonblocking operations and persistent collective operations work ‘in the background’ of user processes once all processes in the communicator’s group have performed the starting step for the operation. Overall, strong progress is meant to enhance the potential for overlap of communication and computation and improve predictability of procedure execution times by eliminating progress effort from user threads. This paper posits that strong progress is desirable as an MPI implementation property and examines whether strong progress: This paper explores such possibilities and sets forth principles that underpin MPI and interactions with normal and fault modes of operation. The key contribution of this paper is the conclusion that, whether measured by absolute performance, by performance portability, or by interface simplicity, strong progress in MPI is no worse than weak progress and, in most scenarios, has more potential to fulfil the aforementioned desirable attributes. Daniel J. Holmes, Anthony Skjellum, Derek Schafer |
EuroMPI | 2 |
| 2020 | Foreword to the Special Issue of the Workshop on Exascale MPI (ExaMPI 2017)abstractThe aim of the Workshop on Exascale MPI (ExaMPI 2017), held in conjunction with SC17: The International Conference for High Performance Computing, Networking, Storage and Analysis, was to bring together researchers and developers to present and discuss innovative algorithms and concepts in the Message Passing programming model and to create a forum for open and potentially controversial discussions on the future of MPI in the Exascale era. This special issue includes selected papers from this workshop that include innovative algorithms for collective operations, extensions to MPI, including datacentric models, scheduling/routing to avoid network congestion, fault-tolerant communication, interoperability of MPI and PGAS models, and use of MPI in large-scale simulations. The first paper, titled “A survey of MPI usage in the US Exascale Computing Project,” provides an analysis of the survey that was conducted to understand how MPI is currently used and intended to be used by different applications that are part of the Exascale Computing Project (ECP).1 The results of analysis provide specific recommendations for MPI implementors, tool developers, and the MPI Forum. The next three papers focus on the issue of message matching in MPI and provide different options to address message matching at exascale. The paper titled “Tail Queues: A Multi-threaded Matching Architecture” introduces a novel parallel matching architecture and prototype implementation based on MPICH to improve the performance of message matching.2 The paper titled “Communication-Aware Message Matching in MPI” uses a novel message queue architecture that allocates dedicated message queues based on the frequency of communication between various processes to reduce the queue search time and also reduces memory consumption.3 These performance improvements result in a speedup of 5 times on the FDS application. The next paper, titled “Hardware MPI Message Matching: Insights into MPI Matching Behavior to Inform Design,” explores what hardware features are needed to support efficient message matching through the evaluation of message matching characteristics of major MPI implementations.4 The next two papers consider support for fault tolerance. The first paper, titled “EReinit: Scalable and Efficient Fault-Tolerance for Bulk-Synchronous MPI Applications,” describes a global-restart model to improve the recovery time of applications when dealing with faults.5 The second paper, titled “The Unexpected Virtue of Almost: Exploiting MPI Collective Operations to Approximately Coordinate Checkpoints,” describes an uncoordinated checkpointing mechanism that makes use of collective operations already used in an application to force checkpoints.6 The paper titled “Optimizing Point-to-Point Communication between Adaptive MPI Endpoints in Shared Memory” describes an approach to optimize point-to-point communication in a shared memory environment and hence improve MPI multithreading support.7 The paper titled “On the Memory Attribution Problem: A Solution and Case Study Using MPI” describes a solution to capture and analyze memory usage by an MPI application and the MPI library.8 Lastly, the paper titled “Twister2: Design of a Big Data Toolkit” describes an architecture to support different types of data-intensive applications in a unified framework.9 Overall, these nine papers contribute to the knowledge base and advancement of the Message Passing Interface in diverse and useful ways. While illustrating the staying power of MPI after a quarter century, they also point to areas of opportunity for enhancement at Exascale as well as to new potential application areas. Anthony Skjellum, Purushotham V. Bangalore, Ryan E. Grant |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | MPI Sessions: Evaluation of an Implementation in Open MPIabstractThe recently proposed MPI Sessions extensions to the MPI standard present a new paradigm for applications to use with MPI. MPI Sessions has the potential to address several limitations of MPI's current specification: MPI cannot be initialized within an MPI process from different application components without a priori knowledge or coordination; MPI cannot be initialized more than once; and, MPI cannot be reinitialized after MPI finalization. MPI Sessions also offers the possibility for more flexible ways for individual components of an application to express the capabilities they require from MPI at a finer granularity than is presently possible.At this time, MPI Sessions has reached sufficient maturity for implementation and evaluation, which are the focuses of this paper. This paper presents a prototype implementation of MPI Sessions, discusses certain of its performance characteristics, and describes its successful use in a large-scale production MPI application. Overall, MPI Sessions is shown to be implementable, integrable with key infrastructure, and effective, but with certain overheads involving the initialization of MPI as well as communicator construction. Small impacts on message-passing latency and throughput are noted. Open MPI was used as the implementation vehicle, but results here are also relevant to other middleware stacks. Nathan T. Hjelm, Howard Pritchard, Samuel K. Gutierrez, Daniel J. Holmes, Ralph Castain, Anthony Skjellum |
CLUSTER | 6 |
| 2019 | GhostSZ: A Transparent FPGA-Accelerated Lossy Compression FrameworkabstractHigh-performance computing (HPC) applications often generate enormous amounts of data that must be transferred for check-pointing, in situ processing, or post-execution analysis. To reduce the related network traffic and storage consumption, lossy compression schemes that target scientific data are often used. SZ compression emerged three years ago and has gained much attention because of its high compression ratio. However, performing SZ compression can take half a day per Terabyte of data; this could be a drawback to adoption. We propose GhostSZ an FPGA framework for accelerating tasks in SZ at line rate, and so transparently. The critical problem to be overcome is the tight data dependence central to SZ. GhostSZ solves this with a data transfer path having novel staged hardware. We test our implementation with both synthetic and real HPC application data and show 9.5×-80× core versus pipeline speedup over the optimized production version running on a state-of-the-art CPU and 8.2× per chip. Much of the variance in performance is due to the FPGA already running at line rate and so benefiting less from optimizations applicable to the CPU only on the most favorable data sets. The significance of this work is the possibility of a major reduction in required networking and storage in HPC installations. For example, using GhostSZ, fewer than 10 FPGAs would be sufficient to handle the entire I/O bandwidth of the top entry on the latest IO-500 list. Qingqing Xiong, Rushi Patel, Chen Yang 0010, Tong Geng, Anthony Skjellum, Martin C. Herbordt |
FCCM | 5 |
| 2019 | User-Level Scheduled Communications for MPIabstractComposability is one of seven reasons for the long-standing and continuing success of MPI. Extending MPI by composing its operations with user-level operations provides useful integration with the progress engine and completion notification methods of MPI. However, the existing extensibility mechanism in MPI (generalized requests) is not widely utilized and has significant drawbacks. MPI can be generalized via scheduled communication primitives, for example, by utilizing implementation techniques from existing MPI-3 nonblocking collectives and from forthcoming MPI-4 persistent and partitioned APIs. Non-trivial schedules are used internally in some MPI libraries; but, they are not accessible to end-users. Message-based communication patterns can be built as libraries on top of MPI. Such libraries can have comparable implementation maturity and potentially higher performance than MPI library code, but do not require intimate knowledge of the MPI implementation. Libraries can provide performance-portable interfaces that cross MPI implementation boundaries. The ability to compose additional user-defined operations using the same progress engine benefits all kinds of general purpose HPC libraries. We propose a definition for MPI schedules: a user-level programming model suitable for creating persistent collective communication composed with new application-specific sequences of user-defined operations managed by MPI and fully integrated with MPI progress and completion notification. The API proposed offers a path to standardization for extensible communication schedules involving user-defined operations. Our approach has the potential to introduce event-driven programming into MPI (beyond the tools interface), although connecting schedules with events comprises future work. Early performance results described here are promising and indicate strong overlap potential. Derek Schafer, Sheikh K. Ghafoor, Daniel J. Holmes, Martin Ruefenacht, Anthony Skjellum |
HiPC | 5 |
| 2019 | Exposition, clarification, and expansion of MPI semantic terms and conventions: is a nonblocking MPI function permitted to block?abstractThis paper offers a timely study and proposed clarifications, revisions, and enhancements to the Message Passing Interface's (MPI's) Semantic Terms and Conventions. To enhance MPI, a clearer understanding of the meaning of the key terminology has proven essential, and, surprisingly, important concepts remain underspecified, ambiguous and, in some cases, inconsistent and/or conflicting despite 26 years of standardization. This work addresses these concerns comprehensively and usefully informs MPI developers, implementors, those teaching and learning MPI, and power users alike about key aspects of existing conventions, syntax, and semantics. This paper will also be a useful driver for great clarity in current and future standardization and implementation efforts for MPI. Purushotham V. Bangalore, Rolf Rabenseifner, Daniel J. Holmes, Julien Jaeger, Guillaume Mercier, Claudia Blaas-Schenner, Anthony Skjellum |
EuroMPI | 7 |
| 2019 | A large-scale study of MPI usage in open-source HPC applicationsabstractUnderstanding the state-of-the-practice in MPI usage is paramount for many aspects of supercomputing, including optimizing the communication of HPC applications and informing standardization bodies and HPC systems procurements regarding the most important MPI features. Unfortunately, no previous study has characterized the use of MPI on applications at a significant scale; previous surveys focus either on small data samples or on MPI jobs of specific HPC centers. This paper presents the first comprehensive study of MPI usage in applications. We survey more than one hundred distinct MPI programs covering a significantly large space of the population of MPI applications. We focus on understanding the characteristics of MPI usage with respect to the most used features, code complexity, and programming models and languages. Our study corroborates certain findings previously reported on smaller data samples and presents a number of interesting, previously un-reported insights. Ignacio Laguna, Ryan J. Marshall, Kathryn Mohror, Martin Ruefenacht, Anthony Skjellum, Nawrin Sultana |
SC | 5 |
| 2019 | Planning for performance: Enhancing achievable performance for MPI through persistent collective operations
Daniel J. Holmes, Bradley Morgan, Anthony Skjellum, Purushotham V. Bangalore, Srinivas Sridharan 0002 |
Parallel Comput. | 3 |
| 2019 | Failure recovery for bulk synchronous applications with MPI stages
Nawrin Sultana, Martin Ruefenacht, Anthony Skjellum, Ignacio Laguna, Kathryn Mohror |
Parallel Comput. | 3 |
| 2018 | Accelerating MPI Message Matching through FPGA OffloadabstractThe Message Passing Interface (MPI) is the de facto communication standard for distributed-memory High-Performance Computing (HPC) systems. Ultra-low latency communication in HPC is difficult to achieve because of MPI processing requirements, in particular matching requests and messages done by traversing the corresponding queues. Many researchers have addressed this issue by redesigning queues or by offloading them to hardware accelerators. However, state-of-art software approaches cannot free CPUs “from the misery” and hardware approaches either lack scalability or still leave substantial room for further improvement. With the emergence of numerous tightly coupled CPU-FPGA computing architectures, offload of MPI functionality to user-controlled hardware is now becoming viable; we find it productive to revisit hardware approaches. To maintain the generality necessary to support MPI while preventing high resource utilization, we design our MPI queue processing offload based on a recent analysis of performance characteristics in HPC applications. We propose a novel, two-level message queue design: a content addressable memory (CAM) coupled with a resource-saving hardware linked-list. We also propose an optimization that maintains high speed in the cases when the queue is long. To test our design, we create an SOC-based testbed consisting of softcore processors and hardware implementations of the MPI communication stacks. Even while using only a small fraction of the Stratix-V logic, our design can be one to two orders of magnitude faster than two well-known hardware designs. Qingqing Xiong, Anthony Skjellum, Martin C. Herbordt |
FPL | 2 |
| 2018 | MPI Stages: Checkpointing MPI State for Bulk Synchronous ApplicationsabstractWhen an MPI program experiences a failure, the most common recovery approach is to restart all processes from a previous checkpoint and to re-queue the entire job. A disadvantage of this method is that, although the failure occurred within the main application loop, live processes must start again from the beginning of the program, along with new replacement processes---this incurs unnecessary overhead for live processes. To avoid such overheads and concomitant delays, we introduce the concept of "MPI Stages." MPI Stages saves internal MPI state in a separate checkpoint in conjunction with application state. Upon failure, both MPI and application state are recovered, respectively, from their last synchronous checkpoints and continue without restarting the overall MPI job. Live processes roll back only a few iterations within the main loop instead of rolling back to the beginning of the program, while a replacement of failed process restarts and reintegrates, thereby achieving faster failure recovery. This approach integrates well with large-scale, bulk synchronous applications and checkpoint/restart. Nawrin Sultana, Anthony Skjellum, Ignacio Laguna, Matthew Shane Farmer, Kathryn Mohror, Murali Emani |
EuroMPI | 2 |
| 2018 | MPI Derived Datatypes: Performance and Portability IssuesabstractThis paper addresses performance-portability and overall performance issues when derived datatypes are used with four MPI implementations: Open MPI, MPICH, MVAPICH2, and Intel MPI. These comparisons are particularly relevant today since most vendor implementations are now based on Open MPI or MPICH rather than on vendor proprietary code as was more prevalent in the past. Our findings are that, within a single MPI implementation, there are significant differences in performance as a function of it reasonable encodings of derived datatypes as supported by the MPI standard. While this finding may not be surprising, it is important to understand how fundamental vs. arbitrary choices made in early implementation impact the use of derived datatypes to date. Qingqing Xiong, Purushotham V. Bangalore, Anthony Skjellum, Martin C. Herbordt |
EuroMPI | 3 |
| 2016 | Using machine learning to secure IoT systemsabstractThe Internet of Things (IoT) is a massive group of devices containing sensors or actuators connected together over wired or wireless networks. With an estimate of over 25 billion devices connected together by 2020, IoT has been rapidly growing over the past decade. During the growth, security has been identified as one of the weakest areas in IoT. When implementing security within an IoT network, there are several challenges including heterogeneity within the system as well as the quantity of devices that need to be addressed. To approach the challenges in securing IoT devices, we propose using machine learning within an IoT gateway to help secure the system. We investigate using Artificial Neural Networks in a gateway to detect anomalies in the data sent from the edge devices. We are convinced that this approach can improve the security of IoT systems. Janice Canedo, Anthony Skjellum |
PST | 2 |
| 2016 | Provenance threat modelingabstractProvenance systems are used to capture history metadata, applications include ownership attribution and determining the quality of a particular data set. Provenance systems are also used for debugging, process improvement, understanding data proof of ownership, certification of validity, etc. The provenance of data includes information about the processes and source data that leads to the current representation. In this paper we study the security risks provenance systems might be exposed to and recommend security solutions to better protect the provenance information. Oluwakemi Hambolu, Lu Yu 0001, Jon Oakley 0001, Richard R. Brooks, Ujan Mukhopadhyay, Anthony Skjellum |
PST | 6 |
| 2016 | A brief survey of Cryptocurrency systemsabstractCryptocurrencies have emerged as important financial software systems. They rely on a secure distributed ledger data structure; mining is an integral part of such systems. Mining adds records of past transactions to the distributed ledger known as Blockchain, allowing users to reach secure, robust consensus for each transaction. Mining also introduces wealth in the form of new units of currency. Cryptocurrencies lack a central authority to mediate transactions because they were designed as peer-to-peer systems. They rely on miners to validate transactions. Cryptocurrencies require strong, secure mining algorithms. In this paper we survey and compare and contrast current mining techniques as used by major Cryptocurrencies. We evaluate the strengths, weaknesses, and possible threats to each mining strategy. Overall, a perspective on how Cryptocurrencies mine, where they have comparable performance and assurance, and where they have unique threats and strengths are outlined. Ujan Mukhopadhyay, Anthony Skjellum, Oluwakemi Hambolu, Jon Oakley 0001, Lu Yu 0001, Richard R. Brooks |
PST | 2 |
| 2016 | MPI Sessions: Leveraging Runtime Infrastructure to Increase Scalability of Applications at ExascaleabstractMPI includes all processes in MPI_COMM_WORLD; this is untenable for reasons of scale, resiliency, and overhead. This paper offers a new approach, extending MPI with a new concept called Sessions, which makes two key contributions: a tighter integration with the underlying runtime system; and a scalable route to communication groups. This is a fundamental change in how we organise and address MPI processes that removes well-known scalability barriers by no longer requiring the global communicator MPI_COMM_WORLD. Daniel J. Holmes, Kathryn Mohror, Ryan E. Grant, Anthony Skjellum, Martin Schulz 0001, Wesley Bland, Jeffrey M. Squyres |
EuroMPI | 4 |
| 2015 | OCF: An Open Cloud Forensics Model for Reliable Digital ForensicsabstractThe rise of cloud computing has changed the way computing services and resources are used. However, existing digital forensics science cannot cope with the black-box nature of clouds nor with multi-tenant cloud models. Because of the fundamental characteristics of clouds, many assumptions of digital forensics are invalidated in clouds. In the digital forensics process involving clouds, the role of cloud service providers (CSP) is utterly important, a role which needs to be considered in the science of cloud forensics. In this paper, we define cloud forensics considering the role of the CSP and propose the Open Cloud Forensics (OCF) model. Based on this OCF model, we propose a cloud computing architecture and validate our proposed model using a case study, which is inspired from an actual civil lawsuit. Shams Zawoad, Ragib Hasan, Anthony Skjellum |
CLOUD | 3 |
| 2014 | Design and Evaluation of FA-MPI, a Transactional Resilience Scheme for Non-blocking MPIabstractWith the rapid scale out of supercomputers comes a corresponding higher failure frequency. Fault-tolerant methods have evolved to adapt to high rates of failure, but the behavior of MPI, the most widely used scalable programming middleware, is insufficient when confronting such failures. We present FA-MPI (Fault-Aware MPI), a set of extensions to the MPI standard designed to enable applications to implement a wide range of fault-tolerant methods. FA-MPI introduces transactional concepts to the MPI programming model for the first time to address failure detection, isolation, mitigation, and recovery via application-driven policies. To reach the maximum achievable performance of these scalable machines, overlapping communication and I/O with computation through non-blocking operations (while reducing jitter) are design themes of growing importance. Therefore, we emphasize fault tolerant, non-blocking communication operations combined with a set of nest able lightweight transactional Try Block API extensions architected to exploit system and application hierarchy both for failure detection and recovery. This is to enable applications to run to completion with higher probability than otherwise. Scaling up and out and fault-free overhead are key concerns that can be managed by tuning transaction granularity, we provide a simulation of FA-MPI in a stencil 3D program to illustrate this. Supported failure models include but are not limited to process failures, a key difference from other proposed fault-tolerant extensions to MPI. Restriction to non-blocking operations is a current limitation as compared to other proposed approaches insofar as legacy applications are concerned, but FA-MPI aligns well with future-looking applications emphasizing Exascale. And, tools to evolve legacy MPI programs to this fault-aware paradigm will soon bridge that portability gap. Amin Hassani, Anthony Skjellum, Ron Brightwell |
DSN | 2 |
| 2014 | A Lightweight Data Location Service for Nondeterministic Exascale Storage SystemsabstractIn this article, we present LWDLS, a lightweight data location service designed for Exascale storage systems (storage systems with order of 10 18 bytes) and geo-distributed storage systems (large storage systems with physically distributed locations). LWDLS provides a search-based data location solution, and enables free data placement, movement, and replication. In LWDLS, probe and prune protocols are introduced that reduce topology mismatch, and a heuristic flooding search algorithm (HFS) is presented that achieves higher search efficiency than pure flooding search while having comparable search speed and coverage to the pure flooding search. LWDLS is lightweight and scalable in terms of incorporating low overhead, high search efficiency, no global state, and avoiding periodic messages. LWDLS is fully distributed and can be used in nondeterministic storage systems and in deterministic storage systems to deal with cases where search is needed. Extensive simulations modeling large-scale High Performance Computing (HPC) storage environments provide representative performance outcomes. Performance is evaluated by metrics including search scope, search efficiency, and average neighbor distance. Results show that LWDLS is able to locate data efficiently with low cost of state maintenance in arbitrary network environments. Through these simulations, we demonstrate the effectiveness of protocols and search algorithm of LWDLS. Anthony Skjellum, Lee Ward, Matthew L. Curry |
ACM Trans. Storage | 2 |
| 2013 | Design, implementation, and performance evaluation of MPI 3.0 on portals 4.0abstractThe latest version of the Portals interconnect programming interface contains several improvements intended for MPI point-to-point and one-sided communication operations, including new functionality defined in MPI-3. This paper discusses the rationale for these improvements to Portals and describes how they can be used in an MPI implementation. We provide preliminary micro-benchmark performance results using a reference implementation of Portals over InfiniBand Verbs and the Open MPI implementation. Amin Hassani, Anthony Skjellum, Ron Brightwell, Brian W. Barrett |
EuroMPI | 2 |
| 2011 | Gibraltar: A Reed-Solomon coding library for storage applications on programmable graphics processorsabstractSUMMARY Reed–Solomon coding is a method for generating arbitrary amounts of erasure correction information from original data via matrix–vector multiplication in finite fields. Previous work has shown that modern CPUs are not well‐matched to this type of computation, requiring applications that depend on Reed–Solomon coding at high speeds (such as high‐performance storage arrays) to use hardware implementations. This work demonstrates that high performance is possible with current cost‐effective graphics processing units across a wide range of operating conditions and describes how performance will likely evolve in similar architectures. It describes the characteristics of the graphics processing unit architecture that enable high‐speed Reed–Solomon coding. A high‐performance practical library, Gibraltar, has been prototyped that performs Reed–Solomon coding on graphics processors in a manner suitable for storage arrays, along with applications with similar data resiliency needs. This library enables variably resilient erasure correcting codes to be used in a broad range of applications. Its performance is compared with that of a widely available CPU implementation, and a rationale for its API is presented. Its practicality is demonstrated through a usage example. Copyright © 2011 John Wiley & Sons, Ltd. Matthew L. Curry, Anthony Skjellum, Lee Ward, Ron Brightwell |
Concurr. Comput. Pract. Exp. | 2 |
| 2010 | A Lightweight, GPU-Based Software RAID SystemabstractWhile RAID is the prevailing method of creating reliable secondary storage infrastructure, many users desire more flexibility than offered by current implementations. Traditionally, RAID capabilities have been implemented largely in hardware in order to achieve the best performance possible, but hardware RAID has rigid designs that are costly to change. Software implementations are much more flexible, but software RAID has historically been viewed as much less capable of high throughput than hardware RAID controllers. This work presents a system, Gibraltar RAID, that attains high RAID performance by offloading the calculations related to error correcting codes to GPUs. This paper describes the architecture, performance, and qualities of the system. A comparison to a well-known software RAID implementation, the md driver included with the Linux operating system, is presented. While this work is presented in the context of high performance computing, these findings also apply to a general RAID market. Matthew L. Curry, Lee Ward, Anthony Skjellum, Ron Brightwell |
ICPP | 3 |
| 2008 | Accelerating Reed-Solomon coding in RAID systems with GPUsabstractGraphical Processing Units (GPUs) have been applied to more types of computations than just graphics processing for several years. Until recently, however, GPU hardware has not been capable of efficiently performing general data processing tasks. With the advent of more general-purpose extensions to GPUs, many more types of computations are now possible. One such computation that we have identified as being suitable for the GPU’s unique architecture is Reed-Solomon coding in a manner appropriate for RAID-type systems. In this paper, we motivate the need for RAID with triple-disk parity and describe a pipelined architecture for using a GPU for this purpose. Performance results show that the GPU can outperform a modern CPU on this problem by an order of magnitude and also confirm that a GPU can be used to support a system with at least three parity disks with no performance penalty. Matthew L. Curry, Anthony Skjellum, Lee Ward, Ron Brightwell |
IPDPS | 2 |
| 2006 | Grid-Flow: a Grid-enabled scientific workflow system with a Petri-net-based interfaceabstractAbstract Advances in computer technologies have enabled scientists to explore research issues in their respective domains at scales greater and finer than ever before. The availability of efficient data collection and analysis tools presents researchers with vast opportunities to process heterogeneous data within a distributed environment. To support the opportunities enabled by massive computation, a suitable scientific workflow system is needed to help the users to manage data and programs, and to design reusable procedures of scientific experimental tasks. In this paper, the design and prototype implementation of a scientific workflow infrastructure, called Grid‐Flow, is presented. Grid‐Flow assists researchers in specifying scientific experiments using a Petri‐net‐based interface. The Grid‐Flow infrastructure is designed as a Service Oriented Architecture with multi‐layer component models. The contributions of Grid‐Flow are as follows: (1) a new, lightweight, programmable Grid workflow language, Grid‐Flow Description Language, is provided to describe the workflow process in a Grid environment; (2) a Petri‐net‐based user interface, based on the Generic Modeling Environment, is demonstrated to help the user design the workflow process with a Petri‐net model; and (3) a program integration component of the Grid‐Flow system is presented to integrate all possible programs into the system. Copyright © 2005 John Wiley & Sons, Ltd. Zhijie Guan, Francisco Hernández, Purushotham V. Bangalore, Jeffrey G. Gray, Anthony Skjellum, Vijay Velusamy |
Concurr. Comput. Pract. Exp. | 5 |
| 2005 | inAspect: interfacing Java and VSIPL applicationsabstractAbstract In this paper, we discuss the origin, design, performance, and directions of the inAspect high‐performance signal‐ and image‐processing package for Java. The Vector Signal and Image Processing Library (VSIPL) community provides a standardized application programmer interface (API) for high‐performance signal and image processing plus linear algebra with a C emphasis and object‐based design framework. Java programmers need high‐performance and/or portable APIs for this broad base of functionality as well. inAspect addresses PDAs, embedded Java boards, workstations, and servers, with emphasis on embedded systems at present. Efforts include supporting integer precisions and utilizing coordinate rotation digital computer (CORDIC) algorithms—both aimed at added relevance for limited‐performance environments, such as present‐day PDAs. Copyright © 2005 John Wiley & Sons, Ltd. Torey Alford, Vijay P. Shah, Anthony Skjellum, Nicolas H. Younan, Clayborne D. Taylor |
Concurr. Pract. Exp. | 3 |
| 2005 | Lightweight monitoring of MPI programs in real timeabstractCurrent technologies allow efficient data collection by several sensors to determine an overall evaluation of the status of a cluster. However, no previous work of which we are aware analyzes the behavior of the parallel programs themselves in real time. In this paper, we perform a comparison of different artificial intelligence techniques that can be used to implement a lightweight monitoring and analysis system for parallel applications on a cluster of Linux workstations. We study the accuracy and performance of deterministic and stochastic algorithms when we observe the flow of both library-function and operating-system calls of parallel programs written with C and MPI. We demonstrate that monitoring of MPI programs can be achieved with high accuracy and in some cases with a false-positive rate near 0% in real time, and we show that the added computational load on each node is small. As an example, the monitoring of function calls using a hidden Markov model generates less than 5% overhead. The proposed system is able to automatically detect deviations of a process from its expected behavior in any node of the cluster, and thus it can be used as an anomaly detector, for performance monitoring to complement other systems or as a debugging tool. Copyright © 2005 John Wiley & Sons, Ltd. German Florez, Susan M. Bridges, Anthony Skjellum, Rayford B. Vaughn |
Concurr. Comput. Pract. Exp. | 4 |
| 2005 | A spectral estimation toolkit for Java applications
Vijay P. Shah, Nicolas H. Younan, Torey Alford, Anthony Skjellum |
Sci. Comput. Program. | 4 |
| 2004 | The Real-Time Message Passing Interface Standard (MPI/RT-1.1) abstractThe Real-Time Message Passing Interface (MPI/RT) standard is the product of the work of many people working in an open community standards group over a period of over six years. The purpose of this archival publication is to preserve the significant knowledge and experience that was developed in real-time message-passing systems as a consequence of the research and development effort as well as in the specification of the standard. Interestingly, several implementations of MPI/RT (as well as comprehensive test suites) have been created in industry and academia over the period during which the standard was created. MPI/RT is likely to gain adoption interest over time, and this adoption may be driven by the promulgation of the standard including this publication. We expect that, when people are interested in understanding options for reliable, quality of service (QoS)-oriented parallel computing with message passing, MPI/RT will serve as a foundation for such a study, whether or not its complete formalism is accepted into other systems or standards. MPI/RT is an offshoot of MPI-1, and retains many of the communication patterns of MPI-1. However, MPI/RT has investigated issues of fine-grain concurrency and highest achievable performance in many ways that were evidently inappropriate for MPI-1 in the scientific computing space in which it resides, with its much broader audience. MPI/RT focuses on early-binding (planned transfer), concurrent message passing, while integrating multiple real-time models: time-based, event-driven, and priority-oriented channels. Group admission control and declarative (deferred early binding) semantics support the goal of hard-real-time for the message-passing component of computation. Importantly, MPI/RT emphasizes the decoupling of message transfer and process/thread scheduling as part of its contribution to parallel processing with QoS. Buffer management and state transition diagrams are also integral to the notion of streaming data into and out of processors in a way that is consistent with QoS, and friendly to zero-copy approaches to communication. MPI/RT has also made strides in the direction of a parallel middleware specification by emphasizing an object-oriented design for the application programmer interface (API) compared with an object-based API or ad hoc API. The advantages of these are plain in the standard, in that the functionality has useful polymorphic adaptations where needed. Furthermore, the concepts that derive from MPI-1 (such as collective operations) appear as objects in MPI/RT. This modification has allowed for the removal of certain constructs in MPI-1 (such as the communicator), in favor of a specification and implementation phase for objects that describe communication in MPI/RT. Overall, a cleaner, more extensible design exists, which does not utilize more resources per se than those which are needed to admit the required channels for a program. Both offline and online admission control is contemplated, and multiple modes are supported, albeit weakly. MPI/RT-1.1, the standard version described here, does not cover all possible real-time parallel programming possibilities. It is silent concerning process/thread scheduling, so, in some sense, still has strong aspects of ‘best effort’, in terms of process scheduling. Leaving process scheduling as an orthogonal concern was intentional, so that the best concepts in these areas would be used in concert with MPI/RT, rather than offering a monolith. Furthermore, MPI/RT-1.1 does not explicitly address mode changes (with guaranteed mode-change QoS) between sets of channels, with invariants and non-invariants among the resources consumed. This remains important work for the future. The object-oriented, resource-conscious approach of MPI/RT naturally extends to multiple modes. Work to realize this in a standard or in prototypes remains for future work, although it was discussed and prototyped extensively during standardization. It is interesting to consider whether the connection-oriented, but limited QoS and resource specification of MPI/RT leads to a more-scalable or less-scalable system environment than that posed by MPI and similar middleware. While connections themselves indicate that resources will be assigned per connection, only those connections that are program-mandated are actually built. By way of contrast, in MPI it is necessary to offer a virtual all-to-all communication topology, and introduce overheads associated with either the static realization of such a topology, or else the dynamic build-up/tear-down of connections in constrained environments seeking to scale. Events over the past six years involving the evolution of networking technology make MPI/RT as interesting as it was when started, and possibly of more ubiquitous application in the long term. Infiniband, Rapid I/O, 3GIO, and other System Area Network standards are likely to offer rudimentary QoS over time in real applications. Likewise, the production of massively concurrent supercomputers (104 nodes or more), is likely to drive the need for predictable message passing in the runtime aspects of such systems. These events are likely to cause the ideas and concepts defined in this standard to have impact in areas far broader than originally anticipated. Copyright © 2004 John Wiley & Sons, Ltd. Contents Preface Six 1 Introduction S1 1.1 General introduction S1 1.1.1 Parallel models S2 1.1.2 ‘Sidedness’ of communication S3 1.1.3 Real-time models and QoS S3 1.1.4 Ontogeny of an MPI/RT application S4 1.1.5 The MPI/RT API S6 1.2 Introduction for users S12 1.3 Introduction for implementors S13 1.3.1 The basics S13 1.3.2 The admission test S14 1.3.3 Other advice to implementors S15 1.4 Error checking and kinds of libraries S15 1.4.1 Erroneous programs S15 1.4.2 Conformance and kinds of libraries S16 1.4.3 Error reporting S17 1.4.4 String representation of error codes S18 1.5 Related work S18 1.5.1 Admission control and resource reservation S18 1.5.2 Access arbitration and transmission control S19 1.5.3 Early results S20 1.6 Summary S23 I Concepts and basic objects S25 2 MPI/RT objects S27 2.1 Overview S27 2.2 Behavior of objects in MPI/RT S32 2.3 Generic operations defined on all MPI/RT objects S33 2.4 Attributes: object decoration S37 2.4.1 Keyval object parameter accessors S41 2.4.2 Object attribute manipulation functions S42 2.5 Containers S44 2.5.1 Container constructors S45 2.5.2 Generic container operations S45 2.5.3 Set operations S47 2.5.4 Vector base operations S48 2.5.5 Container iterators S50 2.6 Groups S53 2.6.1 Group definition S54 2.6.2 Group management S54 2.7 Miscellaneous objects S56 2.7.1 MPIRT_TASK_ADDRESS S56 2.8 Summary S57 3 Dataspecs and types S65 3.1 Overview S65 3.2 Dataspecs S65 3.2.1 Operations on dataspec S65 3.3 Predefined MPI/RT types S66 3.3.1 MPIRT_BOOLEAN S67 3.3.2 MPIRT_STRING_NAME S68 3.3.3 MPIRT_INT64 S68 3.3.4 MPIRT_TIME_SPEC S71 3.3.5 MPIRT_ADDRESS S72 3.3.6 MPIRT_BUFITER_MODE S72 3.4 Summary S73 4 Event delivery abstraction and handlers S79 4.1 Introduction S79 4.2 Event delivery abstraction S79 4.2.1 Event naming S81 4.2.2 Common event delivery abstraction operations S83 4.2.3 Triggers S85 4.2.4 Event receptors S87 4.3 Handlers S94 4.3.1 Handler constructors S94 4.3.2 Light-weight handler functions S96 4.3.3 Handler accessors S98 4.4 Waiting for an event S100 4.5 Summary S101 5 Buffer management S109 5.1 Introduction S109 5.2 Buffer object S109 5.3 Buffer object functions S115 5.3.1 Variable length transfers S115 5.3.2 Variable buffer offsets S116 5.4 Buffer operations S117 5.4.1 Operations on buffer labels S117 5.4.2 Buffer-partitioning operations S118 5.5 Buffer iterator S121 5.6 Buffer iterator accessors S130 5.7 Bufiter modes S132 5.8 Summary S134 II Transfer mechanisms and advanced objects S139 6 Channel overview S141 6.1 Introduction S141 6.2 Common attribute operations for channels S143 6.3 Summary S146 7 Point-to-point channels S149 7.1 Introduction S149 7.2 Operations on the point-to-point channel object S149 7.3 Summary S151 8 Collective channels S153 8.1 Introduction S153 8.2 Broadcast collective channel S153 8.3 Gather operation channel S156 8.4 Scatter operation channel S160 8.5 Reduce operation channel S164 8.5.1 Predefined reduce operations S168 8.6 Barrier operation channel S169 8.7 All-to-all channel S170 8.8 Summary S172 9 Channel operations S181 9.1 Data transfers S181 9.1.1 Performance considerations S181 9.1.2 Channel states and transitions S182 9.1.3 Methods for single message transfer S188 9.1.4 Methods for multiple message transfers S190 9.2 Testing completion and determining the state of data transfers S191 9.2.1 Wait operation S192 9.2.2 Test operation S192 9.3 Summary S193 III Real-time programming models and QoS S195 10 QoS overview S197 10.1 Introduction S197 10.2 Time-driven real-time programming model S198 10.2.1 Scheduling message transfers S199 10.2.2 Schedulable time intervals S199 10.2.3 The MPI/RT time specification S200 10.3 Event-driven real-time programming model S200 10.3.1 Overview S201 10.3.2 Event triggers S202 10.3.3 Event receptors S204 10.4 Priority-driven real-time programming model S204 10.4.1 Channel priority S205 10.4.2 Process priority S206 10.5 Best-effort QoS programming model S206 11 QoS specification S207 11.1 Introduction S207 11.2 Channel QoS specification S207 11.2.1 Specification for time-driven channels S208 11.2.2 Relationship of time-based schedules for different channels S211 11.2.3 Specification for event-driven-with-priority channels S211 11.2.4 Specification for combined event and time-driven with priority channels S219 11.3 Handler QoS specification S225 11.4 Event-delivery abstraction's QoS specification S227 11.4.1 QoS for triggers S227 11.4.2 QoS for receptors S230 11.5 Summary S232 12 Committing objects and resource allocation S239 12.1 Commit operation S239 12.2 Summary S241 IV Environmental mechanisms and functionality S243 13 Initialization and termination S245 13.1 Initialization and termination of MPI/RT S245 13.2 Version information S247 13.3 Summary S248 14 Clocks S251 14.1 Synchronization of clocks S251 14.2 Description of the clocks S251 14.3 Clock synchronization parameters S252 14.3.1 The epoch S252 14.3.2 The MPIRT_TIME type S253 14.3.3 The synchronized time service S253 14.3.4 Parameters S253 14.4 Behavior of the time services S256 14.5 Timed waiting S256 14.6 Summary S257 15 Instrumentation S259 15.1 Introduction S259 15.2 MPI/RT metrics S260 15.3 MPI/RT probes S261 15.4 MPI/RT user metrics S265 15.5 Summary S270 V Appendices S273 A Return codes S275 A.1 Return codes S275 B Deprecated functionality S281 B.1 Functionality deprecated in MPI/RT-1.1 S281 B.1.1 MPIRT_ERR_COMMITTED_OBJECT S281 B.1.2 MPIRT_ERR_INITIALIZED S281 B.1.3 MPIRT_CSET_RETRIEVE_NEXT S281 B.1.4 MPIRT_ERR_ACTIVE_CHANNEL S282 Acknowledgments S283 Glossary S287 Bibliography S295 MPI/RT return code index S299 MPI/RT function index S303 MPI/RT entity index S319 Index S331 Preface PREFACE TO THE CURRENT STANDARD VERSION, MPI/RT-1.1 At the conclusion of MPI/RT-1.0, many of us participating and contributing to the standard recognized the need to continue to improve certain key features, in pursuit of a system with extremely low cost of portability, and to widen the applicability of MPI/RT. A restrained set of extensions have emerged in MPI/RT-1.1, from among a huge set of proposals, ideas, and concepts. While assiduously trying to avoid the ‘second system syndrome’, MPI/RT-1.1 works hard to fix small issues and make the standard easier to use, better, and more applicable. This document conflates the contributions of MPI/RT-1.0, the principal work, with the newly accepted developments of MPI/RT-1.1, plus errata and other improvements designed to keep this document as the main reference for understanding how to implement and use MPI/RT. A subset of the original participants in MPI/RT-1.0, together with some new participants, have built this standard extension, with the view that future extensions (whether termed MPI/RT-1.2 or MPI/RT-2) would come much later, after a period of two or more years of implementation and usage. While initial implementations and experience with MPI/RT-1.0 have driven MPI/RT-1.1 in part, much room remains for further implementation and experience in various application spaces. Historical perspective will of course assess the validity, efficacy, and overall impact of building application programmer interface (API) standards in the way that MPI/RT-1.0 and MPI/RT-1.1 have been done, namely, by a small group of dedicated individuals, supported by a larger group of interested participants from the application, user, and research communities of both private and public sectors. The shoe-string funding associated with MPI/RT over the last three years of its six-year life has actually energized, rather than diminished, the energy for such progress and success. What appears clear, however, is that a tremendous amount of useful computer science related to advanced systems programming of middleware with real-time has been captured in this standards document and its first extension, MPI/RT-1.1. This stable intermediate form will clearly play an important role in the future exploration of real-time middleware for scalable, parallel processing. The efforts of the MPI/RT Forum and its members over the past six years reflect the strong commitment, perseverance, and tireless efforts of its individual participants, for which the chairs offer their sincerest gratitude. Anthony Skjellum, Starkville, MS Arkady Kanevsky, Waltham, MA March 2001 PREFACE TO MPI/RT-1.0 In 1995, several researchers and practitioners became interested in advancing real-time extensions to the then existing Message Passing Interface (MPI) de facto standard and began meeting informally. Later, the group became a sanctioned subcommittee of the MPI-2 de facto standards body, which met regularly in Chicago. People involved with high-performance computing, distributed computing, message-passing systems, and real-time systems were all represented. Researchers and practitioners from industry, academia and defense laboratories were included. The MPI-2 Forum condoned this effort and allowed it to blossom as a Journal of Development activity, with a clear view that it would not formally be part of MPI-2, but nonetheless was a worthwhile working activity. This status was productive and helpful to the work because of the valuable proximity to many interested in messaging, without the compelling deadline faced by MPI-2. From a technical perspective, a lot of issues, approaches, requirements, and techniques evolved, and significant new ideas previously thought about were introduced into the discussion. These issues included provision of key concepts not available or readily addressable within the confines of MPI-1 [1] and MPI-2 [2]: channels, real-time models, predictability, greater support for thread interactions, early-binding strategies, and admission tests. The requirements posed by the subcommittee emerged as follows: achieve highest performance messaging and add the additional constraints of predictability and quality-of-service (QoS), with the additional capability to support relevant memory management to enhance the elimination of data copies and support for the ‘cut-through’ of data. These requirements drove us over time, sometimes systematically and sometimes ad hoc, to re-examine much of what was decided in earlier messaging systems and ultimately to evolve away from explicit upward or downward compatibility with MPI. The outcome of this effort is a ‘lower middleware’ standard called MPI/RT, which strives to offer extremely low cost of portability as compared with any native software architecture for messaging, while providing useful real-time notions of performance, predictability, and QoS. Other major decisions included the elimination of Fortran77 language bindings in favor of C++ language bindings and the thorough and continued use of object-oriented APIs and design methodologies to motivate and support the process. Positioning MPI as conceptually ‘higher middleware’ in the form of a layer on top of MPI/RT establishes a conceptual relationship between this work and the previous standard. In fact, we expect that MPI implementations may actually be layered over MPI/RT on systems where users require both notations, and this ‘layerability’ is mentioned in appropriate parts of the standard. With encouragement and support from DARPA, and from the strong commitment of many of the subcommittee participants, including people from the mainstream of MPI Forum participants, significant progress was made over the first 18 months. The group continued to meet and burgeoned into a full-scale de facto effort of its own after the conclusion of the MPI-2 standards effort. Early versions of this effort appear in the ‘Journal of Development’ of the MPI-2 standard, but the results presented in that snapshot are quite different from what we have ultimately accomplished. This three-year effort has led to quite a satisfactory messaging middleware specification and standard that we expect to see deployed by industry in real-time computing multicomputers and networks of workstations. The group is committed to extending the specification in a limited fashion in 1999 to support channel input/output (I/O), dynamic processes, and a few other features intentionally delayed at present. After that, sincere efforts to introduce MPI/RT to a formal standard body will be undertaken by us and others, in order to help assure its long-term acceptance. In order to facilitate of this we have explicitly all to the MPI standards within the document into in order to additional information of to while the main and of the document of on either MPI-1 or MPI-2 standard knowledge of MPI-1 or MPI-2 is needed to with MPI/RT. In this we issues where is or different decisions have been while on the other and without much from new to MPI/RT. part of this work, we have own Journal of Development that has been into two of is to be in the MPI/RT-1.1 in is ‘best results and issues out during the past years that we to A but not formally best document is a related outcome of this issues as to other and a subset of MPI/RT-1.1 in that other This other body of results remains valuable the community we are seeking to The complete specification for MPI/RT is presented in this 1 a overview of the standard for 2 the object-oriented design of MPI/RT and the objects in describe the basic the and buffer describe the data transfer real-time QoS issues and address the issues for MPI/RT programs such as synchronized and performance of interest to such as a of how to use a of MPI/RT, is by to of interest to implementors of MPI/RT is by to The features are by Arkady Kanevsky, MA Anthony Skjellum, Starkville, MS Anthony Skjellum, Arkady Kanevsky, Yoginder S. Dandass, Jerrell Watts, Steve Paavola, Dennis Cottel, Greg Henley, L. Shane Hebert, Zhenqian Cui, Anna Rounbehler |
Concurr. Pract. Exp. | 1 |
| 2003 | PromisQoS: An Architecture for Delivering QoS to High-Performance Applications on Myrinet ClustersabstractClusters of workstations are being extensively used for solving computationally intensive scientific problems. However, there is limited support for quality of service (QoS) based distributed computing on commercial off- the-shelf (COTS) clusters. This limitation has restricted successful deployment of distributed real-time high-performance computing applications to customized and dedicated embedded multi-processor systems. This paper describes research work that attempts to provide a cluster platform that can guarantee access to computational and communication resources to distributed applications. The authors have developed PromisQoS, an architecture that supports execution of hard real-time distributed applications on a Linux cluster while providing high-throughput and low-latency communication using Myrinet. PromisQoS consists of the following major components - Hare, BDM-RT and Turtle. Hare is a prototype implementation of time-based QoS channels specified by the real-time message passing interface (MPI/RT 1.1) standard. BDM-RT is a low-level messaging library on Myrinet that provides deterministic communication latency and bandwidth on Myrinet. Turtle, a variant of RT-Linux, is the real-time operating system that provides guaranteed computation time. This work demonstrates that it is possible to deploy hard real-time distributed applications on COTS clusters and underlines the significance of the MPI/RT API in the realm of distributed high-performance computing applications that require QoS. Jothi P. Neelamegam, Srigurunath Chakravarthi, Manoj Apte, Anthony Skjellum |
LCN | 4 |
| 2002 | A Fine-Grain Clock Synchronization Mechanism for Myrinet ClustersabstractClock synchronization is a fundamental requirement for any real-time distributed system operating with global schedules. The presence of a global clock permits co-scheduling, enhances the degree of synchronization between cooperating tasks of parallel programs, and permits implementation of quality of service (QoS) based communication. This paper describes the design and implementation of a high accuracy (/spl plusmn/5/spl mu/s) global clock on a Myrinet network of PC with low software overhead. The clock synchronization scheme uses a novel approach to reduce errors arising from variation of latency of transmitted clock values. The contribution of this work is that high-accuracy of clock synchronization is obtained without significant resource consumption. This makes the synchronization scheme equally suitable for high-performance and real-time distributed systems. The new approach achieves clock accuracy of the order of /spl plusmn/5/spl mu/s, compared to the other known clock synchronization algorithm on Myrinet that achieves accuracy of the order of milliseconds. Programmability of the Myrinet interface card and the presence of an on-board real time clock are critical to achieving this accuracy. Srigurunath Chakravarthi, Anand Pillai, Jothi P. Neelamegam, Manoj Apte, Anthony Skjellum |
LCN | 5 |
| 2002 | A framework for high-performance matrix multiplication based on hierarchical abstractions, algorithms and optimized low-level kernelsabstractAbstract Despite extensive research, optimal performance has not easily been available previously for matrix multiplication (especially for large matrices) on most architectures because of the lack of a structured approach and the limitations imposed by matrix storage formats. A simple but effective framework is presented here that lays the foundation for building high‐performance matrix‐multiplication codes in a structured, portable and efficient manner. The resulting codes are validated on three different representative RISC and CISC architectures on which they significantly outperform highly optimized libraries such as ATLAS and other competing methodologies reported in the literature. The main component of the proposed approach is a hierarchical storage format that efficiently generalizes the applicability of the memory hierarchy friendly Morton ordering to arbitrary‐sized matrices. The storage format supports polyalgorithms, which are shown here to be essential for obtaining the best possible performance for a range of problem sizes. Several algorithmic advances are made in this paper, including an oscillating iterative algorithm for matrix multiplication and a variable recursion cutoff criterion for Strassen's algorithm. The authors expose the need to standardize linear algebra kernel interfaces, distinct from the BLAS, for writing portable high‐performance code. These kernel routines operate on small blocks that fit in the L1 cache. The performance advantages of the proposed framework can be effectively delivered to new and existing applications through the use of object‐oriented or compiler‐based approaches. Copyright © 2002 John Wiley & Sons, Ltd. Vinod Valsalam, Anthony Skjellum |
Concurr. Comput. Pract. Exp. | 2 |
| 2001 | MPI/FTTM: Architecture and Taxonomies for Fault-Tolerant, Message-Passing Middleware for Performance-Portable Parallel ComputingabstractMPI has proven effective for parallel applications in situations with neither QoS nor fault handling. Emerging environments motivate fault-tolerant MPI middleware. Environments include space-based, wide-area/web/meta computing and scalable clusters. MPI/FT, the system described in the paper, trades off sufficient MPI fault coverage against acceptable parallel performance, based on mission requirements and constraints. MPI codes are evolved to use MPI/FT features. Non-portable code for event handlers and recovery management is isolated. User-coordinated recovery, checkpointing, transparency and event handling, as well as evolvability of legacy MPI codes form key design criteria. Parallel self-checking threads address four levels of MPI implementation robustness, three of which are portable to any multithreaded MPI. A taxonomy of application types provides six initial fault-relevant models; user-transparent parallel nMR computation is thereby considered. Key concepts from MPI/RT-real-time MPI-are also incorporated into MPI/FT, with further overt support for MPI/RT and MPI/FT in applications possible in future. Rajanikanth Batchu, Anthony Skjellum, Zhenqian Cui, Murali Beddhu, Jothi P. Neelamegam, Yoginder S. Dandass, Manoj Apte |
CCGRID | 2 |
| 2001 | Synchronized Real-Time Linux Based Myrinet Cluster for Deterministic High Performance Computing and MPI/RTabstractThis paper describes the design and implementation of a real-time cluster of PCs that provides globally synchronized scheduling and predictable messaging passing. A high-accuracy, fine-grain global clock implementation has been closely coupled with a time-based scheduler to facilitate finely synchronized global scheduling across the cluster. A real-time messaging layer provides predictable communication latencies over the interconnecting Myrinet network. This system provides a solid framework for QoS based real-time communication, and in particular facilitates efficient layering of high-performance real-time distributed middleware such as MPI/RT. We present experiments and results to demonstrate the degree of predictability and synchronization achieved on an 8-node Myrinet cluster. 1 Introduction It is well known that clusters of commodity workstations are becoming a popular low cost alternative to super-computers for high-performance distributed computing [3]. Such systems also offe... Manoj Apte, Srigurunath Chakravarthi, Jothi Padmanabhan, Anthony Skjellum |
IPDPS | 4 |
| 2001 | Object-oriented analysis and design of the Message Passing InterfaceabstractAbstract The major contribution of this paper is the application of modern analysis techniques to the important Message Passing Interface standard, work done in order to obtain information useful in designing both application programmer interfaces for object‐oriented languages, and message passing systems. Recognition of ‘Design Patterns’ within MPI is an important discernment of this work. A further contribution is a comparative discussion of the design and evolution of three actual object‐oriented designs for the Message Passing Interface ( MPI‐1SF ) application programmer interface (API), two of which have influenced the standardization of C++ explicit parallel programming with MPI‐2, and which strongly indicate the value of a priori object‐oriented design and analysis of such APIs. Knowledge of design patterns is assumed herein. Discussion provided here includes systems developed at Mississippi State University (MPI++), the University of Notre Dame (OOMPI), and the merger of these systems that results in a standard binding within the MPI‐2 standard. Commentary concerning additional opportunities for further object‐oriented analysis and design of message passing systems and APIs, such as MPI‐2 and MPI/RT, are mentioned in conclusion. Connection of modern software design and engineering principles to high performance computing programming approaches is a new and important further contribution of this work. Copyright © 2001 John Wiley & Sons, Ltd. Anthony Skjellum, Diane G. Wooley, Michael Wolf, Purushotham V. Bangalore, Andrew Lumsdaine, Jeffrey M. Squyres, Brian C. McCandless |
Concurr. Comput. Pract. Exp. | 1 |
| 2001 | A Multithreaded Message Passing Interface (MPI) Architecture: Performance and Program Issues
Boris V. Protopopov, Anthony Skjellum |
J. Parallel Distributed Comput. | 2 |
| 2000 | Scaling the Data Mining Step in Knowledge Discovery Using Oceanographic Data
Bruce Wooley, Susan M. Bridges, Julia E. Hodges, Anthony Skjellum |
IEA/AIE | 4 |
| 2000 | MPJ: MPI-like message passing for JavaabstractRecently, there has been a lot of interest in using Java for parallel programming. Efforts have been hindered by lack of standard Java parallel programming APIs. To alleviate this problem, various groups started projects to develop Java message passing systems modelled on the successful Message Passing Interface (MPI). Official MPI bindings are currently defined only for C, Fortran, and C++, so early MPI-like environments for Java have been divergent. This paper relates an effort undertaken by a working group of the Java Grande Forum, seeking a consensus on an MPI-like API, to enhance the viability of parallel programming using Java. Copyright © 2000 John Wiley & Sons, Ltd. Bryan Carpenter, Vladimir Getov, Glenn Judd, Anthony Skjellum, Geoffrey C. Fox |
Concurr. Pract. Exp. | 4 |
| 2000 | Shared-memory communication approaches for an MPI message-passing libraryabstractThis paper discusses several approaches to designing and implementing shared-memory communication protocol modules for the message-passing interface (MPI) libraries, colloquially called ‘shared-memory devices’. The authors present a new taxonomy for classifying designs for shared-memory MPI communication devices and formulate design evaluation criteria. Using these criteria, the authors compare three existing shared-memory devices for MPICH and choose the best one. The authors also present experimental results that support their choice. The contributions of this paper are three-fold. First, the authors present the taxonomy for shared-memory communication devices. Second, they show advantages and potential problems of the devices that belong to different classes of their taxonomy using the formulated design criteria. Third, they analyze communication performance of existing MPICH shared-memory devices, discuss optimizations of their performance, and show the performance gains that these optimizations yield. MPICH is used for comparison, since it is a widely used MPI implementation. Copyright © 2000 John Wiley & Sons, Ltd. Boris V. Protopopov, Anthony Skjellum |
Concurr. Pract. Exp. | 2 |
| 1999 | Time Based Linux for Real-Time NOWs and MPI/RTabstractThe Real-Time Message Passing Interface (MPI/RT) is a communication layer middleware standard that is aimed at providing guaranteed Quality of Service for data transfers on high performance networks. It poses "middleout" requirements both on applications and on the operating system. In this paper, we consider the "middledown" issues by modifying a POSIX compliant operating system in order to support hard real time scheduling, a feature needed for efficient MPI/RT on real-time Networks of Workstations (NOWs). This paper describes the evolution of Turtle, a variant of RT-Linux, that will later support a prototype implementation of time-driven MPI/RT. Analysis, design approach and results for Turtle are discussed. Manoj Apte, Srigurunath Chakravarthi, Anand Pillai, Anthony Skjellum, Xin Yan Zan |
RTSS | 4 |
| 1997 | A poly-algorithm for parallel dense matrix multiplication on two-dimensional process grid topologiesabstractIn this paper, we present several new and generalized parallel dense matrix multiplication algorithms of the form C = αAB + β C on two-dimensional process grid topologies. These algorithms can deal with rectangular matrices distributed on rectangular grids. We classify these algorithms coherently into three categories according to the communication primitives used and thus we offer a taxonomy for this family of related algorithms. All these algorithms are represented in the data distribution independent approach and thus do not require a specific data distribution for correctness. The algorithmic compatibility condition result shown here ensures the correctness of the matrix multiplication. We define and extend the data distribution functions and introduce permutation compatibility and algorithmic compatibility. We also discuss a permutation compatible data distribution (modified virtual 2D data distribution). We conclude that no single algorithm always achieves the best performance on different matrix and grid shapes. A practical approach to resolve this dilemma is to use poly-algorithms. We analyze the characteristics of each of these matrix multiplication algorithms and provide initial heuristics for using the poly-algorithm. All these matrix multiplication algorithms have been tested on the IBM SP2 system. The experimental results are presented in order to demonstrate their relative performance characteristics, motivating the combined value of the taxonomy and new algorithms introduced here. © 1997 by John Wiley & Sons, Ltd. Anthony Skjellum, Robert D. Falgout |
Concurr. Pract. Exp. | 2 |
| 1996 | A High-Performance, Portable Implementation of the MPI Message Passing Interface Standard
William Gropp, Ewing L. Lusk, Nathan E. Doss, Anthony Skjellum |
Parallel Comput. | 4 |
| 1994 | The Design and Evolution of Zipcode
Anthony Skjellum, Steven G. Smith, Nathan E. Doss, Alvin P. Leung, Manfred Morari |
Parallel Comput. | 1 |
| 1993 | Scalable Libraries in a Heterogeneous EnvironmentabstractThis paper concerns itself with efforts to extend multicomputer libraries to a hierarchical, heterogeneous network environment. Two classes of support for such libraries are discussed: first, the message-passing features needed to establish groups of communicating processes, and communication contexts within which libraries can safely work. Second, it discusses message-passing primitives that encapsulate heterogeneity, hiding it from the user program (and library alike), and eliminating it when it proves unnecessary (within a homogeneous invocation, for instance). The multicomputer toolbox first-generation scalable libraries, and zipcode message-passing systems are the means by which the author demonstrates his research, so they are discussed. He relates zipcode syntax and semantics to the emerging MPI standard, when appropriate.> Anthony Skjellum |
HPDC | 1 |
| 1993 | Panel - Software Tools for High-Performance Distributed Computing
Vaidy S. Sunderam, Geoffrey C. Fox, Al Geist, William Gropp, Bob Harrison, Adam Kolawa, Michael J. Quinn, Anthony Skjellum |
HPDC | 8 |