Robert W. Wisniewski

dblp:01/6137 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Security and privacy of machine learning · 50% Authentication and access control · 50%
Computer architecture, parallel and distributed computing, and storage systems
10 papers
High-performance computing · 75% Parallel and multicore computing · 12% Cloud and datacenter computing · 6%
Software engineering, system software, and programming languages
10 papers
Operating systems · 58% Programming languages and type systems · 24% Software maintenance and evolution · 19%
Artificial intelligence
2 papers
Language models and text generation · 94% Motion planning and robot control · 4% Robot manipulation · 1%

Topics — the 23 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Security and privacy of machine learning
AI agent security
1.012026
InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure Agents · AAAI 2026
Authentication and access control › security policy
policy enforcement
1.012026
InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure Agents · AAAI 2026
Natural language and speech › Language models and text generation
large language model
0.312026
InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure Agents · AAAI 2026
High-performance computing
supercomputing
0.222010
Experiences with a Lightweight Supercomputer Kernel: Lessons Learned from Blue Gene's CNK · SC 2010
Providing a cloud network infrastructure on a supercomputer · HPDC 2010
Operating systems › kernel
kernel design
0.112010
Experiences with a Lightweight Supercomputer Kernel: Lessons Learned from Blue Gene's CNK · SC 2010
Operating systems › kernel
lightweight kernel
0.112010
Experiences with a Lightweight Supercomputer Kernel: Lessons Learned from Blue Gene's CNK · SC 2010
Programming languages and type systems › programming paradigms
generic programming
0.112009
Minimizing dependencies within generic classes for faster and smaller programs · OOPSLA 2009
Software maintenance and evolution › dynamic software updating
live patching
0.112007
Reboots Are for Hardware: Challenges and Solutions to Updating an Operating System on the Fly · USENIX ATC 2007
Cloud and datacenter computing
virtualization
0.122010
Providing a cloud network infrastructure on a supercomputer · HPDC 2010
K42: building a complete operating system · EuroSys 2006
Parallel and multicore computing
synchronization
0.031997
Scheduler-Conscious Synchronization · ACM Trans. Comput. Syst. 1997
High Performance Synchronization Algorithms for Multiprogrammed Multiprocessors · PPoPP 1995
Using Scheduler Information to Achieve Optimal Barrier Synchronization Performance · PPoPP 1993
Software maintenance and evolution › software evolution › software adaptation
dynamic reconfiguration
0.012003
System Support for Online Reconfiguration · USENIX ATC, General Track 2003
Operating systems › multiprocessing
multiprocessor operating system
0.012003
Efficient, Unified, and Scalable Performance Monitoring for Multiprocessor Operating Systems · SC 2003
Performance modeling and evaluation
workload characterization
0.012003
Efficient, Unified, and Scalable Performance Monitoring for Multiprocessor Operating Systems · SC 2003
Software maintenance and evolution
software updates
0.022007
Reboots Are for Hardware: Challenges and Solutions to Updating an Operating System on the Fly · USENIX ATC 2007
Providing Dynamic Update in an Operating System · USENIX ATC, General Track 2005
High-performance computing › supercomputing
exascale computing
0.012010
Experiences with a Lightweight Supercomputer Kernel: Lessons Learned from Blue Gene's CNK · SC 2010
Cloud and datacenter computing › virtualization
network virtualization
0.012010
Providing a cloud network infrastructure on a supercomputer · HPDC 2010
Parallel and multicore computing › synchronization
barrier synchronization
0.021997
Scheduler-Conscious Synchronization · ACM Trans. Comput. Syst. 1997
Using Scheduler Information to Achieve Optimal Barrier Synchronization Performance · PPoPP 1993
Parallel and multicore computing › synchronization
mutual exclusion lock
0.011997
Scheduler-Conscious Synchronization · ACM Trans. Comput. Syst. 1997
Robotics › Motion planning and robot control
robot control
0.011995
Adaptable Planner Primitives for Real-World Robotic Applications · IJCAI 1995
Reconfigurable computing and FPGAs
dynamic reconfiguration
0.012003
System Support for Online Reconfiguration · USENIX ATC, General Track 2003
Performance modeling and evaluation › parallel performance evaluation
multicore scalability
0.012003
Efficient, Unified, and Scalable Performance Monitoring for Multiprocessor Operating Systems · SC 2003
Operating systems › resource management › process management
CPU scheduling
0.011997
Scheduler-Conscious Synchronization · ACM Trans. Comput. Syst. 1997
Operating systems › resource management › process management
multiprogramming
0.011997
Scheduler-Conscious Synchronization · ACM Trans. Comput. Syst. 1997

Methods — techniques the papers use, named apart from their topics

policy reasoning · 2.0large language model · 2.0linear programming · 0.8LogGPS model · 0.8system measurement · 0.2experience report · 0.2selective partitioning and replication · 0.1object-oriented structuring · 0.1object-oriented kernel design · 0.1direct hardware access · 0.1performance instrumentation · 0.1performance evaluation · 0.0multiprogramming-aware scheduling · 0.0
YearPublicationVenuePosition
2026 InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure Agents
abstract
The proliferation of Model Context Protocol (MCP) servers in enterprise infrastructure management has revolutionized AI-driven automation while introducing critical multi-layered security vulnerabilities that traditional cybersecurity frameworks cannot adequately address. This paper presents a comprehensive intelligent guardrail system that addresses the unique security challenges of MCP-driven infrastructure management through a novel four-layer defense architecture. Our solution employs a dedicated guardian LLM that interprets natural language policies and applies contextual reasoning to complex infrastructure scenarios, providing dynamic policy enforcement that adapts to user roles, operational timing, and system context. Unlike existing rule-based security systems, our approach implements guardrails at four distinct control points: input message filtering, tool selection validation, execution-time verification, and post-action auditing. The system addresses critical gaps in existing security solutions by providing infrastructure-specific threat modeling, real-time policy adaptation, and comprehensive audit trails with explainable decision-making through confidence scores and detailed reasoning. Our evaluation demonstrates the system's effectiveness in preventing command injection, privilege escalation, and tool poisoning attacks across various enterprise infrastructure scenarios while maintaining operational agility essential for modern data center management.
Aalap Tripathy, Gayathri Saranathan, Martin Foltin, Suparna Bhattacharya, Scott Hinchley, Donald M. Bahls, David Brookshire, Larry Kaplan, Robert W. Wisniewski
AAAI10
2025 EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
abstract
Resource disaggregation is a promising technique for improving the efficiency of large-scale computing systems.However, this comes at the cost of increased memory access latency due to the need to rely on the network fabric to transfer data between remote nodes.As such, it is crucial to ascertain an application's memory latency sensitivity to minimize the overall performance impact.Existing tools for measuring memory latency sensitivity often rely on custom ad-hoc hardware or cycle-accurate simulators, which can be inflexible and time-consuming.To address this, we present EDAN (Execution DAG Analyzer), a novel performance analysis tool that leverages an application's runtime instruction trace to generate its corresponding execution DAG.This approach allows us to estimate the latency sensitivity of sequential programs and investigate the impact of different hardware configurations.EDAN not only provides us with the capability of calculating the theoretical bounds for performance metrics, but it also helps us gain insight into the memorylevel parallelism inherent to HPC applications.We apply
Mikhail Khalilov, Lukas Gianinazzi, Timo Schneider, Marcin Chrapek, Jai Dayal, Manisha Gajbe, Robert W. Wisniewski, Torsten Hoefler
ICS8
2024 LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
abstract
The shift towards high-bandwidth networks driven by AI workloads in data centers and HPC clusters has unintentionally aggravated network latency, adversely affecting the performance of communication-intensive HPC applications. As large-scale applications often exhibit significant differences in their network latency tolerance, it is crucial to determine the extent of network latency an application can withstand without significant performance degradation. Current approaches often rely on specialized hardware or simulators, which can be inflexible and time-consuming. We introduce LLAMP, a novel toolchain that offers an efficient analytical approach to evaluating HPC applications’ network latency tolerance using the LogGPS model and linear programming. Through our validation on a variety of MPI applications such as LULESH and MILC, we demonstrate our tool’s high accuracy, with relative prediction errors below 2%. Additionally, we include a case study of the ICON weather and climate model to illustrate LLAMP’s broad applicability in evaluating collective algorithms and network topologies.
Langwen Huang, Marcin Chrapek, Timo Schneider, Jai Dayal, Manisha Gajbe, Robert W. Wisniewski, Torsten Hoefler
SC7
2018 Performance and Scalability of Lightweight Multi-kernel Based Operating Systems
abstract
Multi-kernels leverage today's multi-core chips to run multiple operating system (OS) kernels, typically a Light Weight Kernel (LWK) and a Linux kernel, simultaneously. The LWK provides high performance and scalability, while the Linux kernel provides compatibility. Multi-kernels show the promise of being able to meet tomorrow's extreme-scale computing needs while providing strong isolation, yielding high performance and scalability needed by classical HPC applications. McKernel and mOS started as independent research initiatives to explore the above potential. Previous work described their design and architecture advantages. This paper deploys the two LWKs and presents results from running them on a 2,048-node system with Intel Xeon Phi processors (KNL) connected by Intel Omni-Path Fabric. We compare the performance of McKernel, mOS, and Linux. Although the two multi-kernel efforts approached the problem from different angles, the results show a median performance improvement of 9% with some applications as high as 280% validating the efficacy of the multi-kernel approach. We provide insight into the performance improvements and discuss the strengths of the two different multi-kernel approaches.
Balazs Gerofi, Rolf Riesen, Masamichi Takagi, Taisuke Boku, Kengo Nakajima, Yutaka Ishikawa, Robert W. Wisniewski
IPDPS7
2012 Evaluating the Impact of TLB Misses on Future HPC Systems
abstract
TLB misses have been considered an important source of system overhead and one of the causes that limit scalability on large supercomputers. This assumption lead to HPC lightweight kernel designs that usually statically map page table entries to TLB entries and do not take TLB misses. While this approach worked for petascale clusters, programming and debugging exascale applications composed of billions of threads is not a trivial task and users have started to explore novel programming models and tools, which require a richer system software support. In this study we present a quantitative analysis of the effect of TLB misses on current and future parallel applications at scale. To provide a fair evaluation, we compare a noiseless OS (CNK) with a custom version of the same OS capable of handling TLB misses on a BG/P system (up to 4096 cores). Our methodology follows a two-step approach: we first analyze the effects of TLB misses with a low-overhead, range-checking TLB miss handler, and then simulate a more complex TLB management system through TLB noise injection. We analyze the system behavior with different page sizes and increasing number of nodes and perform a sensitivity analysis. Our results show that the overhead introduced by TLB misses on complex HPC applications from the LLNL and ANL benchmarks is below 2% if the TLB pressure is contained and/or the TLB miss handler overhead is low, even with 1MB-pages and under large TLB noise injection. These results open the possibility of implementing richer OS memory management services to satisfy the requirements of future applications and users.
Alessandro Morari, Roberto Gioiosa, Robert W. Wisniewski, Bryan S. Rosenburg, Todd Inglett, Mateo Valero
IPDPS3
2012 FusedOS: Fusing LWK Performance with FWK Functionality in a Heterogeneous Environment
abstract
Traditionally, there have been two approaches to providing an operating environment for high performance computing (HPC). A Full-Weight Kernel(FWK) approach starts with a general-purpose operating system and strips it down to better scale up across more cores and out across larger clusters. A Light-Weight Kernel (LWK) approach starts with a new thin kernel code base and extends its functionality by adding more system services needed by applications. In both cases, the goal is to provide end-users with a scalable HPC operating environment with the functionality and services needed to reliably run their applications. To achieve this goal, we propose a new approach, called Fused OS, that combines the FWK and LWK approaches. Fused OS provides an infrastructure capable of partitioning the resources of a multicoreheterogeneous system and collaboratively running different operating environments on subsets of the cores and memory, without the use of a virtual machine monitor. With Fused OS, HPC applications can enjoy both the performance characteristics of an LWK and the rich functionality of an FWK through cross-core system service delegation. This paper presents the Fused OS architecture and a prototype implementation on Blue Gene/Q. The Fused OS prototype leverages Linux with small modifications as a FWK and implements a user-level LWK called Compute Library (CL) by leveraging CNK. We present CL performance results demonstrating low noise and show micro-benchmarks running with performance commensurate with that provided by CNK.
Yoonho Park, Eric Van Hensbergen, Marius Hillenbrand, Todd Inglett, Bryan S. Rosenburg, Kyung Dong Ryu, Robert W. Wisniewski
SBAC-PAD7
2011 A Quantitative Analysis of OS Noise
abstract
Operating system noise is a well-known problem that may limit application scalability on large-scale machines, significantly reducing their performance. Though the problem is well studied, much of the previous work has been qualitative. We have developed a technique to provide a quantitative descriptive analysis for each OS event that contributes to OS noise. The mechanism allows us to detail all sources of OS noise through precise kernel instrumentation and provides frequency and duration analysis for each event. Such a description gives OS developers better guidance for reducing OS noise. We integrated this data with a trace visualizer allowing quicker and more intuitive understanding of the data. Specifically, the contributions of this paper are three-fold. First, we describe a methodology whereby detailed quantitative information may be obtained for each OS noise event. Though not the thrust of the paper, we show how we implemented that methodology by augmenting LTTng. We validate our approach by comparing it to other well-known standard techniques to analyze OS noise. Second, we provide a case study in which we use our methodology to analyze the OS noise when running benchmarks from the LLNL Sequoia applications. Our experiments enrich and expand previous results with our quantitative characterization. Third, we describe how a detailed characterization permits to disambiguate noise signatures of qualitatively similar events, allowing developers to address the true cause of each noise event.
Alessandro Morari, Roberto Gioiosa, Robert W. Wisniewski, Francisco J. Cazorla, Mateo Valero
IPDPS3
2010 Providing a cloud network infrastructure on a supercomputer
abstract
Supercomputers and clouds both strive to make a large number of computing cores available for computation. More recently, similar objectives such as low-power, manageability at scale, and low cost of ownership are driving a more converged hardware and software. Challenges remain, however, of which one is that current cloud infrastructure does not yield the performance sought by many scientific applications. A source of the performance loss comes from virtualization and virtualization of the network in particular. This paper provides an introduction and analysis of a hybrid supercomputer software infrastructure, which allows direct hardware access to the communication hardware for the necessary components while providing the standard elastic cloud infrastructure for other components.
Jonathan Appavoo, Amos Waterland, Dilma Da Silva, Volkmar Uhlig, Bryan S. Rosenburg, Eric Van Hensbergen, Jan Stoess, Robert W. Wisniewski, Udo Steinberg
HPDC8
2010 Experiences with a Lightweight Supercomputer Kernel: Lessons Learned from Blue Gene's CNK
abstract
The Petascale era has recently been ushered in and many researchers have already turned their attention to the challenges of exascale computing. To achieve petascale computing two broad approaches for kernels were taken, a lightweight approach embodied by IBM Blue Gene's CNK, and a more fullweight approach embodied by Cray's CNL. There are strengths and weaknesses to each approach. Examining the current generation can provide insight as to what mechanisms may be needed for the exascale generation. The contributions of this paper are the experiences we had with CNK on Blue Gene/P. We demonstrate it is possible to implement a small lightweight kernel that scales well but still provides a Linux environment and functionality desired by HPC programmers. Such an approach provides the values of reproducibility, low noise, high and stable performance, reliability, and ease of effectively exploiting unique hardware features. We describe the strengths and weaknesses of this approach.
Mark Giampapa, Thomas Gooding, Todd Inglett, Robert W. Wisniewski
SC4
2009 Minimizing dependencies within generic classes for faster and smaller programs
abstract
Generic classes can be used to improve performance by allowing compile-time polymorphism. But the applicability of compile-time polymorphism is narrower than that of runtime polymorphism, and it might bloat the object code. We advocate a programming principle whereby a generic class should be implemented in a way that minimizes the dependencies between its members (nested types, methods) and its generic type parameters. Conforming to this principle (1) reduces the bloat and (2) gives rise to a previously unconceived manner of using the language that expands the applicability of compile-time polymorphism to a wider range of problems. Our contribution is thus a programming technique that generates faster and smaller programs. We apply our ideas to GCC's STL containers and iterators, and we demonstrate notable speedups and reduction in object code size (real application runs 1.2x to 2.1x faster and STL code is 1x to 25x smaller). We conclude that standard generic APIs (like STL) should be amended to reflect the proposed principle in the interest of efficiency and compactness. Such modifications will not break old code, simply increase flexibility. Our findings apply to languages like C++, C#, and D, which realize generic programming through multiple instantiations.
Dan Tsafrir, Robert W. Wisniewski, David F. Bacon, Bjarne Stroustrup
OOPSLA2
2007 Experiences Understanding Performance in a Commercial Scale-Out Environment
Robert W. Wisniewski, Mathieu Desnoyers, Maged M. Michael, José E. Moreira, Doron Shiloach, Livio B. Soares
Euro-Par1
2007 Scale-up x Scale-out: A Case Study using Nutch/Lucene
abstract
Scale-up solutions in the form of large SMPs have represented the mainstream of commercial computing for the past several years. The major server vendors continue to provide increasingly larger and more powerful machines. More recently, scale-out solutions, in the form of clusters of smaller machines, have gained increased acceptance for commercial computing. Scale-out solutions are particularly effective in high-throughput Web-centric applications. In this paper, we investigate the behavior of two competing approaches to parallelism, scale-up and scale-out, in an emerging search application. Our conclusions show that a scale-out strategy can be the key to good performance even on a scale-up machine. Furthermore, scale-out solutions offer better price/performance, although at an increase in management complexity.
Maged M. Michael, José E. Moreira, Doron Shiloach, Robert W. Wisniewski
IPDPS4
2007 Reboots Are for Hardware: Challenges and Solutions to Updating an Operating System on the Fly
Andrew Baumann, Jonathan Appavoo, Robert W. Wisniewski, Dilma Da Silva, Orran Krieger, Gernot Heiser
USENIX ATC3
2007 Libra: a library operating system for a jvm in a virtualized execution environment
abstract
If the operating system could be specialized for every application, many applications would run faster. For example, Java virtual machines (JVMs) provide their own threading model and memory protection, so general-purpose operating system implementations of these abstractions are redundant. However, traditional means of transforming existing systems into specialized systems are difficult to adopt because they require replacing the entire operating system. This paper describes Libra, an execution environment specialized for IBM's J9 JVM. Libra does not replace the entire operating system. Instead, Libra and J9 form a single statically-linked image that runs in a hypervisor partition. Libra provides the services necessary to achieve good performance for the Java workloads of interest but relies on an instance of Linux in another hypervisor partition to provide a networking stack, a filesystem, and other services. The expense of remote calls is offset by the fact that Libra's services can be customized for a particular workload; for example, on the Nutch search engine, we show that two simple customizations improve application throughput by a factor of 2.7.
Glenn Ammons, Jonathan Appavoo, Maria A. Butrico, Dilma Da Silva, David Grove, Kiyokuni Kawachiya, Orran Krieger, Bryan S. Rosenburg, Eric Van Hensbergen, Robert W. Wisniewski
VEE10
2007 Experience distributing objects in an SMMP OS
abstract
Designing and implementing system software so that it scales well on shared-memory multiprocessors (SMMPs) has proven to be surprisingly challenging. To improve scalability, most designers to date have focused on concurrency by iteratively eliminating the need for locks and reducing lock contention. However, our experience indicates that locality is just as, if not more, important and that focusing on locality ultimately leads to a more scalable system. In this paper, we describe a methodology and a framework for constructing system software structured for locality, exploiting techniques similar to those used in distributed systems. Specifically, we found two techniques to be effective in improving scalability of SMMP operating systems: (i) an object-oriented structure that minimizes sharing by providing a natural mapping from independent requests to independent code paths and data structures, and (ii) the selective partitioning, distribution, and replication of object implementations in order to improve locality. We describe concrete examples of distributed objects and our experience implementing them. We demonstrate that the distributed implementations improve the scalability of operating-system-intensive parallel workloads.
Jonathan Appavoo, Dilma Da Silva, Orran Krieger, Marc A. Auslander, Michal Ostrowski, Bryan S. Rosenburg, Amos Waterland, Robert W. Wisniewski, Jimi Xenidis, Michael Stumm, Livio B. Soares
ACM Trans. Comput. Syst.8
2006 K42: building a complete operating system
abstract
K42 is one of the few recent research projects that is examining operating system design structure issues in the context of new whole-system design. K42 is open source and was designed from the ground up to perform well and to be scalable, customizable, and maintainable. The project was begun in 1996 by a team at IBM Research. Over the last nine years there has been a development effort on K42 from between six to twenty researchers and developers across IBM, collaborating universities, and national laboratories. K42 supports the Linux API and ABI, and is able to run unmodified Linux applications and libraries. The approach we took in K42 to achieve scalability and customizability has been successful.The project has produced positive research results, has resulted in contributions to Linux and the Xen hypervisor on Power, and continues to be a rich platform for exploring system software technology. Today, K42, is one of the key exploratory platforms in the DOE's FAST-OS program, is being used as a prototyping vehicle in IBM's PERCS project, and is being used by universities and national labs for exploratory research. In this paper, we provide insight into building an entire system by discussing the motivation and history of K42, describing its fundamental technologies, and presenting an overview of the research directions we have been pursuing.
Orran Krieger, Marc A. Auslander, Bryan S. Rosenburg, Robert W. Wisniewski, Jimi Xenidis, Dilma Da Silva, Michal Ostrowski, Jonathan Appavoo, Maria A. Butrico, Mark F. Mergen, Amos Waterland, Volkmar Uhlig
EuroSys4
2005 Online performance analysis by statistical sampling of microprocessor performance counters
abstract
Hardware performance counters (HPCs) are increasingly being used to analyze performance and identify the causes of performance bottlenecks. However, HPCs are difficult to use for several reasons. Microprocessors do not provide enough counters to simultaneously monitor the many different types of events needed to form an over-all understanding of performance. Moreover, HPCs primarily count low-level micro-architectural events from which it is difficult to extract high-level insight required for identifying causes of performance problems.We describe two techniques that help overcome these difficulties, allowing HPCs to be used in dynamic real-time optimizers. First, statistical sampling is used to dynamically multiplex HPCs and make a larger set of logical HPCs available. Using real programs, we show experimentally that it is possible through this sampling to obtain counts of hardware events that are statistically similar (within 15%) to complete non-sampled counts, thus allowing us to provide a much larger set of logical HPCs. Second, we observe that stall cycles are a primary source of inefficiencies, and hence they should be major targets for software optimization. Based on this observation, we build a simple model in real-time that speculatively associates each stall cycle to a processor component that likely caused the stall. The information needed to produce this model is obtained using our HPC multiplexing facility to monitor a large number of hardware components simultaneously. Our analysis shows that even in an out-of-order superscalar micro-processor such a speculative approach yields a fairly accurate model with run-time overhead for collection and computation of under 2%.These results demonstrate that we can effective analyze on-line performance of application and system code running at full speed. The stall analysis shows where performance is being lost on a given processor.
Michael Stumm, Robert W. Wisniewski
ICS3
2005 Providing Dynamic Update in an Operating System
Andrew Baumann, Gernot Heiser, Jonathan Appavoo, Dilma Da Silva, Orran Krieger, Robert W. Wisniewski, Jeremy Kerr
USENIX ATC, General Track6
2003 Efficient, Unified, and Scalable Performance Monitoring for Multiprocessor Operating Systems
abstract
Programming, understanding, and tuning the performance of large multiprocessor systems is challenging. Experts have difficulty achieving good utilization for applications on large machines. The task of implementing scalable systems such as an operating system or database on large machines is even more challenging. And the importance of achieving good performance on multiprocessor machines is increasing as the number of cores per chip increases and as the size of multiprocessors increases. Crucial to achieving good performance is being able to understand the behavior of the system.
Robert W. Wisniewski, Bryan S. Rosenburg
SC1
2003 System Support for Online Reconfiguration
Craig A. N. Soules, Jonathan Appavoo, Kevin Hui, Robert W. Wisniewski, Dilma Da Silva, Gregory R. Ganger, Orran Krieger, Michael Stumm, Marc A. Auslander, Michal Ostrowski, Bryan S. Rosenburg, Jimi Xenidis
USENIX ATC, General Track4
2001 Supporting Hot-Swappable Components for System Software
abstract
Summary form only given. A hot-swappable component is one that can be replaced with a new or different implementation while the system is running and actively using the component. For example, a component of a TCP/IP protocol stack, when hot-swappable, can be replaced (perhaps to handle new denial-of-service attacks or improve performance), without disturbing existing network connections. The capability to swap components offers a number of potential advantages such as: online upgrades for high availability systems, improved performance due to dynamic adaptability and simplified software structures by allowing distinct policy and implementation options to be implemented in separate components (rather than as a single monolithic component) and dynamically swapped as needed. In order to hot-swap a component, it is necessary to (i) instantiate a replacement component; (ii) establish a quiescent state in which the component is temporarily idle; (iii) transfer state from the old component to the new component; (iv) swap the new component for the old; and (v) deallocate the old component.
Kevin Hui, Jonathan Appavoo, Robert W. Wisniewski, Marc A. Auslander, David Edelsohn, Benjamin Gamsa, Orran Krieger, Bryan S. Rosenburg, Michael Stumm
HotOS3
1997 Scheduler-Conscious Synchronization
abstract
Efficient synchronization is important for achieving good performance in parallel programs, especially on large-scale multiprocessors. Most synchronization algorithms have been designed to run on a dedicated machine, with one application process per processor, and can suffer serious performance degradation in the presence of multiprogramming. Problems arise when running processes block or, worse, busy-wait for action on the part of a process that the scheduler has chosen not to run. We show that these problems are particularly severe for scalable synchronization algorithms based on distributed data structures. We then describe and evaluate a set of algorithms that perform well in the presence of multiprogramming while maintaining good performance on dedicated machines. We consider both large and small machines, with a particular focus on scalability, and examine mutual-exclusion locks, reader-writer locks, and barriers. Our algorithms vary in the degree of support required from the kernel scheduler. We find that while it is possible to avoid pathological performance problems using previously proposed kernel mechanisms, a modest additional widening of the kernel/user interface can make scheduler-conscious synchronization algorithms significantly simpler and faster, with performance on dedicated machines comparable to that of scheduler-oblivious algorithms.
Leonidas I. Kontothanassis, Robert W. Wisniewski, Michael L. Scott
ACM Trans. Comput. Syst.2
1995 Adaptable Planner Primitives for Real-World Robotic Applications
Robert W. Wisniewski
IJCAI1
1995 High Performance Synchronization Algorithms for Multiprogrammed Multiprocessors
abstract
Scalable busy-wait synchronization algorithms are essential for achieving good parallel program performance on large scale multiprocessors. Such algorithms include mutual exclusion locks, reader-writer locks, and barrier synchronization. Unfortunately, scalable synchronization algorithms are particularly sensitive to the effects of multiprogramming: their performance degrades sharply when processors are shared among different applications, or even among processes of the same application. In this paper we describe the design and evaluation of scalable scheduler-conscious mutual exclusion locks, reader-writer locks, and barriers, and show that by sharing information across the kernel/application interface we can improve the performance of scheduler-oblivious implementations by more than an order of magnitude.
Robert W. Wisniewski, Leonidas I. Kontothanassis, Michael L. Scott
PPoPP1
1993 Using Scheduler Information to Achieve Optimal Barrier Synchronization Performance
abstract
Parallel programs frequently use barriers to synchronize successive steps in an algorithm. In the presence of multiprogramming the choice of spinning versus blocking barriers can have a significant impact on performance. We demonstrate how competitive spinning techniques previously designed for locks can be extended to barriers, and we evaluate their performance. We design an additional competitive spinning technique that adapts more quickly in a dynamic environment. We then propose and evaluate a new method that obtains better peformance than previous techniques by using scheduler information to decide between spinning and blocking. The scheduler information technique makes optimal choices incurring little overhead.
Leonidas I. Kontothanassis, Robert W. Wisniewski
PPoPP2