VLDB 2026 Research / reviewers in the wild / expert
Robert W. Wisniewski
dblp:01/6137
· DBLP profile ↗
25ranked-venue papers
4as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
1 paper |
Security and privacy of machine learning · 50% Authentication and access control · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
High-performance computing · 75% Parallel and multicore computing · 12% Cloud and datacenter computing · 6% | |
| Software engineering, system software, and programming languages
10 papers |
Operating systems · 58% Programming languages and type systems · 24% Software maintenance and evolution · 19% | |
| Artificial intelligence
2 papers |
Language models and text generation · 94% Motion planning and robot control · 4% Robot manipulation · 1% |
Topics — the 23 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Security and privacy of machine learning
AI agent security |
1.0 | 1 | 2026 | InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure Agents · AAAI 2026 |
Authentication and access control › security policy
policy enforcement |
1.0 | 1 | 2026 | InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure Agents · AAAI 2026 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2026 | InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure Agents · AAAI 2026 |
High-performance computing
supercomputing |
0.2 | 2 | 2010 | Experiences with a Lightweight Supercomputer Kernel: Lessons Learned from Blue Gene's CNK · SC 2010 Providing a cloud network infrastructure on a supercomputer · HPDC 2010 |
Operating systems › kernel
kernel design |
0.1 | 1 | 2010 | Experiences with a Lightweight Supercomputer Kernel: Lessons Learned from Blue Gene's CNK · SC 2010 |
Operating systems › kernel
lightweight kernel |
0.1 | 1 | 2010 | Experiences with a Lightweight Supercomputer Kernel: Lessons Learned from Blue Gene's CNK · SC 2010 |
Programming languages and type systems › programming paradigms
generic programming |
0.1 | 1 | 2009 | Minimizing dependencies within generic classes for faster and smaller programs · OOPSLA 2009 |
Software maintenance and evolution › dynamic software updating
live patching |
0.1 | 1 | 2007 | Reboots Are for Hardware: Challenges and Solutions to Updating an Operating System on the Fly · USENIX ATC 2007 |
Cloud and datacenter computing
virtualization |
0.1 | 2 | 2010 | Providing a cloud network infrastructure on a supercomputer · HPDC 2010 K42: building a complete operating system · EuroSys 2006 |
Parallel and multicore computing
synchronization |
0.0 | 3 | 1997 | Scheduler-Conscious Synchronization · ACM Trans. Comput. Syst. 1997 High Performance Synchronization Algorithms for Multiprogrammed Multiprocessors · PPoPP 1995 Using Scheduler Information to Achieve Optimal Barrier Synchronization Performance · PPoPP 1993 |
Software maintenance and evolution › software evolution › software adaptation
dynamic reconfiguration |
0.0 | 1 | 2003 | System Support for Online Reconfiguration · USENIX ATC, General Track 2003 |
Operating systems › multiprocessing
multiprocessor operating system |
0.0 | 1 | 2003 | Efficient, Unified, and Scalable Performance Monitoring for Multiprocessor Operating Systems · SC 2003 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2003 | Efficient, Unified, and Scalable Performance Monitoring for Multiprocessor Operating Systems · SC 2003 |
Software maintenance and evolution
software updates |
0.0 | 2 | 2007 | Reboots Are for Hardware: Challenges and Solutions to Updating an Operating System on the Fly · USENIX ATC 2007 Providing Dynamic Update in an Operating System · USENIX ATC, General Track 2005 |
High-performance computing › supercomputing
exascale computing |
0.0 | 1 | 2010 | Experiences with a Lightweight Supercomputer Kernel: Lessons Learned from Blue Gene's CNK · SC 2010 |
Cloud and datacenter computing › virtualization
network virtualization |
0.0 | 1 | 2010 | Providing a cloud network infrastructure on a supercomputer · HPDC 2010 |
Parallel and multicore computing › synchronization
barrier synchronization |
0.0 | 2 | 1997 | Scheduler-Conscious Synchronization · ACM Trans. Comput. Syst. 1997 Using Scheduler Information to Achieve Optimal Barrier Synchronization Performance · PPoPP 1993 |
Parallel and multicore computing › synchronization
mutual exclusion lock |
0.0 | 1 | 1997 | Scheduler-Conscious Synchronization · ACM Trans. Comput. Syst. 1997 |
Robotics › Motion planning and robot control
robot control |
0.0 | 1 | 1995 | Adaptable Planner Primitives for Real-World Robotic Applications · IJCAI 1995 |
Reconfigurable computing and FPGAs
dynamic reconfiguration |
0.0 | 1 | 2003 | System Support for Online Reconfiguration · USENIX ATC, General Track 2003 |
Performance modeling and evaluation › parallel performance evaluation
multicore scalability |
0.0 | 1 | 2003 | Efficient, Unified, and Scalable Performance Monitoring for Multiprocessor Operating Systems · SC 2003 |
Operating systems › resource management › process management
CPU scheduling |
0.0 | 1 | 1997 | Scheduler-Conscious Synchronization · ACM Trans. Comput. Syst. 1997 |
Operating systems › resource management › process management
multiprogramming |
0.0 | 1 | 1997 | Scheduler-Conscious Synchronization · ACM Trans. Comput. Syst. 1997 |
Methods — techniques the papers use, named apart from their topics
policy reasoning · 2.0large language model · 2.0linear programming · 0.8LogGPS model · 0.8system measurement · 0.2experience report · 0.2selective partitioning and replication · 0.1object-oriented structuring · 0.1object-oriented kernel design · 0.1direct hardware access · 0.1performance instrumentation · 0.1performance evaluation · 0.0multiprogramming-aware scheduling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure AgentsabstractThe proliferation of Model Context Protocol (MCP) servers in enterprise infrastructure management has revolutionized AI-driven automation while introducing critical multi-layered security vulnerabilities that traditional cybersecurity frameworks cannot adequately address. This paper presents a comprehensive intelligent guardrail system that addresses the unique security challenges of MCP-driven infrastructure management through a novel four-layer defense architecture. Our solution employs a dedicated guardian LLM that interprets natural language policies and applies contextual reasoning to complex infrastructure scenarios, providing dynamic policy enforcement that adapts to user roles, operational timing, and system context. Unlike existing rule-based security systems, our approach implements guardrails at four distinct control points: input message filtering, tool selection validation, execution-time verification, and post-action auditing. The system addresses critical gaps in existing security solutions by providing infrastructure-specific threat modeling, real-time policy adaptation, and comprehensive audit trails with explainable decision-making through confidence scores and detailed reasoning. Our evaluation demonstrates the system's effectiveness in preventing command injection, privilege escalation, and tool poisoning attacks across various enterprise infrastructure scenarios while maintaining operational agility essential for modern data center management. Aalap Tripathy, Gayathri Saranathan, Martin Foltin, Suparna Bhattacharya, Scott Hinchley, Donald M. Bahls, David Brookshire, Larry Kaplan, Robert W. Wisniewski |
AAAI | 10 |
| 2025 | EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPCabstractResource disaggregation is a promising technique for improving the efficiency of large-scale computing systems.However, this comes at the cost of increased memory access latency due to the need to rely on the network fabric to transfer data between remote nodes.As such, it is crucial to ascertain an application's memory latency sensitivity to minimize the overall performance impact.Existing tools for measuring memory latency sensitivity often rely on custom ad-hoc hardware or cycle-accurate simulators, which can be inflexible and time-consuming.To address this, we present EDAN (Execution DAG Analyzer), a novel performance analysis tool that leverages an application's runtime instruction trace to generate its corresponding execution DAG.This approach allows us to estimate the latency sensitivity of sequential programs and investigate the impact of different hardware configurations.EDAN not only provides us with the capability of calculating the theoretical bounds for performance metrics, but it also helps us gain insight into the memorylevel parallelism inherent to HPC applications.We apply Mikhail Khalilov, Lukas Gianinazzi, Timo Schneider, Marcin Chrapek, Jai Dayal, Manisha Gajbe, Robert W. Wisniewski, Torsten Hoefler |
ICS | 8 |
| 2024 | LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear ProgrammingabstractThe shift towards high-bandwidth networks driven by AI workloads in data centers and HPC clusters has unintentionally aggravated network latency, adversely affecting the performance of communication-intensive HPC applications. As large-scale applications often exhibit significant differences in their network latency tolerance, it is crucial to determine the extent of network latency an application can withstand without significant performance degradation. Current approaches often rely on specialized hardware or simulators, which can be inflexible and time-consuming. We introduce LLAMP, a novel toolchain that offers an efficient analytical approach to evaluating HPC applications’ network latency tolerance using the LogGPS model and linear programming. Through our validation on a variety of MPI applications such as LULESH and MILC, we demonstrate our tool’s high accuracy, with relative prediction errors below 2%. Additionally, we include a case study of the ICON weather and climate model to illustrate LLAMP’s broad applicability in evaluating collective algorithms and network topologies. Langwen Huang, Marcin Chrapek, Timo Schneider, Jai Dayal, Manisha Gajbe, Robert W. Wisniewski, Torsten Hoefler |
SC | 7 |
| 2018 | Performance and Scalability of Lightweight Multi-kernel Based Operating SystemsabstractMulti-kernels leverage today's multi-core chips to run multiple operating system (OS) kernels, typically a Light Weight Kernel (LWK) and a Linux kernel, simultaneously. The LWK provides high performance and scalability, while the Linux kernel provides compatibility. Multi-kernels show the promise of being able to meet tomorrow's extreme-scale computing needs while providing strong isolation, yielding high performance and scalability needed by classical HPC applications. McKernel and mOS started as independent research initiatives to explore the above potential. Previous work described their design and architecture advantages. This paper deploys the two LWKs and presents results from running them on a 2,048-node system with Intel Xeon Phi processors (KNL) connected by Intel Omni-Path Fabric. We compare the performance of McKernel, mOS, and Linux. Although the two multi-kernel efforts approached the problem from different angles, the results show a median performance improvement of 9% with some applications as high as 280% validating the efficacy of the multi-kernel approach. We provide insight into the performance improvements and discuss the strengths of the two different multi-kernel approaches. Balazs Gerofi, Rolf Riesen, Masamichi Takagi, Taisuke Boku, Kengo Nakajima, Yutaka Ishikawa, Robert W. Wisniewski |
IPDPS | 7 |
| 2012 | Evaluating the Impact of TLB Misses on Future HPC SystemsabstractTLB misses have been considered an important source of system overhead and one of the causes that limit scalability on large supercomputers. This assumption lead to HPC lightweight kernel designs that usually statically map page table entries to TLB entries and do not take TLB misses. While this approach worked for petascale clusters, programming and debugging exascale applications composed of billions of threads is not a trivial task and users have started to explore novel programming models and tools, which require a richer system software support. In this study we present a quantitative analysis of the effect of TLB misses on current and future parallel applications at scale. To provide a fair evaluation, we compare a noiseless OS (CNK) with a custom version of the same OS capable of handling TLB misses on a BG/P system (up to 4096 cores). Our methodology follows a two-step approach: we first analyze the effects of TLB misses with a low-overhead, range-checking TLB miss handler, and then simulate a more complex TLB management system through TLB noise injection. We analyze the system behavior with different page sizes and increasing number of nodes and perform a sensitivity analysis. Our results show that the overhead introduced by TLB misses on complex HPC applications from the LLNL and ANL benchmarks is below 2% if the TLB pressure is contained and/or the TLB miss handler overhead is low, even with 1MB-pages and under large TLB noise injection. These results open the possibility of implementing richer OS memory management services to satisfy the requirements of future applications and users. Alessandro Morari, Roberto Gioiosa, Robert W. Wisniewski, Bryan S. Rosenburg, Todd Inglett, Mateo Valero |
IPDPS | 3 |
| 2012 | FusedOS: Fusing LWK Performance with FWK Functionality in a Heterogeneous EnvironmentabstractTraditionally, there have been two approaches to providing an operating environment for high performance computing (HPC). A Full-Weight Kernel(FWK) approach starts with a general-purpose operating system and strips it down to better scale up across more cores and out across larger clusters. A Light-Weight Kernel (LWK) approach starts with a new thin kernel code base and extends its functionality by adding more system services needed by applications. In both cases, the goal is to provide end-users with a scalable HPC operating environment with the functionality and services needed to reliably run their applications. To achieve this goal, we propose a new approach, called Fused OS, that combines the FWK and LWK approaches. Fused OS provides an infrastructure capable of partitioning the resources of a multicoreheterogeneous system and collaboratively running different operating environments on subsets of the cores and memory, without the use of a virtual machine monitor. With Fused OS, HPC applications can enjoy both the performance characteristics of an LWK and the rich functionality of an FWK through cross-core system service delegation. This paper presents the Fused OS architecture and a prototype implementation on Blue Gene/Q. The Fused OS prototype leverages Linux with small modifications as a FWK and implements a user-level LWK called Compute Library (CL) by leveraging CNK. We present CL performance results demonstrating low noise and show micro-benchmarks running with performance commensurate with that provided by CNK. Yoonho Park, Eric Van Hensbergen, Marius Hillenbrand, Todd Inglett, Bryan S. Rosenburg, Kyung Dong Ryu, Robert W. Wisniewski |
SBAC-PAD | 7 |
| 2011 | A Quantitative Analysis of OS NoiseabstractOperating system noise is a well-known problem that may limit application scalability on large-scale machines, significantly reducing their performance. Though the problem is well studied, much of the previous work has been qualitative. We have developed a technique to provide a quantitative descriptive analysis for each OS event that contributes to OS noise. The mechanism allows us to detail all sources of OS noise through precise kernel instrumentation and provides frequency and duration analysis for each event. Such a description gives OS developers better guidance for reducing OS noise. We integrated this data with a trace visualizer allowing quicker and more intuitive understanding of the data. Specifically, the contributions of this paper are three-fold. First, we describe a methodology whereby detailed quantitative information may be obtained for each OS noise event. Though not the thrust of the paper, we show how we implemented that methodology by augmenting LTTng. We validate our approach by comparing it to other well-known standard techniques to analyze OS noise. Second, we provide a case study in which we use our methodology to analyze the OS noise when running benchmarks from the LLNL Sequoia applications. Our experiments enrich and expand previous results with our quantitative characterization. Third, we describe how a detailed characterization permits to disambiguate noise signatures of qualitatively similar events, allowing developers to address the true cause of each noise event. Alessandro Morari, Roberto Gioiosa, Robert W. Wisniewski, Francisco J. Cazorla, Mateo Valero |
IPDPS | 3 |
| 2010 | Providing a cloud network infrastructure on a supercomputerabstractSupercomputers and clouds both strive to make a large number of computing cores available for computation. More recently, similar objectives such as low-power, manageability at scale, and low cost of ownership are driving a more converged hardware and software. Challenges remain, however, of which one is that current cloud infrastructure does not yield the performance sought by many scientific applications. A source of the performance loss comes from virtualization and virtualization of the network in particular. This paper provides an introduction and analysis of a hybrid supercomputer software infrastructure, which allows direct hardware access to the communication hardware for the necessary components while providing the standard elastic cloud infrastructure for other components. Jonathan Appavoo, Amos Waterland, Dilma Da Silva, Volkmar Uhlig, Bryan S. Rosenburg, Eric Van Hensbergen, Jan Stoess, Robert W. Wisniewski, Udo Steinberg |
HPDC | 8 |
| 2010 | Experiences with a Lightweight Supercomputer Kernel: Lessons Learned from Blue Gene's CNKabstractThe Petascale era has recently been ushered in and many researchers have already turned their attention to the challenges of exascale computing. To achieve petascale computing two broad approaches for kernels were taken, a lightweight approach embodied by IBM Blue Gene's CNK, and a more fullweight approach embodied by Cray's CNL. There are strengths and weaknesses to each approach. Examining the current generation can provide insight as to what mechanisms may be needed for the exascale generation. The contributions of this paper are the experiences we had with CNK on Blue Gene/P. We demonstrate it is possible to implement a small lightweight kernel that scales well but still provides a Linux environment and functionality desired by HPC programmers. Such an approach provides the values of reproducibility, low noise, high and stable performance, reliability, and ease of effectively exploiting unique hardware features. We describe the strengths and weaknesses of this approach. Mark Giampapa, Thomas Gooding, Todd Inglett, Robert W. Wisniewski |
SC | 4 |
| 2009 | Minimizing dependencies within generic classes for faster and smaller programsabstractGeneric classes can be used to improve performance by allowing compile-time polymorphism. But the applicability of compile-time polymorphism is narrower than that of runtime polymorphism, and it might bloat the object code. We advocate a programming principle whereby a generic class should be implemented in a way that minimizes the dependencies between its members (nested types, methods) and its generic type parameters. Conforming to this principle (1) reduces the bloat and (2) gives rise to a previously unconceived manner of using the language that expands the applicability of compile-time polymorphism to a wider range of problems. Our contribution is thus a programming technique that generates faster and smaller programs. We apply our ideas to GCC's STL containers and iterators, and we demonstrate notable speedups and reduction in object code size (real application runs 1.2x to 2.1x faster and STL code is 1x to 25x smaller). We conclude that standard generic APIs (like STL) should be amended to reflect the proposed principle in the interest of efficiency and compactness. Such modifications will not break old code, simply increase flexibility. Our findings apply to languages like C++, C#, and D, which realize generic programming through multiple instantiations. Dan Tsafrir, Robert W. Wisniewski, David F. Bacon, Bjarne Stroustrup |
OOPSLA | 2 |
| 2007 | Experiences Understanding Performance in a Commercial Scale-Out Environment
Robert W. Wisniewski, Mathieu Desnoyers, Maged M. Michael, José E. Moreira, Doron Shiloach, Livio B. Soares |
Euro-Par | 1 |
| 2007 | Scale-up x Scale-out: A Case Study using Nutch/LuceneabstractScale-up solutions in the form of large SMPs have represented the mainstream of commercial computing for the past several years. The major server vendors continue to provide increasingly larger and more powerful machines. More recently, scale-out solutions, in the form of clusters of smaller machines, have gained increased acceptance for commercial computing. Scale-out solutions are particularly effective in high-throughput Web-centric applications. In this paper, we investigate the behavior of two competing approaches to parallelism, scale-up and scale-out, in an emerging search application. Our conclusions show that a scale-out strategy can be the key to good performance even on a scale-up machine. Furthermore, scale-out solutions offer better price/performance, although at an increase in management complexity. Maged M. Michael, José E. Moreira, Doron Shiloach, Robert W. Wisniewski |
IPDPS | 4 |
| 2007 | Reboots Are for Hardware: Challenges and Solutions to Updating an Operating System on the Fly
Andrew Baumann, Jonathan Appavoo, Robert W. Wisniewski, Dilma Da Silva, Orran Krieger, Gernot Heiser |
USENIX ATC | 3 |
| 2007 | Libra: a library operating system for a jvm in a virtualized execution environmentabstractIf the operating system could be specialized for every application, many applications would run faster. For example, Java virtual machines (JVMs) provide their own threading model and memory protection, so general-purpose operating system implementations of these abstractions are redundant. However, traditional means of transforming existing systems into specialized systems are difficult to adopt because they require replacing the entire operating system. This paper describes Libra, an execution environment specialized for IBM's J9 JVM. Libra does not replace the entire operating system. Instead, Libra and J9 form a single statically-linked image that runs in a hypervisor partition. Libra provides the services necessary to achieve good performance for the Java workloads of interest but relies on an instance of Linux in another hypervisor partition to provide a networking stack, a filesystem, and other services. The expense of remote calls is offset by the fact that Libra's services can be customized for a particular workload; for example, on the Nutch search engine, we show that two simple customizations improve application throughput by a factor of 2.7. Glenn Ammons, Jonathan Appavoo, Maria A. Butrico, Dilma Da Silva, David Grove, Kiyokuni Kawachiya, Orran Krieger, Bryan S. Rosenburg, Eric Van Hensbergen, Robert W. Wisniewski |
VEE | 10 |
| 2007 | Experience distributing objects in an SMMP OSabstractDesigning and implementing system software so that it scales well on shared-memory multiprocessors (SMMPs) has proven to be surprisingly challenging. To improve scalability, most designers to date have focused on concurrency by iteratively eliminating the need for locks and reducing lock contention. However, our experience indicates that locality is just as, if not more, important and that focusing on locality ultimately leads to a more scalable system. In this paper, we describe a methodology and a framework for constructing system software structured for locality, exploiting techniques similar to those used in distributed systems. Specifically, we found two techniques to be effective in improving scalability of SMMP operating systems: (i) an object-oriented structure that minimizes sharing by providing a natural mapping from independent requests to independent code paths and data structures, and (ii) the selective partitioning, distribution, and replication of object implementations in order to improve locality. We describe concrete examples of distributed objects and our experience implementing them. We demonstrate that the distributed implementations improve the scalability of operating-system-intensive parallel workloads. Jonathan Appavoo, Dilma Da Silva, Orran Krieger, Marc A. Auslander, Michal Ostrowski, Bryan S. Rosenburg, Amos Waterland, Robert W. Wisniewski, Jimi Xenidis, Michael Stumm, Livio B. Soares |
ACM Trans. Comput. Syst. | 8 |
| 2006 | K42: building a complete operating systemabstractK42 is one of the few recent research projects that is examining operating system design structure issues in the context of new whole-system design. K42 is open source and was designed from the ground up to perform well and to be scalable, customizable, and maintainable. The project was begun in 1996 by a team at IBM Research. Over the last nine years there has been a development effort on K42 from between six to twenty researchers and developers across IBM, collaborating universities, and national laboratories. K42 supports the Linux API and ABI, and is able to run unmodified Linux applications and libraries. The approach we took in K42 to achieve scalability and customizability has been successful.The project has produced positive research results, has resulted in contributions to Linux and the Xen hypervisor on Power, and continues to be a rich platform for exploring system software technology. Today, K42, is one of the key exploratory platforms in the DOE's FAST-OS program, is being used as a prototyping vehicle in IBM's PERCS project, and is being used by universities and national labs for exploratory research. In this paper, we provide insight into building an entire system by discussing the motivation and history of K42, describing its fundamental technologies, and presenting an overview of the research directions we have been pursuing. Orran Krieger, Marc A. Auslander, Bryan S. Rosenburg, Robert W. Wisniewski, Jimi Xenidis, Dilma Da Silva, Michal Ostrowski, Jonathan Appavoo, Maria A. Butrico, Mark F. Mergen, Amos Waterland, Volkmar Uhlig |
EuroSys | 4 |
| 2005 | Online performance analysis by statistical sampling of microprocessor performance countersabstractHardware performance counters (HPCs) are increasingly being used to analyze performance and identify the causes of performance bottlenecks. However, HPCs are difficult to use for several reasons. Microprocessors do not provide enough counters to simultaneously monitor the many different types of events needed to form an over-all understanding of performance. Moreover, HPCs primarily count low-level micro-architectural events from which it is difficult to extract high-level insight required for identifying causes of performance problems.We describe two techniques that help overcome these difficulties, allowing HPCs to be used in dynamic real-time optimizers. First, statistical sampling is used to dynamically multiplex HPCs and make a larger set of logical HPCs available. Using real programs, we show experimentally that it is possible through this sampling to obtain counts of hardware events that are statistically similar (within 15%) to complete non-sampled counts, thus allowing us to provide a much larger set of logical HPCs. Second, we observe that stall cycles are a primary source of inefficiencies, and hence they should be major targets for software optimization. Based on this observation, we build a simple model in real-time that speculatively associates each stall cycle to a processor component that likely caused the stall. The information needed to produce this model is obtained using our HPC multiplexing facility to monitor a large number of hardware components simultaneously. Our analysis shows that even in an out-of-order superscalar micro-processor such a speculative approach yields a fairly accurate model with run-time overhead for collection and computation of under 2%.These results demonstrate that we can effective analyze on-line performance of application and system code running at full speed. The stall analysis shows where performance is being lost on a given processor. Michael Stumm, Robert W. Wisniewski |
ICS | 3 |
| 2005 | Providing Dynamic Update in an Operating System
Andrew Baumann, Gernot Heiser, Jonathan Appavoo, Dilma Da Silva, Orran Krieger, Robert W. Wisniewski, Jeremy Kerr |
USENIX ATC, General Track | 6 |
| 2003 | Efficient, Unified, and Scalable Performance Monitoring for Multiprocessor Operating SystemsabstractProgramming, understanding, and tuning the performance of large multiprocessor systems is challenging. Experts have difficulty achieving good utilization for applications on large machines. The task of implementing scalable systems such as an operating system or database on large machines is even more challenging. And the importance of achieving good performance on multiprocessor machines is increasing as the number of cores per chip increases and as the size of multiprocessors increases. Crucial to achieving good performance is being able to understand the behavior of the system. Robert W. Wisniewski, Bryan S. Rosenburg |
SC | 1 |
| 2003 | System Support for Online Reconfiguration
Craig A. N. Soules, Jonathan Appavoo, Kevin Hui, Robert W. Wisniewski, Dilma Da Silva, Gregory R. Ganger, Orran Krieger, Michael Stumm, Marc A. Auslander, Michal Ostrowski, Bryan S. Rosenburg, Jimi Xenidis |
USENIX ATC, General Track | 4 |
| 2001 | Supporting Hot-Swappable Components for System SoftwareabstractSummary form only given. A hot-swappable component is one that can be replaced with a new or different implementation while the system is running and actively using the component. For example, a component of a TCP/IP protocol stack, when hot-swappable, can be replaced (perhaps to handle new denial-of-service attacks or improve performance), without disturbing existing network connections. The capability to swap components offers a number of potential advantages such as: online upgrades for high availability systems, improved performance due to dynamic adaptability and simplified software structures by allowing distinct policy and implementation options to be implemented in separate components (rather than as a single monolithic component) and dynamically swapped as needed. In order to hot-swap a component, it is necessary to (i) instantiate a replacement component; (ii) establish a quiescent state in which the component is temporarily idle; (iii) transfer state from the old component to the new component; (iv) swap the new component for the old; and (v) deallocate the old component. Kevin Hui, Jonathan Appavoo, Robert W. Wisniewski, Marc A. Auslander, David Edelsohn, Benjamin Gamsa, Orran Krieger, Bryan S. Rosenburg, Michael Stumm |
HotOS | 3 |
| 1997 | Scheduler-Conscious SynchronizationabstractEfficient synchronization is important for achieving good performance in parallel programs, especially on large-scale multiprocessors. Most synchronization algorithms have been designed to run on a dedicated machine, with one application process per processor, and can suffer serious performance degradation in the presence of multiprogramming. Problems arise when running processes block or, worse, busy-wait for action on the part of a process that the scheduler has chosen not to run. We show that these problems are particularly severe for scalable synchronization algorithms based on distributed data structures. We then describe and evaluate a set of algorithms that perform well in the presence of multiprogramming while maintaining good performance on dedicated machines. We consider both large and small machines, with a particular focus on scalability, and examine mutual-exclusion locks, reader-writer locks, and barriers. Our algorithms vary in the degree of support required from the kernel scheduler. We find that while it is possible to avoid pathological performance problems using previously proposed kernel mechanisms, a modest additional widening of the kernel/user interface can make scheduler-conscious synchronization algorithms significantly simpler and faster, with performance on dedicated machines comparable to that of scheduler-oblivious algorithms. Leonidas I. Kontothanassis, Robert W. Wisniewski, Michael L. Scott |
ACM Trans. Comput. Syst. | 2 |
| 1995 | Adaptable Planner Primitives for Real-World Robotic Applications
Robert W. Wisniewski |
IJCAI | 1 |
| 1995 | High Performance Synchronization Algorithms for Multiprogrammed MultiprocessorsabstractScalable busy-wait synchronization algorithms are essential for achieving good parallel program performance on large scale multiprocessors. Such algorithms include mutual exclusion locks, reader-writer locks, and barrier synchronization. Unfortunately, scalable synchronization algorithms are particularly sensitive to the effects of multiprogramming: their performance degrades sharply when processors are shared among different applications, or even among processes of the same application. In this paper we describe the design and evaluation of scalable scheduler-conscious mutual exclusion locks, reader-writer locks, and barriers, and show that by sharing information across the kernel/application interface we can improve the performance of scheduler-oblivious implementations by more than an order of magnitude. Robert W. Wisniewski, Leonidas I. Kontothanassis, Michael L. Scott |
PPoPP | 1 |
| 1993 | Using Scheduler Information to Achieve Optimal Barrier Synchronization PerformanceabstractParallel programs frequently use barriers to synchronize successive steps in an algorithm. In the presence of multiprogramming the choice of spinning versus blocking barriers can have a significant impact on performance. We demonstrate how competitive spinning techniques previously designed for locks can be extended to barriers, and we evaluate their performance. We design an additional competitive spinning technique that adapts more quickly in a dynamic environment. We then propose and evaluate a new method that obtains better peformance than previous techniques by using scheduler information to decide between spinning and blocking. The scheduler information technique makes optimal choices incurring little overhead. Leonidas I. Kontothanassis, Robert W. Wisniewski |
PPoPP | 2 |