Larry Kaplan

dblp:11/10722 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0007-8807-1426ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Security and privacy of machine learning · 50% Authentication and access control · 50%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 55% Distributed systems · 23% Performance modeling and evaluation · 14%
Computer networks
1 paper
Datacenter networks · 50% Network measurement and analytics · 50%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › supercomputing
exascale computing
1.022025
Breaking the System Noise Barrier at Exascale · SC 2025
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
Security and privacy of machine learning
AI agent security
1.012026
InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure Agents · AAAI 2026
Authentication and access control › security policy
policy enforcement
1.012026
InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure Agents · AAAI 2026
Network measurement and analytics › network performance measurement
congestion measurement
0.412020
Measuring Congestion in High-Performance Datacenter Interconnects · NSDI 2020
Natural language and speech › Language models and text generation
large language model
0.312026
InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure Agents · AAAI 2026
Distributed systems › fault tolerance
checkpointing
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
Interconnection networks and networks-on-chip › error control
error detection and recovery
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
Distributed systems
fault tolerance
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
Distributed systems › fault tolerance
resilience
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012

Methods — techniques the papers use, named apart from their topics

policy reasoning · 2.0large language model · 2.0performance analysis · 1.7trace-driven simulation · 0.1analytical modeling · 0.1
YearPublicationVenuePosition
2026 InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure Agents
abstract
The proliferation of Model Context Protocol (MCP) servers in enterprise infrastructure management has revolutionized AI-driven automation while introducing critical multi-layered security vulnerabilities that traditional cybersecurity frameworks cannot adequately address. This paper presents a comprehensive intelligent guardrail system that addresses the unique security challenges of MCP-driven infrastructure management through a novel four-layer defense architecture. Our solution employs a dedicated guardian LLM that interprets natural language policies and applies contextual reasoning to complex infrastructure scenarios, providing dynamic policy enforcement that adapts to user roles, operational timing, and system context. Unlike existing rule-based security systems, our approach implements guardrails at four distinct control points: input message filtering, tool selection validation, execution-time verification, and post-action auditing. The system addresses critical gaps in existing security solutions by providing infrastructure-specific threat modeling, real-time policy adaptation, and comprehensive audit trails with explainable decision-making through confidence scores and detailed reasoning. Our evaluation demonstrates the system's effectiveness in preventing command injection, privilege escalation, and tool poisoning attacks across various enterprise infrastructure scenarios while maintaining operational agility essential for modern data center management.
Aalap Tripathy, Gayathri Saranathan, Martin Foltin, Suparna Bhattacharya, Scott Hinchley, Donald M. Bahls, David Brookshire, Larry Kaplan, Robert W. Wisniewski
AAAI9
2025 GPU Stream-Aware Communication for Effective Pipelining
abstract
Modern heterogeneous supercomputing systems consist of CPUs, GPUs, and high-speed network interconnects. Communication libraries that support efficient inter-process data movement between memory buffers, especially those involving GPU memory, typically require the CPU to orchestrate the data transfer operations. This approach necessitates expensive synchronization between the CPU and GPU, and is ineffective for achieving better compute/communication overlap in applications using techniques like pipelining. A new offload-friendly communication strategy, stream-triggered (ST) communication, is explored to offload the synchronization and data movement operations from the CPU to the GPU. A Message Passing Interface (MPI) one-sided active target synchronization-based implementation is used to illustrate the proposed strategy. A latency-sensitive nearest-neighbor microbenchmark was used to examine various performance characteristics of the implementation. The offloaded implementation showed significant performance improvements both between nodes (inter-node) and within a single node (intranode) when compared to standard MPI active RMA (33% and $\mathbf{2 7 \%}$, respectively) and point-to-point communication ($\mathbf{9 \%}$ and 38%, respectively).
Naveen Namashivayam, Krishna Kandalla, Pen-Chung Yew, Trey White, Larry Kaplan, Mark Pagel
PACT5
2025 Breaking the System Noise Barrier at Exascale
abstract
To meet the increasing demands of parallel scientific applications, supercomputers continue to grow in both scale and complexity. The fastest supercomputer in the world, El Capitan, features over a million CPU cores and tens of thousands of GPUs. Applications running on such large-scale systems are particularly susceptible to system noise or interference caused by the operating system (OS) and other services running on the same compute nodes as the application.
Edgar A. León, Joseph Glenski, Mark J. Stock, Kim H. McMahon, William Loewe, Clark Snyder, Larry Kaplan, Srinath Vadlamani, Timothy I. Mattox, Trent D'Hooge, Brian Behlendorf, Nathan Hanford, Ramesh Pankajakshan, Matthew L. Leininger
SC7
2020 Measuring Congestion in High-Performance Datacenter Interconnects
Saurabh Jha, Archit Patke, Jim M. Brandt, Ann C. Gentile, Benjamin Lim, Michael T. Showerman, Gregory H. Bauer, Larry Kaplan, Zbigniew T. Kalbarczyk, William T. Kramer, Ravishankar K. Iyer
NSDI8
2017 Holistic Measurement-Driven System Assessment
abstract
In high-performance computing systems, application performance and throughput are dependent on a complex interplay of hardware and software subsystems and variable workloads with competing resource demands. Data-driven insights into the potentially widespread scope and propagationof impact of events, such as faults and contention for shared resources, can be used to drive more effective use of resources, for improved root cause diagnosis, and for predicting performance impacts. We present work developing integrated capabilities for holistic monitoring and analysis to understand and characterize propagation of performance-degrading events. These characterizations can be used to determine and invoke mitigating responses by system administrators, applications, and system software.
Saurabh Jha, Jim M. Brandt, Ann C. Gentile, Zbigniew T. Kalbarczyk, Gregory H. Bauer, Jeremy Enos, Michael T. Showerman, Larry Kaplan, Brett M. Bode, Annette Greiner, Amanda Bonnie, Mike Mason, Ravishankar K. Iyer, William T. Kramer
CLUSTER8
2012 Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems
abstract
This paper describes and evaluates a scalable and efficient resilience scheme based on the concept of containment domains. Containment domains are a programming construct that enable applications to express resilience needs and to interact with the system to tune and specialize error detection, state preservation and restoration, and recovery schemes. Containment domains have weak transactional semantics and are nested to take advantage of the machine and application hierarchies and to enable hierarchical state preservation, restoration, and recovery. We evaluate the scalability and efficiency of containment domains using generalized trace-driven simulation and analytical analysis and show that containment domains are superior to both checkpoint restart and redundant execution approaches.
Jinsuk Chung, Ikhwan Lee, Michael B. Sullivan 0001, Jeeho Ryoo, Dong-Wan Kim, Doe Hyun Yoon, Larry Kaplan, Mattan Erez
SC7