Larry Dennison

dblp:54/10926 · also Larry R. Dennison · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0001-5533-1083ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 ACTINA: Adapting Circuit-Switching Techniques for AI Networking Architectures
abstract
While traditional datacenters rely on static, electrically switched fabrics, Optical Circuit Switch (OCS)-enabled reconfigurable networks offer dynamic bandwidth allocation and lower power consumption. This work introduces a quantitative framework for evaluating reconfigurable networks in large-scale AI systems, guiding the adoption of various OCS and link technologies by analyzing trade-offs in reconfiguration latency, link bandwidth provisioning, and OCS placement. Using this framework, we develop two in-workload reconfiguration strategies and propose an OCS-enabled, multi-dimensional all-to-all topology that supports hybrid parallelism with improved energy efficiency. Our evaluation demonstrates that with state-of-the-art per-GPU bandwidth, the optimal in-workload strategy achieves up to 2.3 × improvement over the commonly used one-shot approach when reconfiguration latency is low (<100 μ s). However, with sufficiently high bandwidth, one-shot reconfiguration can achieve comparable performance without requiring in-workload reconfiguration. Additionally, our proposed topology improves performance–power efficiency, achieving up to 1.75 × better trade-offs than Fat-Tree and 3D-Torus–based OCS network architectures.
Zhenguo Wu, Benjamin Klenk, Larry Dennison, Keren Bergman
SC3
2023 Calculon: a methodology and tool for high-level co-design of systems and large language models
abstract
This paper presents a parameterized analytical performance model of transformer-based Large Language Models (LLMs) for guiding high-level algorithm-architecture codesign studies. This model derives from an extensive survey of performance optimizations that have been proposed for the training and inference of LLMs; the model's parameters capture application characteristics, the hardware system, and the space of implementation strategies. With such a model, we can systematically explore a joint space of hardware and software configurations to identify optimal system designs under given constraints, like the total amount of system memory. We implemented this model and methodology in a Python-based open-source tool called Calculon. Using it, we identified novel system designs that look significantly different from current inference and training systems, showing quantitatively the estimated potential to achieve higher efficiency, lower cost, and better scalability.
Mikhail Isaev, Nic McDonald, Larry Dennison, Richard W. Vuduc
SC3
2022 A Case For Intra-rack Resource Disaggregation in HPC
abstract
The expected halt of traditional technology scaling is motivating increased heterogeneity in high-performance computing (HPC) systems with the emergence of numerous specialized accelerators. As heterogeneity increases, so does the risk of underutilizing expensive hardware resources if we preserve today’s rigid node configuration and reservation strategies. This has sparked interest in resource disaggregation to enable finer-grain allocation of hardware resources to applications. However, there is currently no data-driven study of what range of disaggregation is appropriate in HPC. To that end, we perform a detailed analysis of key metrics sampled in NERSC’s Cori, a production HPC system that executes a diverse open-science HPC workload. In addition, we profile a variety of deep-learning applications to represent an emerging workload. We show that for a rack (cabinet) configuration and applications similar to Cori, a central processing unit with intra-rack disaggregation has a 99.5% probability to find all resources it requires inside its rack. In addition, ideal intra-rack resource disaggregation in Cori could reduce memory and NIC resources by 5.36% to 69.01% and still satisfy the worst-case average rack utilization.
George Michelogiannakis, Benjamin Klenk, Brandon Cook 0001, Min Yee Teh, Madeleine Glick, Larry Dennison, Keren Bergman, John Shalf
ACM Trans. Archit. Code Optim.6
2020 An In-Network Architecture for Accelerating Shared-Memory Multiprocessor Collectives
abstract
The slowdown of single-chip performance scaling combined with the growing demands of computing ever larger problems efficiently has led to a renewed interest in distributed architectures and specialized hardware. Dedicated accelerators for common or critical operations are becoming cost-effective additions to processors, peripherals, and networks. In this paper we focus on one such operation, the All-Reduce, which is both a common and critical feature of neural network training. All-Reduce is impossible to fully parallelize and difficult to amortize, so it benefits greatly from hardware acceleration. We are proposing an accelerator-centric, shared-memory network that improves All-Reduce performance through in-network reductions, as well as accelerating other collectives like Multicast. We propose switch designs to support in-network computation, including two reduction methods that offer trade-offs in implementation complexity and performance. Additionally, we propose network endpoint modifications to further improve collectives. We present simulation results for a 16 GPU system showing that our collective acceleration design improves the All-Reduce operation by up to 2x for large messages and up to 18x for small messages when compared with a state-of-the-art software algorithm, leading up to 1.4x faster DL training times for networks like Transformer. We demonstrate that this design is scalable to large systems and present results for up to 128 GPUs.
Benjamin Klenk, Nan Jiang 0009, Greg Thorson, Larry Dennison
ISCA4
2019 Bandwidth steering in HPC using silicon nanophotonics
abstract
As bytes-per-FLOP ratios continue to decline, communication is becoming a bottleneck for performance scaling. This paper describes bandwidth steering in HPC using emerging reconfigurable silicon photonic switches. We demonstrate that placing photonics in the lower layers of a hierarchical topology efficiently changes the connectivity and consequently allows operators to recover from system fragmentation that is otherwise hard to mitigate using common task placement strategies. Bandwidth steering enables efficient utilization of the higher layers of the topology and reduces cost with no performance penalties. In our simulations with a few thousand network endpoints, bandwidth steering reduces static power consumption per unit throughput by 36% and dynamic power consumption by 14% compared to a reference fat tree topology. Such improvements magnify as we taper the bandwidth of the upper network layer. In our hardware testbed, bandwidth steering improves total application execution time by 69%, unaffected by bandwidth tapering.
George Michelogiannakis, Yiwen Shen 0002, Min Yee Teh, Xiang Meng 0003, Benjamin Aivazi, Taylor L. Groves, John Shalf, Madeleine Glick, Manya Ghobadi, Larry Dennison, Keren Bergman
SC10
2018 Exploiting idle resources in a high-radix switch for supplemental storage
Matthias A. Blumrich, Nan Jiang 0009, Larry Dennison
SC3
2018 Light-weight protocols for wire-speed ordering
Hans Eberle, Larry Dennison
SC2
2017 Relaxations for High-Performance Message Passing on Massively Parallel SIMT Processors
abstract
Accelerators, such as GPUs, have proven to be highly successful in reducing execution time and power consumption of compute-intensive applications. Even though they are already used pervasively, they are typically supervised by general-purpose CPUs, which results in frequent control flow switches and data transfers as CPUs are handling all communication tasks. However, we observe that accelerators are recently being augmented with peer-to-peer communication capabilities that allow for autonomous traffic sourcing and sinking. While appropriate hardware support is becoming available, it seems that the right communication semantics are yet to be identified. Maintaining the semantics of existing communication models, such as the Message Passing Interface (MPI), seems problematic as they have been designed for the CPU’s execution model, which inherently differs from such specialized processors. In this paper, we analyze the compatibility of traditional message passing with massively parallel Single Instruction Multiple Thread (SIMT) architectures, as represented by GPUs, and focus on the message matching problem. We begin with a fully MPI-compliant set of guarantees, including tag and source wildcards and message ordering. Based on an analysis of exascale proxy applications, we start relaxing these guarantees to adapt message passing to the GPU’s execution model. We present suitable algorithms for message matching on GPUs that can yield matching rates of 60M and 500M matches/s, depending on the constraints that are being relaxed. We discuss our experiments and create an understanding of the mismatch of current message passing protocols and the architecture and execution model of SIMT processors.
Benjamin Klenk, Holger Fröning, Hans Eberle, Larry Dennison
IPDPS4
2015 Network endpoint congestion control for fine-grained communication
abstract
Endpoint congestion in HPC networks creates tree saturation that is detrimental to performance. Endpoint congestion can be alleviated by reducing the injection rate of traffic sources, but requires fast reaction time to avoid congestion buildup. Congestion control becomes more challenging as application communication shift from traditional two-sided model to potentially fine-grained, one-sided communication embodied by various global address space programming models. Existing hardware solutions, such as Explicit Congestion Notification (ECN) and Speculative Reservation Protocol (SRP), either react too slowly or incur too much overhead for small messages.
Nan Jiang 0009, Larry Dennison, William J. Dally
SC2
2012 Compiling high throughput network processors
abstract
Gorilla is a methodology for generating FPGA-based solutions especially well suited for data parallel applications with fine grain irregularity. Irregularity simultaneously destroys performance and increases power consumption on many data parallel processors such as General Purpose Graphical Processor Units (GPGPUs). Gorilla achieves high performance and low power through the use of FPGA-tailored parallelization techniques and application-specific hardwired accelerators, processing engines, and communication mechanisms. Automatic compilation from a stylized C language and templates that define the hardware structure coupled with the intrinsic flexibility of FPGAs provide high performance, low power, and programmability.
Maysam Lavasani, Larry Dennison, Derek Chiou
FPGA2
1990 Simultaneous bidirectional signalling for IC systems
abstract
A design and implementation of an input/output (I/O) circuit capable of simultaneous bidirectional transmission in CMOS integrated circuits are presented. Conventional techniques for improving communication bandwidth between VLSI chips use the tri-state bidirectional drivers and/or add more I/O pins to a given chip. The design of simultaneous bidirectional transmission of signals between two chips results in optimal performance while maintaining the same pin count. The circuitry has a low output-voltage swing and occupies relatively small silicon area, so that both high speed and low power dissipation are achieved. Included also is the design of a current driver that can properly terminate a given transmission line without the use of any off-chip termination resistors. The design remains functional under significant process variation. Simulation results show that the I/O circuit can be clocked with a frequency over 50 MHz under the worst-case condition for a typical 2- mu m CMOS process. This implies a bandwidth of more than 100 Mb/s per I/O pin.>
Kevin Lam, Larry Dennison, William J. Dally
ICCD2