Duncan Roweth

dblp:78/21 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Flowcut Switching: High-Performance Adaptive Routing With In-Order Delivery Guarantees
abstract
Network latency severely impacts the performance of applications running on supercomputers. Adaptive routing algorithms route packets over different available paths to reduce latency and improve network utilization. However, if a switch routes packets belonging to the same network flow on different paths, they might arrive at the destination out-of-order due to differences in the latency of these paths. For some transport protocols like TCP, QUIC, and RoCE, out-of-order (OOO) packets might cause large performance drops or significantly increase CPU utilization. In this work, we proposeFlowcut switching, a new adaptive routing algorithm that provides high-performance in-order packet delivery. Differently from existing solutions likeFlowlet switching, which are based on the assumption of bursty traffic and that might still reorder packets,Flowcut switchingguarantees in-order delivery under any network conditions, and is effective also for non-bursty traffic, as it is often the case for RDMA. On top of this, Flowcut can be implemented either at the switch or NIC level providing flexibility and different tradeoffs.
Tommaso Bonato, Daniele De Sensi, Salvatore Di Girolamo, Abdulla Bataineh, David Hewson, Duncan Roweth, Torsten Hoefler
IEEE Trans. Netw.6
2024 Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
abstract
Multi-GPU nodes are increasingly common in the rapidly evolving landscape of exascale supercomputers. On these systems, GPUs on the same node are connected through dedicated networks, with bandwidths up to a few terabits per second. However, gauging performance expectations and maximizing system efficiency is challenging due to different technologies, design options, and software layers. This paper comprehensively characterizes three supercomputers — Alps, Leonardo, and LUMI — each with a unique architecture and design. We focus on performance evaluation of intra-node and inter-node interconnects on up to 4,096 GPUs, using a mix of intra-node and inter-node benchmarks. By analyzing its limitations and opportunities, we aim to offer practical guidance to researchers, system architects, and software developers dealing with multi-GPU supercomputing. Our results show that there is untapped bandwidth, and there are still many opportunities for optimization, ranging from network to software optimization.
Daniele De Sensi, Lorenzo Pichetti, Flavio Vella, Tiziano De Matteis, Zebin Ren, Luigi Fusco, Matteo Turisini, Daniele Cesarini, Kurt Lust, Animesh Trivedi, Duncan Roweth, Filippo Spiga, Salvatore Di Girolamo, Torsten Hoefler
SC11
2023 Not all applications have boring communication patterns: Profiling message matching with BMM
abstract
Summary Message matching within MPI is an important performance consideration for applications that utilize two‐sided semantics. In this work, we present an instrumentation of the CrayMPI library that allows the collection of detailed message‐matching statistics as well as an implementation of hashed matching in software. We use this functionality to profile key DOE applications with complex communication patterns to determine under what circumstances an application might benefit from hardware offload capabilities within the NIC to accelerate message matching. We find that there are several applications and libraries that exhibit sufficiently long match list lengths to motivate a Binned Message Matching approach.
Taylor L. Groves, Naveen Ravichandrasekaran, Brandon Cook 0001, Noel Keen, David Trebotich, Nicholas J. Wright, Robert Alverson, Duncan Roweth, Keith D. Underwood
Concurr. Comput. Pract. Exp.8
2021 Future of HPC: Diversifying Heterogeneity
abstract
After the end of Dennard scaling and with the imminent end of Moore's Law, it has become challenging to continue scaling HPC systems within a given power envelope. This is exacerbated most in large systems, such as high end supercomputers. To alleviate this problem, general purpose is no longer sufficient, and HPC systems and components are being augmented with special-purpose hardware. By definition, because of the narrow applicability of specialization, broad supercomputing adoption requires using different heterogeneous components, each optimized for a specific application domain. In this paper, we discuss the impact of the introduced heterogeneity of specialization across the HPC stack: interconnects including memory models, accelerators including power and cooling, use cases and applications including AI, and delivery models, such as traditional, as-a-Service, and federated. We believe that a stack that supports diversification across hardware and software is required to continue scaling performance and maintaining energy efficiency.
Dejan S. Milojicic, Paolo Faraboschi, Nicolas Dubé, Duncan Roweth
DATE4
2020 An in-depth analysis of the slingshot interconnect
abstract
The interconnect is one of the most critical components in large scale computing systems, and its impact on the performance of applications is going to increase with the system size. In this paper, we will describe SLINGSHOT, an interconnection network for large scale computing systems. SLINGSHOT is based on high-radix switches, which allow building exascale and hyper-scale datacenters networks with at most three switch-to-switch hops. Moreover, SLINGSHOT provides efficient adaptive routing and congestion control algorithms, and highly tunable traffic classes. SLINGSHOT uses an optimized Ethernet protocol, which allows it to be interoperable with standard Ethernet devices while providing high performance to HPC applications. We analyze the extent to which SLINGSHOT provides these features, evaluating it on microbenchmarks and on several applications from the datacenter and AI worlds, as well as on HPC applications. We find that applications running on SLINGSHOT are less affected by congestion compared to previous generation networks.
Daniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth, Torsten Hoefler
SC4
2019 Network-accelerated non-contiguous memory transfers
abstract
Applications often communicate data that is non-contiguous in the send- or the receive-buffer, e.g., when exchanging a column of a matrix stored in row-major order. While non-contiguous transfers are well supported in HPC (e.g., MPI derived datatypes), they can still be up to 5x slower than contiguous transfers of the same size. As we enter the era of network acceleration, we need to investigate which tasks to offload to the NIC: In this work we argue that non-contiguous memory transfers can be transparently network-accelerated, truly achieving zero-copy communications. We implement and extend sPIN, a packet streaming processor, within a Portals 4 NIC SST model, and evaluate strategies for NIC-offloaded processing of MPI datatypes, ranging from datatype-specific handlers to general solutions for any MPI datatype. We demonstrate up to 8x speedup in the unpack throughput of real applications, demonstrating that non-contiguous memory transfers are a first-class candidate for network acceleration.
Salvatore Di Girolamo, Konstantin Taranov, Andreas Kurth, Michael Schaffner, Timo Schneider, Jakub Beránek, Maciej Besta, Luca Benini, Duncan Roweth, Torsten Hoefler
SC9
2012 Cray cascade: a scalable HPC system based on a Dragonfly network
abstract
Higher global bandwidth requirement for many applications and lower network cost have motivated the use of the Dragonfly network topology for high performance computing systems. In this paper we present the architecture of the Cray Cascade system, a distributed memory system based on the Dragonfly [1] network topology. We describe the structure of the system, its Dragonfly network and the routing algorithms. We describe a set of advanced features supporting both mainstream high performance computing applications and emerging global address space programing models. We present a combination of performance results from prototype systems and simulation data for large systems. We demonstrate the value of the Dragonfly topology and the benefits obtained through extensive use of adaptive routing.
Greg Faanes, Abdulla Bataineh, Duncan Roweth, Tom Court, Edwin Froese, Robert Alverson, Tim Johnson, Joe Kopnick, Mike Higgins, James Reinhard
SC3
2006 High performance interconnects - High performance networks for the future
abstract
QsNetIII and 10 Gbit/s Ethernet, two networks for high perfomance computing. While the proprietary QsNet will continue to provide supercomputing funcionalities to clusters of commodity based servers, the second will establish itself as preferred choice in capacity class systems.Quadrics, who is actively developing products based on the two technologies, will present early performance results and comparisons.
Duncan Roweth, Moray McLaren
SC1
1993 The Meiko CS-2 System Architecture
abstract
No abstract available.
Duncan Roweth
SPAA1
1988 Neural network models
B. M. Forrest, Duncan Roweth, N. Stroud, D. J. Wallace, Gregory V. Wilson
Parallel Comput.2
1987 Implementing Neural Network Models on Parallel Computers
abstract
The remarkable processing capabilities of the nervous system must derive from the large numbers of neurons participating (roughly 1010), since the time-scales involved are of the order of a millisecond, rather than the nanoseconds of modern computers. The neural network models which attempt to capture this behaviour are inherently parallel. We review the implementation of a range of neural network models on SIMD and MIMD computers. On the ICL Distributed Array Processor (DAP), a 4096-processor SIMD machine, we have studied training algorithms in the context of the Hopfield net, with specific applications including the storage of words and continuous text in content-addressable memory. The Hopfield and Tank analogue neural net has been used for image restoration with the Geman and Geman algorithm. We compare the performance of this scheme on the DAP and on a Meiko Computing Surface, a reconfigurable MIMD array of transputers. We describe also the strategies which we have used to implement the Durbin and Willshaw elastic net model on the Computing Surface.
B. M. Forrest, Duncan Roweth, N. Stroud, D. J. Wallace, Gregory V. Wilson
Comput. J.2
1986 Design and performance analysis of Transputer arrays
Duncan Roweth
J. Syst. Softw.1