Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Henry Schuh

dblp:246/9302 · also Henry N. Schuh · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0004-8430-4104ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Memory systems · 30% Cloud and datacenter computing · 22% Hardware accelerators and domain-specific architectures · 17%
Computer networks
2 papers
Datacenter networks · 74% Transport protocols and congestion control · 26%
Databases, data mining, and information retrieval
1 paper
Transaction processing and concurrency control · 100%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › network accelerator
SmartNIC
0.922021
Xenic: SmartNIC-Accelerated Distributed Transactions · SOSP 2021
Offloading distributed applications onto smartNICs using iPipe · SIGCOMM 2019
Transport protocols and congestion control
delay-based congestion control
0.912025
Falcon: A Reliable, Low Latency Hardware Transport · SIGCOMM 2025
Datacenter networks › load balancing
multipath load balancing
0.912025
Falcon: A Reliable, Low Latency Hardware Transport · SIGCOMM 2025
Memory systems › cache coherence
cache-coherent interconnect
0.812024
CC-NIC: a Cache-Coherent Interface to the NIC · ASPLOS (1) 2024
Cloud and datacenter computing
datacenter application performance
0.812024
Understanding the Host Network · SIGCOMM 2024
Performance modeling and evaluation
workload characterization
0.812024
Understanding the Host Network · SIGCOMM 2024
Transaction processing and concurrency control
distributed transaction processing
0.512021
Xenic: SmartNIC-Accelerated Distributed Transactions · SOSP 2021
Storage systems › file systems
distributed file system
0.412020
Assise: Performance and Availability via Client-local NVM in a Distributed File System · OSDI 2020
Memory systems
cache coherence
0.212024
CC-NIC: a Cache-Coherent Interface to the NIC · ASPLOS (1) 2024
Memory systems › cache coherence
cache coherence protocol
0.212024
CC-NIC: a Cache-Coherent Interface to the NIC · ASPLOS (1) 2024
Memory systems
memory interconnect
0.212024
Understanding the Host Network · SIGCOMM 2024
Interconnection networks and networks-on-chip
remote direct memory access
0.112021
Xenic: SmartNIC-Accelerated Distributed Transactions · SOSP 2021
Distributed systems › distributed communication
remote memory access
0.112021
Xenic: SmartNIC-Accelerated Distributed Transactions · SOSP 2021
Memory systems
non-volatile memory
0.112020
Assise: Performance and Availability via Client-local NVM in a Distributed File System · OSDI 2020

Methods — techniques the papers use, named apart from their topics

production datacenter measurement · 1.5point-to-point communication · 1.0asynchronous aggregated execution · 1.0programmable engine · 0.9hardware retransmission · 0.9characterization · 0.4
YearPublicationVenuePosition
2025 Falcon: A Reliable, Low Latency Hardware Transport
abstract
Hardware transports such as RoCE deliver high performance with minimal host CPU, but are best suited to special-purpose deployments that limit their use, e.g., backend networks or Ethernet with Priority Flow Control (PFC). We introduce Falcon, the first hardware transport that supports multiple Upper Layer Protocols (ULPs) and heterogeneous application workloads in general-purpose Ethernet datacenter environments (with losses and without special switch support). Key design elements include: delay-based congestion control with multipath load balancing; a layered design with a simple request-response transaction interface for multi-ULP support; hardware-based retransmissions and error-handling for scalability; and a programmable engine for flexibility. The first Falcon hardware implementation delivers a peak performance of 200 Gbps, 120 Mops/sec, with near-optimal operation completion times that are up to 8× lower than CX-7 RoCE under network congestion, and up to 65% higher goodput under lossy conditions.
Arjun Singhvi, Nandita Dukkipati, Prashant Chandra, Hassan M. G. Wassel, Naveen Kr. Sharma, Anthony Rebello, Henry Schuh, Praveen Kumar 0003, Behnam Montazeri, Neelesh Bansod, Sarin Thomas, Inho Cho, Hyojeong Lee Seibert, Baijun Wu, Rui Yang 0034, Qianwen Yin, Srinivas Vaduvatha, Weihuang Wang, Masoud Moshref, David Wetherall, Amin Vahdat
SIGCOMM7
2024 CC-NIC: a Cache-Coherent Interface to the NIC
abstract
Emerging interconnects make peripherals, such as the network interface controller (NIC), accessible through the processor's cache hierarchy, allowing these devices to participate in the CPU cache coherence protocol. This is a fundamental change from the separate I/O data paths and read-write transaction primitives of today's PCIe NICs. Our experiments show that the I/O data path characteristics cause NICs to prioritize CPU efficiency at the expense of inflated latency, an issue that can be mitigated by the emerging low-latency coherent interconnects. But, the coherence abstraction is not suited to current host-NIC access patterns. Applying existing signaling mechanisms and data structure layouts in a cache-coherent setting results in extraneous communication and cache retention, limiting performance. Redesigning the interface is necessary to minimize overheads and benefit from the new interactions coherence enables. This work contributes CC-NIC, a host-NIC interface design for coherent interconnects. We model CC-NIC using Intel's Ice Lake and Sapphire Rapids UPI interconnects, demonstrating the potential of optimizing for coherence. Our results show a maximum packet rate of 1.5Gpps and 980Gbps packet throughput. CC-NIC has 77% lower minimum latency, and 88% lower at 80% load, than today's PCIe NICs. We also demonstrate application-level core savings. Finally, we show that CC-NIC's benefits hold across a range of interconnect performance characteristics.
Henry Schuh, Arvind Krishnamurthy, David E. Culler, Henry M. Levy, Luigi Rizzo, Samira Manabi Khan, Brent E. Stephens
ASPLOS (1)1
2024 Understanding the Host Network
abstract
The host network integrates processor, memory, and peripheral interconnects to enable data transfer within the host. Several recent studies from production datacenters show that contention within the host network can have significant impact on end-to-end application performance. The goal of this paper is to build an in-depth understanding of such contention within the host network.
Midhul Vuppalapati, Saksham Agarwal, Henry Schuh, Baris Kasikci, Arvind Krishnamurthy, Rachit Agarwal 0001
SIGCOMM3
2021 Xenic: SmartNIC-Accelerated Distributed Transactions
abstract
High-performance distributed transactions require efficient remote operations on database memory and protocol metadata. The high communication cost of this workload calls for hardware acceleration. Recent research has applied RDMA to this end, leveraging the network controller to manipulate host memory without consuming CPU cycles on the target server. However, the basic read/write RDMA primitives demand trade-offs in data structure and protocol design, limiting their benefits. SmartNICs are a flexible alternative for fast distributed transactions, adding programmable compute cores and on-board memory to the network interface. Applying measured performance characteristics, we design Xenic, a SmartNIC-optimized transaction processing system. Xenic applies an asynchronous, aggregated execution model to maximize network and core efficiency. Xenic's co-designed data store achieves low-overhead remote object accesses. Additionally, Xenic uses flexible, point-to-point communication patterns between SmartNICs to minimize transaction commit latency. We compare Xenic against prior RDMA- and RPC-based transaction systems with the TPC-C, Retwis, and Smallbank benchmarks. Our results for the three benchmarks show 2.42x, 2.07x, and 2.21x throughput improvement, 59%, 42%, and 22% latency reduction, while saving 2.3, 8.1, and 10.1 threads per server.
Henry Schuh, Weihao Liang, Ming Liu 0027, Jacob Nelson 0001, Arvind Krishnamurthy
SOSP1
2020 Assise: Performance and Availability via Client-local NVM in a Distributed File System
Thomas E. Anderson, Marco Canini, Jongyul Kim 0001, Dejan Kostic, Youngjin Kwon, Simon Peter 0001, Waleed Reda, Henry Schuh, Emmett Witchel
OSDI8
2019 Offloading distributed applications onto smartNICs using iPipe
abstract
Emerging Multicore SoC SmartNICs, enclosing rich computing resources (e.g., a multicore processor, onboard DRAM, accelerators, programmable DMA engines), hold the potential to offload generic datacenter server tasks. However, it is unclear how to use a SmartNIC efficiently and maximize the offloading benefits, especially for distributed applications. Towards this end, we characterize four commodity SmartNICs and summarize the offloading performance implications from four perspectives: traffic control, computing capability, onboard memory, and host communication.
Ming Liu 0027, Tianyi Cui, Henry Schuh, Arvind Krishnamurthy, Simon Peter 0001
SIGCOMM3