VLDB 2026 Research / reviewers in the wild / expert
Jack Tigar Humphries
dblp:237/0910
· DBLP profile ↗
7ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0001-5163-516XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 3 first-author · 5 since 2021Computer networks · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Wave: Offloading Resource Management to SmartNIC CoresabstractSmartNICs are increasingly deployed in datacenters to offload tasks from server CPUs, improving the efficiency and flexibility of datacenter security, networking and storage. Optimizing cloud server efficiency in this way is critically important to ensure that virtually all server resources are available to paying customers. Userspace system software, specifically, decision-making tasks performed by various operating system subsystems, is particularly well suited for execution on mid-tier SmartNIC ARM cores. To this end, we introduce Wave, a framework for offloading userspace system software to processes/agents running on the SmartNIC. Wave uses Linux userspace systems to better align system functionality with SmartNIC capabilities. It also introduces a new host-SmartNIC communication API that enables offloading of even μs-scale system software. To evaluate Wave, we offloaded preexisting userspace system software including kernel thread scheduling, memory management, and an RPC stack to SmartNIC ARM cores, which showed a performance degradation of 1.1%-7.4% in an apples-to-apples comparison with on-host implementations. Wave recovered host resources consumed by on-host system software for memory management (saving 16 host cores), RPCs (saving 8 host cores), and virtual machines (an 11.2% performance improvement). Wave highlights the potential for rethinking system software placement in modern datacenters, unlocking new opportunities for efficiency and scalability. Jack Tigar Humphries, Neel Natu, Kostis Kaffes, Stanko Novakovic, Henry M. Levy, David E. Culler, Christoforos E. Kozyrakis |
ASPLOS (3) | 1 |
| 2025 | Towards ML System ExtensibilityabstractWith the rise of large language models, distributed execution across multiple accelerators has become commonplace. Current ML systems must adopt complex distributed execution strategies for efficiency, but do so at the cost of extensibility. We believe that it is time to introduce a general-purpose distributed runtime for programming clusters of accelerators that enables: (1) placement flexibility, and (2) interoperability, without sacrificing (3) codesign. We propose using the DAFT API: distributed actors, futures, and tasks. To enable a smooth tradeoff between flexibility vs. performance, we introduce two execution modes: interpreted vs. compiled. We show how current applications in LLM inference and training can be executed as interpreted and compiled DAFT programs and discuss open questions and challenges. Weixin Deng, Andy Ruan, Megan Frisella, Kai-Hsun Chen, SangBin Cho, Jack Tigar Humphries, Stephanie Wang |
HotOS | 6 |
| 2021 | A case against (most) context switchesabstractMultiplexing software threads onto hardware threads and serving interrupts, VM-exits, and system calls require frequent context switches, causing high overheads and significant kernel and application complexity. We argue that context switching is an idea whose time has come and gone, and propose eliminating it through a radically different hardware threading model targeted to solve software rather than hardware problems. The new model adds a large number of hardware threads to each physical core - making thread multiplexing unnecessary - and lets software manage them. The only state change directly triggered in hardware by system calls, exceptions, and asynchronous hardware events will be blocking and unblocking hardware threads. We also present ISA extensions to allow kernel and user software to exploit this new threading model. Developers can use these extensions to eliminate interrupts and implement fast I/O without polling, exception-less system and hypervisor calls, practical microkernels, simple distributed programming models, and untrusted but fast hypervisors. Finally, we suggest practical hardware implementations and discuss the hardware and software challenges toward realizing this novel approach. Jack Tigar Humphries, Kostis Kaffes, David Mazières, Christoforos E. Kozyrakis |
HotOS | 1 |
| 2021 | ghOSt: Fast & Flexible User-Space Delegation of Linux SchedulingabstractWe present ghOSt, our infrastructure for delegating kernel scheduling decisions to userspace code. ghOSt is designed to support the rapidly evolving needs of our data center workloads and platforms. Jack Tigar Humphries, Neel Natu, Ashwin Chaugule, Ofir Weisse, Barret Rhoden, Josh Don, Luigi Rizzo, Oleg Rombakh, Christoforos E. Kozyrakis |
SOSP | 1 |
| 2021 | Syrup: User-Defined Scheduling Across the StackabstractSuboptimal scheduling decisions in operating systems, networking stacks, and application runtimes are often responsible for poor application performance, including higher latency and lower throughput. These poor decisions stem from a lack of insight into the applications and requests the scheduler is handling and a lack of coherence and coordination between the various layers of the stack, including NICs, kernels, and applications. Kostis Kaffes, Jack Tigar Humphries, David Mazières, Christoforos E. Kozyrakis |
SOSP | 2 |
| 2019 | Mind the Gap: A Case for Informed Request Scheduling at the NICabstractRecent research in high-throughput networked systems has established the need for centralized and preemptive request scheduling in order to achieve good hardware utilization and low tail latency for a wide variety of workloads. However, this approach is expensive to scale as it requires an increasing number of CPU cores dedicated to scheduling. Moreover, passing every request through a scheduling core introduces latency for inter-core communication and reduces the effectiveness of data preloading and caching optimizations. Jack Tigar Humphries, Kostis Kaffes, David Mazières, Christoforos E. Kozyrakis |
HotNets | 1 |
| 2019 | Shinjuku: Preemptive Scheduling for μsecond-scale Tail Latency
Kostis Kaffes, Timothy Chong, Jack Tigar Humphries, Adam Belay, David Mazières, Christoforos E. Kozyrakis |
NSDI | 3 |