EDBT 2026 Demo / reviewers in the wild / expert
Hanjiang Wu
dblp:393/5005
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0003-3718-5272ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCALE: Tackling Communication Bottlenecks in Confidential Distributed Machine LearningabstractMachine Learning (ML) has become a cornerstone in numerous applications, creating the need for secure and efficient distributed ML frameworks. However, maintaining data privacy in these systems poses significant challenges, particularly in distributed environments where user data and model parameters must frequently be transmitted between GPUs. Confidential GPU computing technologies, such as NVIDIA's Confidential Computing (CC) mode, offer hardware-based enterprise solutions designed to protect ML workloads in untrusted environments (e.g., public clouds). These technologies leverage heterogeneous systems that combine Confidential Virtual Machines (CVMs) with GPU-based Trusted Execution Environments (TEEs). Nevertheless, confidential computing introduces considerable performance overhead due to its complex heterogeneous architecture and the high-throughput data flows required across TEE security boundaries. For example, encrypted communication occurs both between CVMs and GPU TEEs, and among multiple GPU TEEs, resulting in significant latency compared to native PCIe or high-speed interconnects such as NVLink. Our extensive evaluation shows that these overheads become particularly severe during collective communication operations, which suffer from encryption-induced delays that negatively impact end-to-end training performance. To address this, we propose a co-encryption design that leverages underutilized GPU resources, optimizes encryption and authentication, and introduces a communication algorithm tailored for confidential settings. We evaluate our design using real ML workloads and execution traces collected from four HGX H100/H200 clusters. While CC mode was not available on current NVIDIA software stacks, we incorporate encryptionaware modeling based on hardware specifications to estimate secure communication overheads. Our results demonstrate a$\mathbf{4 0 - 7 0 \%}$reduction in communication-related security costs. Joongun Park, Yongqin Wang, Hanjiang Wu, Tushar Krishna |
HPCA | 4 |
| 2026 | Closing the Efficiency Gap: AI Datacenter Co-design Roadmap for Scalable Training of LLMsabstractThe massive compute, memory, and networking needs for LLM training necessitate a fundamental rethinking of datacenter architectures to ensure scalability, efficiency, and cost-effectiveness. In particular, the design of the network fabric for AI datacenters for emerging LLMs (such as MoEs) remains a crucial and challenging open question, spanning technology choices (that determine the size of the high-bandwidth domain), topology, and software optimizations (collective algorithms and overlap strategies). This necessitates an agile framework to traverse the co-design space. This work introduces Calculon-MoE, a tool that jointly explores FLOPS, HBM bandwidth and capacity, multiple network topologies (Two-tiered vs. FullFlat optical), the size of scale-up domain, and popular parallelism/optimization strategies used in LLMs. Our validation studies demonstrate that our LLM/MoE runtime predictions are within 10% of real-world measurements. Using Calculon-MoE, we conduct a suite of case studies to develop an actionable roadmap for data centers. For example, the results point to the promise of Fullflat network architectures, which provide uniform high-bandwidth, low-latency connectivity between all nodes and demonstrate their positive impacts on performance and scalability. We also quantify the benefits of overlapping compute and communication, hardware-accelerated collectives, widening the scale-up domain, and higher memory bandwidth and capacity. Our study spans both sparse (mixture of experts) and dense transformer-based LLMs, revealing how system design and optimization choices affect system efficiency and overall throughput in both cases. Jesmin Jahan Tithi, Hanjiang Wu, Joongun Park, Avishaii Abuhatzera, Fabrizio Petrini, Tushar Krishna |
ICS | 2 |
| 2026 | Scalable Synthesis of Distributed Llm Workloads Through Symbolic Tensor Graphs
Changhai Man, Joongun Park, Hanjiang Wu, Srinivas Sridharan 0002, Tushar Krishna |
ISCA | 3 |
| 2026 | DynamoServe: A Distributed Tiered Memory System for Multi-tenant LLM ServingabstractThe rapid adoption of large language models (LLMs) has increased the need for efficient multi-tenant inference systems that maximize GPU utilization. However, existing frameworks struggle to scale due to the high memory demands of model weights and key-value (KV) caches. We present DynamoServe, a multi-tenant LLM serving framework that addresses these challenges through three key innovations: (1) leveraging stranded GPU memory to offload model weights and KV caches, (2) mitigating resource fragmentation in multi-workload environments, and (3) improving memory locality through coordinated data placement and demand-driven weight migration across GPUs. Together, these techniques enable high-throughput, low-latency inference. Experiments on state-of-the-art models show that DynamoServe significantly improves memory efficiency without sacrificing latency. Diman Zad Tootaghaj, Khaled Diab 0001, Bob Lantz, Hanjiang Wu, K. K. Ramakrishnan, Md Ashfaqur Rahaman, Ryan Stutsman, Puneet Sharma 0001, Tushar Krishna |
SIGCOMM | 4 |
| 2025 | Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal PerspectiveabstractThe rapid scaling of Large Language Models (LLMs) has pushed training workloads far beyond the limits of single-node analysis, demanding a deeper understanding of how these models behave across large-scale, multi-GPU systems.In this paper, we present a comprehensive characterization of LLM training across diverse real-world workloads and hardware platforms, including NVIDIA H100/H200 and AMD MI250 GPUs.We analyze dense and sparse models under various parallelism strategies -tensor, pipeline, data, and expert -and evaluate their effects on hardware utilization, power consumption, and thermal behavior.We further evaluate the effectiveness of optimizations such as activation recomputation and compute-communication overlap.Our findings show that performance is not determined solely by scaling hardware capacity.Scale-up systems with fewer, higher-memory GPUs can outperform scale-out systems in communication-bound regimes, but only under carefully tuned configurations; in other cases, scale-out deployments achieve superior throughput.We also show that certain parallelism combinations, such as tensor with pipeline, lead to bandwidth underutilization due to inefficient data chunking, while increasing microbatch sizes beyond a certain point induces bursty execution and peak power excursions that worsen thermal throttling.These insights reveal how training performance is shaped by complex interactions between hardware, system topology, and model execution.We conclude by offering recommendations for system and hardware design to improve the scalability and reliability of future LLM systems and workloads.The source code of this project is available at https:/ Seokjin Go, Joongun Park, Spandan More, Hanjiang Wu, Irene Wang, Aaron Jezghani, Tushar Krishna, Divya Mahajan 0001 |
MICRO | 4 |