Tarannum Khan

dblp:286/1997 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0004-6744-6186ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 75% Interconnection networks and networks-on-chip · 25%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Computer networks
1 paper
Content delivery and video streaming · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
0.712023
Shockwave: Fair and Efficient Cluster Scheduling for Dynamic Adaptation in Machine Learning · NSDI 2023
Content delivery and video streaming › content delivery network
CDN caching
0.712023
Darwin: Flexible Learning-based CDN Caching · SIGCOMM 2023
Cloud and datacenter computing › cluster resource management and scheduling
cluster scheduling
0.712023
Shockwave: Fair and Efficient Cluster Scheduling for Dynamic Adaptation in Machine Learning · NSDI 2023
Cloud and datacenter computing
datacenter network
0.712023
RingLeader: Efficiently Offloading Intra-Server Orchestration to NICs · NSDI 2023
Interconnection networks and networks-on-chip › network interface
network interface card
0.712023
RingLeader: Efficiently Offloading Intra-Server Orchestration to NICs · NSDI 2023
Cloud and datacenter computing › computation offloading › network function offloading
SmartNIC offload
0.712023
RingLeader: Efficiently Offloading Intra-Server Orchestration to NICs · NSDI 2023

Methods — techniques the papers use, named apart from their topics

fair scheduling · 1.3unsupervised clustering · 0.7reinforcement learning · 0.7neural bandit · 0.7
YearPublicationVenuePosition
2026 Reforge: Low-Latency Distributed GNN Serving with Selective Embedding Recomputation
Geon-Woo Kim, Donghyun Kim 0002, Jeongyoon Moon, Henry Liu, Tarannum Khan, Anand Iyer, Daehyeok Kim, Aditya Akella
IPDPS5
2023 RingLeader: Efficiently Offloading Intra-Server Orchestration to NICs
Adney Cardoza, Tarannum Khan, Yeonju Ro, Brent E. Stephens, Hassan M. G. Wassel, Aditya Akella
NSDI3
2023 Shockwave: Fair and Efficient Cluster Scheduling for Dynamic Adaptation in Machine Learning
Rui Pan 0003, Tarannum Khan, Shivaram Venkataraman, Aditya Akella
NSDI3
2023 Darwin: Flexible Learning-based CDN Caching
abstract
Cache management is critical for Content Delivery Networks (CDNs), impacting their performance and operational costs. Most production CDNs apply static, hand-tuned caching policy parameters at cache servers, such as admission frequency or size thresholds for the Hot Object Caches (HOC) of their system. However, these static policies fall short when a server is faced with unpredictable traffic pattern changes, even when policies employ multiple control parameters/knobs. Recent approaches have proposed learning-based solutions to dynamically adjust policy parameters, but they are limited in action space, caching objectives, or impose high overhead. We propose Darwin, a CDN cache management system that is robust to traffic pattern changes and can flexibly optimize different caching objectives with unrestricted action spaces. Darwin employs a three-stage pipeline involving traffic pattern feature collection, unsupervised clustering for classification, and neural bandit expert selection to choose the optimal caching policy. Through extensive simulations, experiments using an Apache Traffic Server (ATS)-based prototype, and theoretical analysis, we show that Darwin achieves significant performance gain w.r.t. different objectives such as maximizing object hit rates and minimizing disk writes, while simultaneously adapting to traffic pattern shifts. Darwin imposes negligible overhead and achieves high throughput compared to the state-of-the-art.
Nihal Sharma, Tarannum Khan, Brian Chang, Aditya Akella, Sanjay Shakkottai, Ramesh K. Sitaraman
SIGCOMM3
2022 Impact of RoCE Congestion Control Policies on Distributed Training of DNNs
abstract
Ahstract-RDMA over Converged Ethernet (RoCE) has gained significant attraction for datacenter networks due to its compatibility with conventional Ethernet-based fabric. However, the RDMA protocol is efficient only on (nearly) lossless networks, emphasizing the vital role of congestion control on RoCE networks. Unfortunately, the native RoCE congestion control scheme, based on Priority Flow Control (PFC), suffers from many drawbacks such as unfairness, head-of-line-blocking, and deadlock. Therefore, in recent years many schemes have been proposed to provide additional congestion control for RoCE networks to minimize PFC drawbacks. However, these schemes are proposed for general datacenter environments. In contrast to the general datacenters that are built using commodity hardware and run general-purpose workloads, high-performance distributed training platforms deploy high-end accelerators and network components and exclusively run training workloads using collectives (All-Reduce, All-To-All) communication libraries for communication. Furthermore, these platforms usually have a private network, separating their communication traffic from the rest of the datacenter traffic. Scalable topology-aware collective algorithms are inherently designed to avoid incast patterns and balance traffic optimally. These distinct features necessitate revisiting previously proposed congestion control schemes for general-purpose datacenter environments. In this paper, we thoroughly analyze some of the state-of-the-art RoCE congestion control schemes (DCQCN, DCTCP, TIMELY, and HPCC) vs. PFC when running on distributed training platforms. Our results indicate that pre-viously proposed RoCE congestion control schemes have little impact on the end-to-end performance of training workloads, motivating the necessity of designing an optimized, yet low-overhead, congestion control scheme based on the characteristics of distributed training platforms and workloads.
Tarannum Khan, Saeed Rashidi, Srinivas Sridharan 0002, Pallavi Shurpali, Aditya Akella, Tushar Krishna
HOTI1