VLDB 2026 Research / reviewers in the wild / expert
Saeed Rashidi
dblp:217/4121
· DBLP profile ↗
10ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-6472-9920ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 4 first-author · 6 since 2021Systems, architecture and hardware · 5 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FRED: A Wafer-scale Fabric for 3D Parallel DNN TrainingabstractWafer-scale systems are an emerging technology that tightly integrates high-end accelerator chiplets with high-speed wafer-scale interconnects, enabling low-latency and high-bandwidth connectivity.This makes them a promising platform for deep neural network (DNN) training.However, current network-on-wafer topologies, such as 2D Meshes, lack the flexibility needed to support various parallelization strategies effectively.In this paper, we propose Fred, a wafer-scale fabric architecture tailored to the unique communication needs of DNN training.Fred creates a distributed on-wafer topology with tiny microswitches, providing nonblocking connectivity for collective communications between arbitrary groups of accelerators and enabling in-switch collective support.Our results show that for sample parallelization strategies, Fred can improve the average end-to-end training time of ResNet-152, Transformer-17B, GPT-3, and Transformer-1T by 1.76×, 1.87×, 1.34×, and 1.4×, respectively, compared to a baseline wafer-scale Mesh. Saeed Rashidi, William Won, Sudarshan Srinivasan, Puneet Gupta 0001, Tushar Krishna |
ISCA | 1 |
| 2024 | Leveraging Memory Expansion to Accelerate Large-Scale DL TrainingabstractModern Deep Learning (DL) models require massive clusters of specialized, high-end nodes to train. Designing such clusters to maximize both performance and utilization is a challenging task requiring careful balance of compute, memory, and network resources. Moreover, a plethora of each model's tuning knobs drastically affect a cluster's performance, with optimal values often depending on the underlying hardware characteristics, thus necessitating a complex cluster-workload co-design process. To facilitate design space exploration, we develop Come T,a holistic cluster design methodology and workflow to jointly study the impact of parallelization strategies and key cluster resource provisioning on the performance of distributed DL training. In this short paper, we use Comet to focus on studying the impact of each compute node's memory system on overall cluster performance. We show that per-node memory expansion using technologies like CXL is a promising technique to boost a cluster's DL training performance. Divya Kiran Kadiyala, Saeed Rashidi, Taekyung Heo, Abhimanyu Bambhaniya, Tushar Krishna, Alexandros Daglis |
ISPASS | 2 |
| 2024 | LIBRA: Enabling Workload-Aware Multi-Dimensional Network Topology Optimization for Distributed Training of Large AI ModelsabstractAs model sizes in machine learning continue to scale, distributed training is necessary to accommodate model weights within each device and to reduce training time. However, this comes with the expense of increased communication overhead due to the exchange of gradients and activations, which become the critical bottleneck of the end-to-end training process. In this work, we motivate the design of multi-dimensional networks within machine learning systems as a cost-efficient mechanism to enhance overall network bandwidth. We also identify that optimal bandwidth allocation is pivotal for multi-dimensional networks to ensure efficient resource utilization. We introduce Libra, a framework specifically focused on optimizing multi-dimensional fabric architectures. Through case studies, we demonstrate the value of Libra, both in architecting optimized fabrics under diverse constraints and in enabling co-optimization opportunities. William Won, Saeed Rashidi, Sudarshan Srinivasan, Tushar Krishna |
ISPASS | 2 |
| 2023 | Efficient Distributed Inference of Deep Neural Networks via Restructuring and PruningabstractIn this paper, we consider the parallel implementation of an already-trained deep model on multiple processing nodes (a.k.a. workers). Specifically, we investigate as to how a deep model should be divided into several parallel sub-models, each of which is executed efficiently by a worker. Since latency due to synchronization and data transfer among workers negatively impacts the performance of the parallel implementation, it is desirable to have minimum interdependency among parallel sub-models. To achieve this goal, we propose to rearrange the neurons in the neural network, partition them (without changing the general topology of the neural network), and modify the weights such that the interdependency among sub-models is minimized under the computations and communications constraints of the workers while minimizing its impact on the performance of the model. We propose RePurpose, a layer-wise model restructuring and pruning technique that guarantees the performance of the overall parallelized model. To efficiently apply RePurpose, we propose an approach based on L0 optimization and the Munkres assignment algorithm. We show that, compared to the existing methods, RePurpose significantly improves the efficiency of the distributed inference via parallel implementation, both in terms of communication and computational complexity. Afshin Abdi, Saeed Rashidi, Faramarz Fekri, Tushar Krishna |
AAAI | 2 |
| 2023 | ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at ScaleabstractAs deep learning models and input data continue to scale at an unprecedented rate, it has become inevitable to move towards distributed training platforms to fit the models and increase training throughput. State-of-the-art distributed training systems are adopting emerging approaches and techniques such as wafer-scale nodes, multi-dimensional network topologies, disaggregated memory systems, and optimized parallelization strategies. This results in a complex software/hardware co-design stack, necessitating a modeling/simulation infrastructure for design-space exploration. This paper introduces ASTRA-sim2.0, which extends the open-source ASTRA-sim infrastructure with capabilities to model state-of-the-art and emerging distributed training models and platforms. Specifically, we enable ASTRAsim to (i) support arbitrary model parallelization strategies via a graph-based training-loop implementation, (ii) implement a parameterizable multi-dimensional heterogeneous topology generation infrastructure with the capability to simulate target systems at scale through analytical performance estimation, and (iii) enhance memory system modeling to support accurate modeling of in-network collective communication and disaggregated memory systems. With these capabilities, we conduct comprehensive case studies targeting emerging distributed models and platforms. ASTRA-sim2.0 enables system designers to swiftly traverse the complex co-design stack and gain meaningful insights when designing and deploying distributed training platforms at scale. William Won, Taekyung Heo, Saeed Rashidi, Srinivas Sridharan 0002, Sudarshan Srinivasan, Tushar Krishna |
ISPASS | 3 |
| 2022 | Impact of RoCE Congestion Control Policies on Distributed Training of DNNsabstractAhstract-RDMA over Converged Ethernet (RoCE) has gained significant attraction for datacenter networks due to its compatibility with conventional Ethernet-based fabric. However, the RDMA protocol is efficient only on (nearly) lossless networks, emphasizing the vital role of congestion control on RoCE networks. Unfortunately, the native RoCE congestion control scheme, based on Priority Flow Control (PFC), suffers from many drawbacks such as unfairness, head-of-line-blocking, and deadlock. Therefore, in recent years many schemes have been proposed to provide additional congestion control for RoCE networks to minimize PFC drawbacks. However, these schemes are proposed for general datacenter environments. In contrast to the general datacenters that are built using commodity hardware and run general-purpose workloads, high-performance distributed training platforms deploy high-end accelerators and network components and exclusively run training workloads using collectives (All-Reduce, All-To-All) communication libraries for communication. Furthermore, these platforms usually have a private network, separating their communication traffic from the rest of the datacenter traffic. Scalable topology-aware collective algorithms are inherently designed to avoid incast patterns and balance traffic optimally. These distinct features necessitate revisiting previously proposed congestion control schemes for general-purpose datacenter environments. In this paper, we thoroughly analyze some of the state-of-the-art RoCE congestion control schemes (DCQCN, DCTCP, TIMELY, and HPCC) vs. PFC when running on distributed training platforms. Our results indicate that pre-viously proposed RoCE congestion control schemes have little impact on the end-to-end performance of training workloads, motivating the necessity of designing an optimized, yet low-overhead, congestion control scheme based on the characteristics of distributed training platforms and workloads. Tarannum Khan, Saeed Rashidi, Srinivas Sridharan 0002, Pallavi Shurpali, Aditya Akella, Tushar Krishna |
HOTI | 2 |
| 2022 | Themis: a network bandwidth-aware collective scheduling policy for distributed training of DL modelsabstractDistributed training is a solution to reduce DNN training time by splitting the task across multiple NPUs (e.g., GPU/TPU). However, distributed training adds communication overhead between the NPUs in order to synchronize the gradients and/or activation, depending on the parallelization strategy. In next-generation platforms for training at scale, NPUs will be connected through multidimensional networks with diverse, heterogeneous bandwidths. This work identifies a looming challenge of keeping all network dimensions busy and maximizing the network BW within the hybrid environment if we leverage scheduling techniques for collective communication on systems today. We propose Themis, a novel collective scheduling scheme that dynamically schedules collectives (divided into chunks) to balance the communication loads across all dimensions, further improving the network BW utilization. Our results show that on average, Themis can improve the network BW utilization of the single All-Reduce by 1.72× (2.70× max), and improve the end-to-end training iteration performance of real workloads such as ResNet-152, GNMT, DLRM, and Transformer-1T by 1.49× (2.25× max), 1.30× (1.78× max), 1.30× (1.77× max), and 1.25× (1.53× max), respectively. Saeed Rashidi, William Won, Sudarshan Srinivasan, Srinivas Sridharan 0002, Tushar Krishna |
ISCA | 1 |
| 2021 | Enabling Compute-Communication Overlap in Distributed Deep Learning Training PlatformsabstractDeep Learning (DL) training platforms are built by interconnecting multiple DL accelerators (e.g., GPU/TPU) via fast, customized interconnects with 100s of gigabytes (GBs) of bandwidth. However, as we identify in this work, driving this bandwidth is quite challenging. This is because there is a pernicious balance between using the accelerator’s compute and memory for both DL computations and communication.This work makes two key contributions. First, via real system measurements and detailed modeling, we provide an understanding of compute and memory bandwidth demands for DL compute and comms. Second, we propose a novel DL collective communication accelerator called Accelerator Collectives Engine (ACE) that sits alongside the compute and networking engines at the accelerator endpoint. ACE frees up the endpoint’s compute and memory resources for DL compute, which in turn reduces the required memory BW by 3.5× on average to drive the same network BW compared to state-of-the-art baselines. For modern DL workloads and different network sizes, ACE, on average, increases the effective network bandwidth utilization by 1.44× (up to 2.67×), resulting in an average of 1.41× (up to 1.51×), 1.12× (up to 1.17×), and 1.13× (up to 1.19×) speedup in iteration time for ResNet-50, GNMT and DLRM when compared to the best baseline configuration, respectively. Saeed Rashidi, Matthew Denton, Srinivas Sridharan 0002, Sudarshan Srinivasan, Amoghavarsha Suresh, Jade Nie, Tushar Krishna |
ISCA | 1 |
| 2020 | ASTRA-SIM: Enabling SW/HW Co-Design Exploration for Distributed DL Training PlatformsabstractModern Deep Learning systems heavily rely on distributed training over high-performance accelerator (e.g., TPU, GPU)-based hardware platforms. Examples today include Google's Cloud TPU and Facebook's Zion. DNN training involves a complex interplay between the DNN model architecture, parallelization strategy, scheduling strategy, collective communication algorithm, network topology, and the end-point accelerator. As innovation in AI/ML models continues to grow at an accelerated rate, there is a need for a comprehensive methodology to understand and navigate this complex SW/HW design-space for future systems to support efficient training of future DNN models. In this work, we make the following contributions (i) establish the SW/HW design-space for Distributed Training over a hierarchical scale-up fabric, (ii) develop a network simulator for navigating the design-space, and (iii) demonstrate the promise of algorithm-topology co-design for speeding up end to end training. Saeed Rashidi, Srinivas Sridharan 0002, Sudarshan Srinivasan, Tushar Krishna |
ISPASS | 1 |
| 2018 | Improving MLC PCM Performance through Relaxed Write and Read for Intermediate Resistance LevelsabstractPhase Change Memory (PCM) is one of the most promising candidates to be used at the main memory level of the memory hierarchy due to poor scalability, considerable leakage power, and high cost/bit of DRAM. PCM is a new resistive memory that is capable of storing data based on resistance values. The wide resistance range of PCM allows for storing multiple bits per cell (MLC) rather than a single bit per cell (SLC). Unfortunately, higher density of MLC PCM comes at the expense of longer read/write latency, higher soft error rate, higher energy consumption, and earlier wearout compared to the SLC PCM. Some studies suggest removing the most error-prone level to mitigate soft error and write latency of MLC PCM, hence introducing a less dense memory called Tri-Level memory. Another scheme, called M-Metric, proposes a new read metric to address the soft error problem in MLC PCM. In order to deal with the limited lifetime of PCM, some extra storage per memory line is required to correct permanent hard errors (stuck-at faults). Since the extra storage is used only when permanent faults occur, it has a low utilization for a long time before hard errors start to occur. In this article, we utilize the extra storage to improve the read/write latency in a 2-bit MLC PCM using a relaxation scheme for reading and writing the cells for intermediate resistance levels. More specifically, we combine the most time-consuming levels (intermediate resistance levels) to reduce the number of resistance levels (making a Tri-Level PCM) and therefore improve write latency. We then store some error correction metadata in the extra storage section to successfully retrieve the exact data values in the read operation. We also modify the Tri-Level PCM cell to increase its read latency when the M-Metric scheme is used. Evaluation results show that the proposed scheme improves read latency by 57.2%, write latency by 56.1%, and overall system performance (IPC) by 26.9% over the baseline. It is noteworthy that combining the proposed scheme and FPC compression method improves read latency by 75.2%, write latency by 67%, and overall system performance (IPC) by 37.4%. Saeed Rashidi, Majid Jalili 0001, Hamid Sarbazi-Azad |
ACM Trans. Archit. Code Optim. | 1 |