Darius Bunandar

dblp:265/0978 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-8218-5656ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reconfigurable Torus Fabrics for Multi-tenant ML
abstract
We develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML datacenters with Morphlux can improve the bandwidth of tenant compute allocations by up to 66%, reduce compute fragmentation by up to 70%, and minimize the blast radius of accelerator failures. We develop a novel end-to-end hardware prototype of Morphlux to demonstrate these performance benefits which translate to 1.72x improvement in finetuning throughput of ML models. By rapidly programming the server-scale fabric in our hardware testbed, Morphlux can replace a failed accelerator with a healthy one in 1.2 seconds.
Abhishek Vijaya Kumar, Ding Ding 0005, Arjun Devraj, Darius Bunandar, Rachee Singh
ASPLOS (2)4
2025 Passage M1000 : A 3D Photonic Interposer for AI
Darius Bunandar
HCS1
2025 Waferscale Silicon Photonics Systems: A Cost-Benefit Analysis and Optimization
abstract
Silicon photonics holds considerable promise for reducing long reach communication overheads in future computing systems. Similarly, waferscale integration promises dramatic improvements in performance and energy efficiency for scale out systems, but suffers from the long reach limitations of electrical interconnects. No prior work has looked at the performance benefits of silicon photonics over electrical interconnects to address the long reach challenges of waferscale integration, or at the overheads of silicon photonics for such systems across multiple implementations. In this work, we study a tile-based silicon photonics waferscale system for different implementations of waveguide networks and topologies, and across multiple applications and number of tiles. We find that the performance benefits of using silicon photonics instead of electrical interconnects at waferscale are highly application-dependent - benefits primarily come from reduced communication latency. The power and area overheads of implementation are high, especially for high connectivity topologies and when reconfigurability is considered. Some implementations are infeasible - the microring resonator maximum power limits are exceeded for these implementations. Custom waveguide networks address the problem and limit the overheads when supporting high connectivity topologies and reconfigurability. Overall, this is the first paper to analyze the performance benefits of silicon photonics vs electrical interconnects at waferscale and optimize the implementation overheads of waferscale silicon photonics systems.
Robert Bao, Zongrui Cai, Shuangliang Chen, Ajay Joshi, Darius Bunandar, Rakesh Kumar 0002
ICCAD5
2025 EPiCarbon: A Carbon Modeling Tool for Electro-Photonic Accelerators
abstract
The escalating carbon emissions driven by the growing computational demands of Artificial Intelligence (AI) have made energy-efficient and sustainable hardware design a high priority. Photonic computing has emerged as a promising solution, delivering orders of magnitude higher throughput and energy efficiency than CMOS for deep neural network inferences, thereby lowering operational carbon. However, studies have shown that the carbon emission from manufacturing, i.e., embodied carbon, constitutes a substantial and often dominant portion of the total carbon footprint of a computing system. Hence, it is crucial to consider both operational and embodied carbon to determine the true benefits of photonic computing. While the embodied carbon of CMOS chips and CMOS-based systems has been studied extensively, there is currently no model available for estimating the embodied carbon of photonic chips.In this work, we develop the first-ever model to estimate the embodied carbon of photonic chips. Our findings show that photonic chips can reduce the embodied carbon of computing systems with at least 4.1× less fabrication energy and significantly higher yield than CMOS. Building on our model, we introduce EPiCarbon, an open-source tool to evaluate the carbon footprint of Electro-Photonic (EPiC) accelerators, incorporating both operational and embodied carbon. Using EPiCarbon, we analyze the carbon footprint of state-of-the-art EPiC accelerators, demonstrating their potential as carbon-sustainable solutions for computationally demanding AI applications. Finally, through a case study on a comprehensive EPiC accelerator, ADEPT, we demonstrate key strategies to further reduce the carbon footprint of EPiC accelerators, guiding future sustainable hardware design.
Farbin Fayza, Cansu Demirkiran, Satyavolu Papa Rao, Darius Bunandar, Ajay Joshi
ICCAD4
2025 TrioSim: A Lightweight Simulator for Large-Scale DNN Workloads on Multi-GPU Systems
abstract
Deep Neural Networks (DNNs) have become increasingly capable of performing tasks ranging from image recognition to content generation.The training and inference of DNNs heavily rely on GPUs, as GPUs' massively parallel architecture delivers extremely high computing capability.With the growing complexity of DNNs and the size of training datasets, training DNNs with a large number of GPUs is becoming a prevalent strategy.Researchers have been exploring how to design software and hardware systems for GPU farms to achieve the best utilization, efficiency, and DNN accuracy during training or inference.However, when designing and deploying such systems, designers usually rely on testing on physical hardware platforms equipped with many GPUs, incurring high costs that are almost prohibitive for system designers to test different configurations and designs, even for highly resourceful companies.While an alternative solution is to test on GPU simulators, they are often too slow for these large-scale systems and depend on profiling details collected from real distributed systems to initiate the simulation.To address these challenges, we present TrioSim, a novel lightweight simulator for DNNs on multi-GPU systems.TrioSim combines performance modeling techniques and simulation methods to achieve high flexibility, high simulation speed, and * Part of this work was done while Yuhui Bao and Pranav Vaid were interns at Lightmatter.
Ying Li 0049, Yuhui Bao, Gongyu Wang, Xinxin Mei, Pranav Vaid, Anandaroop Ghosh, Adwait Jog, Darius Bunandar, Ajay Joshi, Yifan Sun 0002
ISCA8
2024 A case for server-scale photonic connectivity
abstract
The commoditization of machine learning is fuelling the demand for compute required to both train large models and infer from them. At the same time, scaling the performance of individual microprocessors to satisfy the demand for compute has become increasingly difficult since the end of Moore's law and Dennard scaling. As a result, compute resources in modern servers are distributed across multiple accelerators on the server board. In this work, we make the case for using optics to interconnect accelerators within a server. A key benefit of on-board chip-to-chip optical connectivity is its ability to dynamically allocate bandwidth between accelerators, where necessary, rather than the common practice of statically dividing bandwidth among links within the topology of a multi-accelerator server, as seen in popular direct-connect architectures. This property prevents bandwidth under-utilization in state-of-the-art rack-scale multi-accelerator deployments. Moreover, server-scale optical connectivity can reduce the blast radius of individual accelerator failures in rack-scale ML deployments. Our early experiments with the prototype of a newly commercialized server-scale photonic interconnect show how the capability of the hardware can enable our vision.
Abhishek Vijaya Kumar, Arjun Devraj, Darius Bunandar, Rachee Singh
HotNets3
2024 Mirage: An RNS-Based Photonic Accelerator for DNN Training
abstract
Photonic computing is a compelling avenue for performing highly efficient matrix multiplication, a crucial operation in Deep Neural Networks (DNNs). While this method has shown great success in DNN inference, meeting the high precision demands of DNN training proves challenging due to the precision limitations imposed by costly data converters and the analog noise inherent in photonic hardware. This paper proposes Mirage, a photonic DNN training accelerator that overcomes the precision challenges in photonic hardware using the Residue Number System (RNS). RNS is a numeral system based on modular arithmetic-allowing us to perform high-precision operations via multiple low-precision modular operations. In this work, we present a novel micro-architecture and dataflow for an RNS-based photonic tensor core performing modular arithmetic in the analog domain. By combining RNS and photonics, Mirage provides high energy efficiency without compromising precision and can successfully train state-of-the-art DNNs achieving accuracy comparable to FP32 training. Our study shows that on average across several DNNs when compared to systolic arrays, Mirage achieves more than $23.8 \times$ faster training and $32.1 \times$ lower EDP in an iso-energy scenario and consumes $42.8 \times$ lower power with comparable or better EDP in an iso-area scenario.
Cansu Demirkiran, Guowei Yang 0005, Darius Bunandar, Ajay Joshi
ISCA3
2023 An Electro-Photonic System for Accelerating Deep Neural Networks
abstract
The number of parameters in deep neural networks (DNNs) is scaling at about 5× the rate of Moore’s Law. To sustain this growth, photonic computing is a promising avenue, as it enables higher throughput in dominant general matrix-matrix multiplication (GEMM) operations in DNNs than their electrical counterpart. However, purely photonic systems face several challenges including lack of photonic memory and accumulation of noise. In this article, we present an electro-photonic accelerator, ADEPT, which leverages a photonic computing unit for performing GEMM operations, a vectorized digital electronic application-specific integrated circuits for performing non-GEMM operations, and SRAM arrays for storing DNN parameters and activations. In contrast to prior works in photonic DNN accelerators, we adopt a system-level perspective and show that the gains while large are tempered relative to prior expectations. Our goal is to encourage architects to explore photonic technology in a more pragmatic way considering the system as a whole to understand its general applicability in accelerating today’s DNNs. Our evaluation shows that ADEPT can provide, on average, 5.73× higher throughput per watt compared to the traditional systolic arrays in a full-system, and at least 6.8× and 2.5× better throughput per watt, compared to state-of-the-art electronic and photonic accelerators, respectively.
Cansu Demirkiran, Furkan Eris, Gongyu Wang, Jonathan Elmhurst, Nick Moore, Nicholas C. Harris, Ayon Basumallik, Vijay Janapa Reddi, Ajay Joshi, Darius Bunandar
ACM J. Emerg. Technol. Comput. Syst.10