EDBT 2026 Demo / reviewers in the wild / expert
Dharanidhar Dang
dblp:161/2547
· DBLP profile ↗
13ranked-venue papers
11as first author
6since 2021 · last 2025
0000-0002-3802-381XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 11 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SOFTONIC: A Photonic Design Approach to Softmax Activation for High-Speed Fully Analog AI Acceleration
Priyabrata Dash, Anxiao Jiang, Dharanidhar Dang |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | SUSTAINPHOT: Sustainable Large-Scale AI Training Using Analog Silicon Photonic AcceleratorsabstractArtificial intelligence (AI) has transformed domains from language to vision, but its rapid growth has led to soaring computational and energy demands that threaten sustainability. At the heart of these workloads are matrix-vector multiplications (MVMs) and nonlinear functions such as Softmax Activation (SMA), which dominate both runtime and power. Transformer models exemplify this challenge, where massive parameter counts and long sequence lengths make MVM and SMA critical bottlenecks. While silicon photonics (SiPh) promises massive bandwidth, low latency, and low power, existing accelerators remain limited to unsigned MVM and rely on CMOS for SMA, incurring repeated E/O conversions and lost efficiency. We present SUSTAINPHOT, a sustainable analog silicon photonic accelerator for large-scale AI training/inference that unifies two key innovations. First, SOFTONIC implements the first fully photonic Softmax engine, with decomposition, polynomial, scaling, and division units that natively compute SMA. Second, MIRAGE realizes a microring-based signed MVM architecture with built-in phase-error compensation, enabling accurate, fullrange computation without duplicating weight banks or bulky equalizers. Together, these components form an end-to-end photonic pipeline that reduces energy and carbon footprint while sustaining accuracy and scalability. Co-simulations with commercial CAD tools demonstrate up to 80% lower power,$109 \times$lower latency, and$199 \times$higher compute density for SMA, while MVM achieves 34 ps latency, 39 fJ/MAC energy, and BER$<3 \times 10^{-4}$under variations. By combining efficiency with resilience, SUSTAINPHOT establishes a path toward light-speed, energy-aware, and environmentally sustainable AI training. Dharanidhar Dang |
ICCD | 1 |
| 2024 | P-ReTI: Silicon Photonic Accelerator for Greener and Real-Time AIabstractComputing deep AI algorithms on traditional CPUs and GPUs brings several performance and energy pitfalls. Most of the emerging AI accelerators target only the inference phase of deep learning. There have been very limited attempts to design a full-fledged AI accelerator capable of both training and inference in real-time. It is due to the highly compute and memory intensive nature of the training phase. In this paper, we propose P-ReTI, a novel analog photonics AI accelerator. P-ReTI uses silicon microdisk-based convolution, photonic phase change memory-based memory, and dense-wavelength-division-multiplexing for energy-efficient and ultrafast deep learning in real-time. We evaluate P-ReTI using a commercial CAD framework (IPKISS) on deep learning benchmark models including LeNet and VGG-Net. Compared to the state-of-the-art, P-ReTI improves the CNN throughput, energy-efficiency, and computational efficiency by up to two orders of magnitude with trivial accuracy degradation. Dharanidhar Dang, Priyabrata Dash, Ahmedullah Aziz |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | Co-designing 2.5D Silicon Photonic Accelerators for Distributed Transformer at the EdgeabstractThe efficient execution of attention-based transformers and large language models on traditional CPUs and GPUs presents significant challenges related to performance and energy efficiency. While innovative solutions like ASICs, FPGAs, and ReRAMs have been explored, the field of silicon photonics has emerged as a promising avenue for developing energy-efficient accelerators for deep AI models. Notably, existing endeavors in silicon photonics have predominantly concentrated on inference for deep AI algorithms, leaving a limited number of initiatives focused on creating comprehensive deep learning accelerators capable of real-time training for transformer-like algorithms. This paper utilizes the superior merits of silicon photonics to realize a full-fledged transformer accelerator equipped for both inference and training. Introducing PHOTRAN, an AI analog photonics accelerator, we harness silicon microdisk-based convolution, photonic phase-change memory-based cache, and dense-wavelength-division-multiplexing to achieve energy-efficient and ultrafast transformer acceleration. Through evaluations using a commercial CAD framework on benchmark models, including Vision Transformers and Large Language models, our results showcase the superior performance of PHOTRAN. This work underscores the significant potential of photonic computing for on-chip training of large deep AI models. Dharanidhar Dang, Priyabrata Dash, Luqi Zheng, Haitong Li |
ICCAD | 1 |
| 2022 | LiteCON: An All-photonic Neuromorphic Accelerator for Energy-efficient Deep LearningabstractDeep learning is highly pervasive in today's data-intensive era. In particular, convolutional neural networks (CNNs) are being widely adopted in a variety of fields for superior accuracy. However, computing deep CNNs on traditional CPUs and GPUs brings several performance and energy pitfalls. Several novel approaches based on ASIC, FPGA, and resistive-memory devices have been recently demonstrated with promising results. Most of them target only the inference (testing) phase of deep learning. There have been very limited attempts to design a full-fledged deep learning accelerator capable of both training and inference. It is due to the highly compute- and memory-intensive nature of the training phase. In this article, we propose LiteCON , a novel analog photonics CNN accelerator. LiteCON uses silicon microdisk-based convolution, memristor-based memory, and dense-wavelength-division-multiplexing for energy-efficient and ultrafast deep learning. We evaluate LiteCON using a commercial CAD framework (IPKISS) on deep learning benchmark models including LeNet and VGG-Net. Compared to the state of the art, LiteCON improves the CNN throughput, energy efficiency, and computational efficiency by up to 32×, 37×, and 5×, respectively, with trivial accuracy degradation. Dharanidhar Dang, Bill Lin 0001, Debashis Sahoo |
ACM Trans. Archit. Code Optim. | 1 |
| 2021 | BPLight-CNN: A Photonics-Based Backpropagation Accelerator for Deep LearningabstractTraining deep learning networks involves continuous weight updates across the various layers of the deep network while using a backpropagation (BP) algorithm. This results in expensive computation overheads during training. Consequently, most deep learning accelerators today employ pretrained weights and focus only on improving the design of the inference phase. The recent trend is to build a complete deep learning accelerator by incorporating the training module. Such efforts require an ultra-fast chip architecture for executing the BP algorithm. In this article, we propose a novel photonics-based backpropagation accelerator for high-performance deep learning training. We present the design for a convolutional neural network (CNN), BPLight-CNN , which incorporates the silicon photonics-based backpropagation accelerator. BPLight-CNN is a first-of-its-kind photonic and memristor-based CNN architecture for end-to-end training and prediction. We evaluate BPLight-CNN using a photonic CAD framework (IPKISS) on deep learning benchmark models, including LeNet and VGG-Net. The proposed design achieves (i) at least 34× speedup, 34× improvement in computational efficiency, and 38.5× energy savings during training; and (ii) 29× speedup, 31× improvement in computational efficiency, and 38.7× improvement in energy savings during inference compared with the state-of-the-art designs. All of these comparisons are done at a 16-bit resolution, and BPLight-CNN achieves these improvements at a cost of approximately 6% lower accuracy compared with the state-of-the-art. Dharanidhar Dang, Sai Vineel Reddy Chittamuru, Sudeep Pasricha, Rabi N. Mahapatra, Debashis Sahoo |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2020 | MEMTONIC: A Neuromorphic Accelerator for Energy Efficient Deep LearningabstractMost deep learning accelerators in the literature focus only on improving the design of inference phase. We propose a novel photonics-based backpropagation accelerator for high performance deep learning training. The proposed MEMTONIC architecture is a first-of-its-kind memristor-integrated photonics-based deep learning architecture for end-to-end training and prediction. We evaluate the architecture using a photonic CAD framework (IPKISS) on deep learning benchmark models including LeNet and VGG-Net. The proposed design achieves at least 35× acceleration in training time, 31× improvement in computational efficiency, and 45× energy savings compared to the state-of-the-art designs, without any loss of accuracy. Dharanidhar Dang, Sahar Taheri, Bill Lin 0001, Debashis Sahoo |
DAC | 1 |
| 2020 | BPhoton-CNN: An Ultrafast Photonic Backpropagation Accelerator for Deep LearningabstractTraining deep learning networks involves continuous weight updates across its many using a backpropagation algorithm (BP). This results in expensive computation and energy overhead during training. Consequently, most deep learning accelerators today employ pre-trained weights and focus only on improving the design of the inference phase. The recent trend is to develop a complete deep learning accelerator by incorporating the training module. Such efforts require an ultra-fast chip architecture for executing the BP algorithm. In this paper, we introduce a novel photonics-based backpropagation accelerator for high performance deep learning training. We present the design for a convolutional neural network, BPhoton-CNN, which incorporates the silicon photonics-based backpropagation accelerator. BPhoton-CNN is a first-of-its-kind photonic and memristor-based CNN architecture for end-to-end training and prediction. We evaluate BPhoton-CNN using a commercial CAD framework (IPKISS) on deep learning benchmark models including LeNet and VGG-Net. The proposed design achieves at least 35× acceleration in training time, 31× improvement in computational efficiency, and 45× energy savings compared to the state-of-the-art designs, without any loss of accuracy. Dharanidhar Dang, Aurosmita Khansama, Rabi N. Mahapatra, Debashis Sahoo |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | BiGNoC: Accelerating Big Data Computing with Application-Specific Photonic Network-on-Chip ArchitecturesabstractIn the era of big data, high performance data analytics applications are frequently executed on large-scale cluster architectures to accomplish massive data-parallel computations. Often, these applications involve iterative machine learning algorithms to extract information and make predictions from large data sets. Multicast data dissemination is one of the major performance bottlenecks for such data analytics applications in cluster computing, as terabytes of data need to be distributed frequently from a single data source to hundreds of computing nodes. To overcome this bottleneck for big data applications, we proposeBiGNoC, a manycore chip platform with a novel application-specific photonic network-on-chip (PNoC) fabric.BiGNoCis designed for big data computing and exploits multicasting in photonic waveguides. For high performance data analytics applications,BiGNoCimproves throughput by up to${{9.9}}\times$while reducing latency by up to 88 percent and energy-per-bit by up to 98 percent over two state-of-the-art PNoC architectures as well as a broadcast-optimized electrical mesh NoC architecture, and a traditional electrical mesh NoC architecture. Sai Vineel Reddy Chittamuru, Dharanidhar Dang, Sudeep Pasricha, Rabi N. Mahapatra |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Islands of heaters: A novel thermal management framework for photonic NoCsabstractSilicon photonics has become a promising candidate for future networks-on-chip (NoCs) as it can enable high bandwidth density and lower latency with traversal of data at the speed of light. But the operation of photonic NoCs (PNoCs) is very sensitive to temperature variations that frequently occur on a chip. These variations can create significant reliability issues for PNoCs. For example, microring resonators (MRRs) which are the building blocks of PNoCs, may resonate at another wavelength instead of their designated wavelength due to thermal variations, which can lead to bandwidth wastage and data corruption in PNoCs. This paper proposes a novel run-time framework to overcome temperature-induced issues in PNoCs. The framework consists of (i) a PID controlled heater mechanism to nullify the thermal gradient across PNoCs, (ii) a device-level thermal island framework to distribute MRRs across regions of temperatures; and (iii) a system-level proactive thread migration technique to avoid on-chip thermal threshold violations and to reduce MRR tuning/trimming power by migrating threads between cores. Our experimental results with 64-core Corona and Flexishare PNoCs indicate that the proposed approach reliably satisfies on-chip thermal thresholds and maintains high network bandwidth while reducing total power by up to 64.1%. Dharanidhar Dang, Sai Vineel Reddy Chittamuru, Rabi N. Mahapatra, Sudeep Pasricha |
ASP-DAC | 1 |
| 2017 | ConvLight: A Convolutional Accelerator with Memristor Integrated Photonic ComputingabstractNeuromorphic computing is a promising candidate to accelerate big data processing. Recently, several attempts have been made to design neuromorphic accelerators for popular machine learning algorithms, such as reservoir computing, deep learning, spiking neurons etc. Deep learning accelerator which involves convolutional neural networks (CNNs) have received widespread attention for their accuracy and efficiency. This paper proposes ConvLight, a novel deep learning accelerator based on memristor integrated photonic computing framework. While the use of on-chip photonic circuits for analog computing is well known, no prior work has demonstrated a full-fledged accelerator based on photonic components. In particular, this paper makes the following novel contributions: (i) A multilayer deep learning architecture design is proposed using compute efficient memristors and photonic components for the first time. (ii) A pipelined design for each CNN layer is presented for maximizing throughput and enabling parallelism across the layers. (iii) Simulation of ConvLight architecture with standard photonic tools for demonstrating the execution of DNN and CNN workloads yielding 25X, 60X, and 40X improvements in computational efficiency, throughput, and energy efficiency (respectively) compared to state-of-the-art design. Dharanidhar Dang, Jyotikrishna Dass, Rabi N. Mahapatra |
HiPC | 1 |
| 2015 | A Multilayered Design Approach for Efficient Hybrid 3D Photonics Network-on-chipabstractIn Chip Multiprocessors, traditional metallic interconnects will soon reach their bandwidth and energy dissipation limits. Photonic NoC (PNoC) is a promising alternative to renew higher performance in the advent of rising number of cores on chip. Efficient PNoC architectures are needed to reduce laser related energy consumption and maintain high performance. In this work we propose a novel sandwich layered approach to design a 3D PNoC architecture that is able to reduce no of hops, cross over points, and no of laser sources using multiplexing techniques. The 3D hybrid PNoC uses high performance 5X5 photonic routers incorporating mode division multiplexing (MDM) along with wavelength division multiplexing (WDM) and time division multiplexing (TDM). Experimental results demonstrates an increase in aggregated bandwidth up to 4x while reducing average energy consumption per router by 83\% as compared to the recently reported results. Dharanidhar Dang, Biplab Patra, Rabi N. Mahapatra |
ACM Great Lakes Symposium on VLSI | 1 |
| 2015 | PID controlled thermal management in photonic network-on-chipabstractThe communication bandwidth and power consumption of network-on-chip (NoC) are going to meet their limits soon because of traditional metallic interconnects. Photonic NoC (PNoC) is emerging as a promising alternative to address these bottlenecks. However, PNoCs are highly susceptible to thermal fluctuations which is highly common in a manycore chip. This paper first introduces a low power, low cost mesh-based PNoC architecture and provides a quantitative analysis of it's power consumption over varying on-chip temperature. The paper then proposes a proportional-integral-derivative (PID) heater mechanism that minimizes the effect of thermal variation on PNoC's performance and power. Experimental results for a 8 × 8 PNoC shows that the proposed technique offers the maximum network-bandwidth considering the thermal effects. Compared to the recently reported results, the proposed design consumes 40% less power and has a temperature variation as low as 1 °C. Dharanidhar Dang, Rabi N. Mahapatra, Eun Jung Kim 0001 |
ICCD | 1 |