Tharini Suresh

dblp:397/9358 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0004-1511-4673ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Focus Session: Accelerating Diffusion Models for Generative AI Applications with Silicon Photonics
abstract
Diffusion models have revolutionized generative AI, with their inherent capacity to generate highly realistic state-of-the-art synthetic data. However, these models employ an iterative denoising process over computationally intensive layers such as UNets and attention mechanisms. This results in high inference energy on conventional electronic platforms, and thus, there is an emerging need to accelerate these models in a sustainable manner. To address this challenge, we present a novel silicon photonics-based accelerator for diffusion models. Experimental evaluations demonstrate that our photonic accelerator achieves at least 3× better energy efficiency and 5.5× throughput improvement compared to state-of-the-art diffusion model accelerators.
Tharini Suresh, Salma Afifi, Sudeep Pasricha
DATE1
2026 MCFlash: bulk bitwise processing in 3D NAND with dynamic sensing and multi-level encoding
Habib Ur Rahman, Tharini Suresh, Sudeep Pasricha, Biswajit Ray
J. Supercomput.2
2025 TCFlash: In-Flash Bulk Bitwise Processing via Dynamic Sensing and TLC Encoding in 3D NAND
abstract
This paper presents TCFlash, a practical and immediately deployable technique for executing bulk bitwise operations directly within commercial off-the-shelf (COTS) 3D NAND flash chips, using only standard user-mode commands. TCFlash enables in-place bitwise computation by combining logical data encoding of triple level-cell (TLC) storage with dynamic read reference voltage shifting. We demonstrate TCFlash across multiple 3D TLC NAND devices spanning both floatinggate and charge-trap technologies from two major vendors. To our knowledge, this is the first on-chip demonstration of error-free bitwise operations in 3D NAND. Experimental evaluation across vertical layers in the 3D NAND stack shows that TCFlash achieves zero raw bit error rate (RBER) for two-operand OR, AND, and XNOR operations, and RBER is below 0.006% for NAND, NOR, and XOR. Additionally, for the first time, we also demonstrate simultaneous three-operand bitwise operations with RBER below 0.008%.
Habib Ur Rahman, Tharini Suresh, Sudeep Pasricha, Biswajit Ray
ICCD2
2025 Sustainable Acceleration of Generative AI Neural Network Models with Silicon Photonics
abstract
Generative AI models such as Generative Adversarial Networks (GANs) and Diffusion Models (DMs), have demonstrated remarkable capabilities in producing high-quality synthetic data for applications ranging from image synthesis and medical imaging to data augmentation. However, the complex model architectures and unique computational operations pose significant challenges for traditional electronic accelerators. To address energy/sustainability bottlenecks with conventional electronic hardware, we present a novel silicon photonic accelerator targeting both GANs and DMs. Experimental evaluations show that our photonic accelerator achieves at least 2.18× lower energy consumption and at least 4.4× throughout improvement compared to several state-of-theart CPU, GPU, FPGA, ReRAM, and ASIC-based accelerators.
Tharini Suresh, Salma Afifi, Sudeep Pasricha
ICCD1