Rakshith Saligram

dblp:132/9014 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0002-7436-9375ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 CryoBoost: A 40nm Cryogenic-CMOS Matrix Multiplication Accelerator for Energy Efficient Computing
abstract
This paper proposes a cryogenic Matrix Multiplication (MATMUL) accelerator chip to address the exponential increase in energy consumption in AI model training. The accelerator, designed in a 40nm CMOS process, leverages a Liquid Nitrogen-based cooling system, operating from 300K to 77K. The design comprises 10 processing elements (PEs) operating on 4x4 matrices at INT8 precision, interconnected by a core data ring. The PEs use a load-store architecture with a 128-bit very long instruction word (VLIW)-based controller. The paper presents a detailed characterization of the supply voltage versus frequency, performance, and power across different temperatures. The results indicate a significant reduction in power at cryogenic temperatures, with up to 45.7% reduction in power at 77K compared to 300K at iso performance. The maximum energy efficiency increases from 1.975GHz/W at 300K to 2.497GHz/W at 77K, yielding a 26.4% gain. This translates to up to 20% reduction in training energy for Large Language Models.
Rakshith Saligram, Samuel Spetalnick, Brian Crafton, Muya Chang, Alec Nordlund, Joshua Gess, Ruslan Nagimov, Arijit Raychowdhury
ACM Great Lakes Symposium on VLSI1
2026 Jitter Reduction in Voltage Controlled Oscillators for Clocking at Cryogenic Temperature
Rakshith Saligram, Suman Datta, Arijit Raychowdhury
ISCAS1
2024 Cooling the Chaos: Mitigating the Effect of Threshold Voltage Variation in Cryogenic CMOS Memories
abstract
Cryogenic CMOS is a promising technology for high performance computing due to its improvement in subthreshold slope, carrier mobilities and reduced wire resistance. The threshold voltage (Vth) increase at 77K can be mitigated by metal gate work function (PHIG) engineering to achieve matched off current (Ioff) further enhancing the device performance allowing us to operate at very low supply voltage thereby reducing the Energy Delay Product (EDP). However, the effect of variation on noise margins of static random access memories (SRAM) deploying these matched Ioff devices is very prominent especially at low supply voltages (Vdd) limiting its scaling. In this work, we propose a framework to perform Vth retargeting for cryogenic SRAM for improving noise margins in high performance cryogenic SRAM cells under variation. The proposed framework comprises of a Monte-Carlo engine which performs statistical analysis and DC characterization and a backend processing engine to analyze noise margins and tune the PHIG. To demonstrate the framework, we use calibrated 14nm FinFET models at 300K and 77K. First, we analyze the logic blocks using iso-Ioff devices, which yield up to 3x improvement in delay at iso-energy and a 4.5x reduction in energy at iso-delay. Next, we study the effect of Vth variation on the device currents. Finally, the framework is deployed to tune PHIG, and results show that it can enhance the noise margins by 23%, 31% and 19% for hold, read and write operations respectively at 77K compared to iso-Ioff devices. Further, a 1kb SRAM array has been simulated using iso-Ioff tuned peripherals and framework tuned SRAM cells, and it shows 5.4x reduction in read/write energies along with 1.2x delay reduction and better noise margins at 77K compared to 300K.
Rakshith Saligram, Amol D. Gaidhane, Yu Cao 0001, Suman Datta, Arijit Raychowdhury
ISLPED1
2024 Cryogenic Operation of Computing-In-Memory based Spiking Neural Network
abstract
This paper introduces a Computing-In-Memory based Spiking Neural Network (SNN) architecture for cryogenic operation of CMOS (Cryo-SNN). The paper demonstrates design strategies to improve energy efficiency of Cryo-SNN by coupling low-voltage operation at cryogenic temperature with innovative design of neuron circuits optimized for cryogenic conditions. By exploiting the enhanced device characteristics of 14 nm FinFET transistors at cryogenic temperatures, our architecture outlines critical adaptations to SNN components for optimal functionality in extreme environments. The circuit simulation using measurement calibrated 14nm FinFET models shows that a Cryo-SNN designed for MNIST classification operates with 4.54X improved energy-delay-product (EDP) over room temperature operation while maintaining similar accuracy. Further, the paper designs an optimized SNN architecture for autonomous health monitoring of miniaturized satellites at cryogenic temperature consuming less than 1mW of power.
Laith A. Shamieh, Wei-Chun Wang 0001, Shida Zhang, Rakshith Saligram, Amol D. Gaidhane, Yu Cao 0001, Arijit Raychowdhury, Suman Datta, Saibal Mukhopadhyay
ISLPED4
2023 Cryogenic CMOS as an Enabler for Low Power Dynamic Logic
abstract
Cryogenic High-Performance Computing (HPC) has gained traction for server and cloud systems which demand large scale, energy efficient and fast computing systems. Dynamic logic satisfies these goals and at cryogenic temperature, its inherent problems of charge leakage are readily addressed thanks to the exponential reduction in subthreshold leakage currents. Fully Depleted Silicon on Insulator (FDSOI) devices present an additional “dial” of back gate biasing which opens multiple design options and solutions to further enhance the circuit power performance metrics. In this paper we present a solution - selective back gate biasing ― applied to dynamic and domino logic circuits, to increase their energy efficiency and/or performance. With the proposed method, we show up to 48% decrease in delay at constant energy and 41% decrease in energy at constant delay at 77K compared to 300K. We further scale up the circuit to a radix-4 sparse-2 64 bit adder where the proposed technique increases energy efficiency by 53% and/or performance by 56% going from 300K to 77K.
Rakshith Saligram, Suman Datta, Arijit Raychowdhury
ISLPED1
2022 Design Space Exploration of Interconnect Materials for Cryogenic Operation: Electrical and Thermal Analyses
abstract
With Copper (Cu) Interconnects causing performance bottleneck at single nanometer nodes due to increase in resistivity size effects viz., grain boundary scattering and surface scattering, there has always been scavenging for alternate interconnect materials. Although the Cu resistivity value decreases at cryogenic temperature, the problems continue to persist. In this work, we study three alternate interconnect materials specifically for 77K High Performance Compute applications. We select the materials based on their resistivity value at 77K for 7nm node computed using Fuchs-Sondheimer-Mayadas-Shatzkes (FS-MS) models. We analyze the delay of the interconnects, understand repeater insertion as a function of wire length, evaluate repeater count and energy at system level and perform IR drop analysis by showing through detailed analytical models that Ru, Rh and Al can provide appreciable improvements over Cu at 77K. The delay of interconnects reduces by 1-3.75% for Ru, 1.5-7.25% for Rh and 4.4-17.8% for Al across the BEOL stack while repeater counts decrease by 10%, 15% and 37% for Ru, Rh and Al respectively at 77K. We investigate thermal and reliability aspects of interconnect design including electromigration, Joule Heating and maximum allowed current densities again proving that Ru (9%), Rh (18%) and Al (63%) outperform Cu at 77K. Finally, we study the effects of various Low-k dielectric materials on the interconnect capacitance and thermal behavior for Cu as well as three alternate materials noting that, even though thermal conductivity of dielectrics decrease at 77K, the Joule Heating will not be as worse as one might expect.
Rakshith Saligram, Suman Datta, Arijit Raychowdhury
IEEE Trans. Circuits Syst. I Regul. Pap.1
2020 A Model Study of Multilevel Signaling for High-Speed Chiplet-to-Chiplet Communication in 2.5D Integration
abstract
The quest for high yield has motivated significant advancement in 2.5D integrated circuits, where chiplets are integrated on a silicon interposer or a package substrate with high-speed parallel communication among them. These channels for 2.5D integrated systems need to have high data bandwidth per unit length (also called shoreline-BW-density and measured in Gb/s/mm) and lower energy per bit area (measured in pJ/b). Typically, NRZ signalling is used but achieving higher data rates continues to be a major challenge. In this paper we explore PAM4 as an alternative to NRZ for signalling the channels. Simulations show that we can achieve up to 63% more energy-efficiency and 27% higher BW density for 2.5D integrated systems.
Rakshith Saligram, Ankit Kaul, Muhannad S. Bakir, Arijit Raychowdhury
VLSI-SOC1