EDBT 2026 Demo / reviewers in the wild / expert
Ragh Kuttappa
dblp:182/8688
· DBLP profile ↗
18ranked-venue papers
12as first author
8since 2021 · last 2025
0000-0003-1022-2187ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 12 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | System-Level Validation Across Multiple Platforms to build a Robust 2.5D Multi Foundry Chiplet Solution
Srivatsa Rangachar Srinivasa, Dileep Kurian, Paolo A. Aseron, Prerna Budhkar, Vinayak Honkote, Dan Lake, Jaykant Timbadiya, Satish Yada, Sureshbabu Kadavakollu, James Greensky, Gauthaman Murali, Anuradha Srinivasan, Ragh Kuttappa, Tanay Karnik |
ACM Great Lakes Symposium on VLSI | 13 |
| 2025 | Invited Paper: System and Technology Co-Optimization Framework for a Disaggregated System with Passive Die 2.5D IntegrationabstractThe semiconductor industry is steadily shifting toward disaggregated system design to overcome the scalability (yield and cost) limitations of monolithic integration. 2.5D integration offers a compelling pathway for realizing such systems, enabling the assembly of heterogeneous chiplets—compute, memory, analog & I/O, with dense interconnects and high bandwidth. Our previous work showcased experimental results from a multi-foundry chiplet design over a large passive silicon base. There were 20 chip slots (CS) on an interposer that can be configured with compute die (CD) or memory die (MD) chiplets. Building on this foundation, this paper introduces a methodology for system technology co-optimization (STCO) across several vectors. These include varying the number of chip slots, configuring slots with different numbers of MDs and CDs, sweeping the inter-die bandwidth, choosing between different technology nodes for the performance limiting MDs. This work demonstrates how the design and technology choices affect system performance, power, and cost across different workloads, empowering designers to select optimal configurations for their specific needs. Gauthaman Murali, Mudit Bhargava, Shairfe Salahuddin, Archana Pandey, Srivatsa Rangachar Srinivasa, Prerna Budhkar, Ragh Kuttappa, Vinayak Honkote, Prashanth Sakthi, Myung-Hee Na, Tanay Karnik |
ICCAD | 8 |
| 2024 | Design Automation for Charge Recovery LogicabstractThis paper introduces a novel design automation methodology for charge recovery logic (CRL). The proposed methodology combines a novel logic compression algorithm with automatic schematic generation to automate the design process of CRL, enabling power and performance simulations for a large number and variety of CRL circuits. As a measure of the effectiveness of the proposed design flow, automated implementations of CRL equivalents of the LGSynth’91 combinational benchmark circuits are compared with their CMOS counterparts. The results demonstrate a trade-off in power for area: Automatically generated CRL circuits dissipate 51.3% less power on average compared to CMOS equivalents, occupying 54.9% larger area. Yilmaz Ege Gonul, Leo Filippini, Junghoon Oh, Ragh Kuttappa, Scott Lerner, Mineo Kaneko, Baris Taskin |
ISCAS | 4 |
| 2024 | High-Speed Phase-Based ComputingabstractThis work presents the utilization of rotary traveling wave oscillators (RTWOs) to implement an Ising machine. Ising machines utilizing ring oscillators have recently been demonstrated on silicon, for instance, for the solution of a max-cut problem. Rotary traveling wave oscillators scale better in frequency compared to ring oscillators, but have increased power consumption. Phase-based computing principles, implemented with the proposed RTWO-based Ising machines, are prime for high speed phase-based computation. The experiments reveal the proposed RTWO-based Ising machines provide significant reduction (5x) in runtime in the solution of the max-cut problem. The power dissipation is two orders of magnitude higher than the minuscule, low power ring-oscillators but RTWO-based Ising machines sub-linear increase with the demonstrated frequency increase from 2GHz to 32GHz for high speed phase-based computing. The accuracy of the solution is significantly improved as well, as demonstrated with respect to two of the D-Wave solvers (tabu and simulated annealing) acting as the baseline for ring oscillator and RTWO based Ising machines. Nicholas Sica, Ragh Kuttappa, Vinayak Honkote, Baris Taskin |
ISCAS | 2 |
| 2022 | A 0.45 pJ/bit 20 Gb/s/Wire Parallel Die-to-Die Interface with Rotary Traveling Wave OscillatorsabstractIn this work, a die-to-die communication architecture with the integration of resonant clocking is presented. The novelty of the architecture are the rotary traveling wave oscillators, designed across the interposer of a 2. 5D multi-die system to provide a synchronous high frequency clock to all chiplets simultaneously. The transmitter and receiver interface circuits of the architecture benefit from the use of the low power, low skew, multiple phase clock signals across the chiplets. In experimentation, a channel length of 4 mm between transceivers is investigated over a 5 mm $\times 5$ mm silicon interposer. SPICE based simulations with post-layout, parasitic extracted models are performed. The proposed architecture demonstrates 20 Gb/s operation at 0.45 pJ/bit over a 4mm channel at a nominal 1 V supply voltage. The overall clock power of the proposed architecture is 56% lower than prior works at 20 Gb/s. Ragh Kuttappa, Baris Taskin |
ISCAS | 1 |
| 2022 | Resonant Rotary Clock Synchronization with Active and Passive Silicon InterposerabstractRotary traveling wave oscillators (RTWO) are designed to provide a high frequency clock signal through the silicon interposer to multiple chiplets in a heterogeneous 2.5D system. In particular, two different RTWO synchronization topologies are presented: 1) Active interposer RTWO and 2) passive interposer RTWO. The proposed topologies are evaluated across a silicon interposer with a dimension of 42 mm × 20 mm. Each topology is implemented with post-layout, parasitic extracted models for a clock frequency of ≈8 GHz. The performance metrics are presented for clock period, skew, rise time, fall time, and oscillation start-up and settling times across the multi-die system (MDS) with SPICE based simulations. Ragh Kuttappa, Baris Taskin, Vinayak Honkote, Satish Yada, Jainaveen Sundaram, Dileep Kurian, Tanay Karnik, Anuradha Srinivasan |
ISCAS | 1 |
| 2022 | Multiphase Digital Low-Dropout RegulatorsabstractIn this work, multiphase digital low-dropout (MP-DLDO) regulators are designed with resonant rotary clocks (ReRoCs) in order to improve on the tradeoff of conventional DLDOs between current efficiency and transient response speed. The proposed DLDOs are multiphased, coined MP-DLDOs, designed with a clock-gated control technique to provide high current efficiencies along with transient response improvements at GHz frequency levels. The multiple phases within the MP-DLDO are served with ReRoCs that provide: 1) a robust high-speed low-power resonant clock distribution solution for the synchronous elements in the multiphase DLDO architecture and 2) improve the transient response characteristics [dynamic voltage scaling (DVS) speed and voltage ripple] while saving power in the controller circuitry. The proposed MP-DLDOs are distributed across the chip to achieve low voltage ripple. SPICE simulations are performed on post-layout, parasitic-extracted models to evaluate the MP-DLDO architecture on open-source digital cores, with performance metrics that include the voltage ripple reduction, transient response speed improvement, and power savings in the control logic. The proposed MP-DLDO architecture, evaluated on an RISC-V design, demonstrates a DVS speed of 6.5 V/$\mu \text{s}$and an output voltage ripple of 21.1 mV (38% reduction when compared to a conventional DLDO) with a sampling frequency of 2 GHz. Ragh Kuttappa, Selçuk Köse, Baris Taskin |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | Resonant Clock Synchronization With Active Silicon Interposer for Multi-Die SystemsabstractThis paper presents the integration of resonant clocking to multi-die architectures to synchronize individual chiplets connected through an active silicon interposer. The proposed inter-chiplet synchronization through the active silicon interposer rotary oscillator array (ASI-ROA) provides a unitary clock domain to the multiple die (i.e. multiple chiplets) in the package with a very low design overhead. System performance analysis is performed with parasitics-extracted, post-layout simulation models of two different sizes of representative heterogeneous multi-die architectures, each with varying number of RISC-V cores per die. Each RISC-V core of the multi-die package belongs to the unitary clock domain, designed with ASI-ROA to operate at a frequency of 2 GHz. The proposed architecture is investigated for robustness in frequency and skew across the multi-die system (MDS) with SPICE based simulations of post layout models, demonstrating variations of only 80 MHz for a 2 GHz target frequency. The power savings are upto 41% for the overall MDS, compared to an equivalent implementation with a contemporary ADPLL used to synchronize the multiple chiplets over the active interposer. The average clock skew of the completely resonant architecture presented in this work is 8.2 ps. Ragh Kuttappa, Baris Taskin, Scott Lerner, Vasil Pano |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | SnackNoC: Processing in the Communication LayerabstractIn this work, we propose and evaluate a Network-on-Chip (NoC) augmented with light-weight processing elements to provide a lean dataflow-style system. We show that contemporary NoC routers can frequently experience long periods of idle time, with less than 10% link utilization in HPC applications. By repurposing the temporal and spatial slack of the NoC, the proposed platform, SnackNoC, is able to compute linear algebra kernels efficiently within the communication layer with minimal additional resource costs. SnackNoC 'Snack' application kernels are programmed with a producer-consumer data model that uses the NoC slack to store and transmit intermediate data between processing elements. SnackNoC is demonstrated in a multi-program environment that continually executes linear algebra kernels on the NoC simultaneously with chip multiprocessor (CMP) applications on the processor cores. Linear algebra kernels are computed up to 14.2x faster on SnackNoC compared to an Intel Haswell EPx86 processing core. The cost of executing 'snack' kernels in parallel to the CMP applications is a minimal runtime impact of 0.01% to 0.83% due to higher link utilization, and an uncore area overhead of 1.1%. Karthik Sangaiah, Michael Lui, Ragh Kuttappa, Baris Taskin, Mark Hempstead |
HPCA | 3 |
| 2020 | Comprehensive Low Power Adiabatic Circuit Design with Resonant Power ClockingabstractIn this paper, the first comprehensive methodology is presented for design of low power adiabatic circuits inclusive of the adiabatic core design and the power-clock generation. Prior works have focused on either designing adiabatic cores or the power clock generation circuit, only. These non-comprehensive views can misrepresent the performance savings and fail to address the opportunities at integration. In this work, a comprehensive solution is presented that also features a unique innovation for the power clock generation circuit in step-charged circuits designed with rotary traveling wave oscillators (RTWO) and adiabatic frequency dividers. In experimentation, SPICE based simulations are performed at 416 MHz and 330 MHz in the 90 nm technology node and compared to CMOS based implementations, as well as other known power-clock generation techniques. A 32-bit CMOS adder consumes 3.5× more power when compared to the proposed 32-bit ECRL adder operating at a frequency of 416 MHz. Furthermore, 1000 32-bit CMOS adders in parallel consumes 3.4× more power when compared to 1000 32-bit ECRL adders in parallel designed with the proposed architecture at a frequency of 416 MHz. Ragh Kuttappa, Steven Khoa, Leo Filippini, Vasil Pano, Baris Taskin |
ISCAS | 1 |
| 2020 | FinFET - Based Low Swing Rotary Traveling Wave OscillatorsabstractFinFET based, low swing clocking with rotary traveling wave oscillators (RTWO) is presented in this paper. It is shown that the low-swing clock signal generation by RTWOs is very effective, thanks to FinFETs accommodating high frequency operation and voltage scaling better than planar CMOS transistors. Low swing clocks are aimed at lowering the power dissipation of the clock networks, while maintaining the full voltage operation of non-clock components (such as logic and memory). In this work shows that robust low swing (LS) RTWOs are designed with FinFET based technologies. To this end, SPICE simulations are performed on the ISPD'10 clock benchmark circuits operating at 2.25 GHz and 3 GHz in the 16 nm FinFET technology node. LS-RTWO based designs are compared to an all digital phase locked loop (ADPLL) based designs operating at the same target frequency. At 3 GHz, the LS-RTWO consumes 36% lower power with 42.7dB better phase noise @10 MHz on comparison to corresponding ADPLL based designs. Ragh Kuttappa, Baris Taskin |
ISCAS | 1 |
| 2019 | Low Swing - Low Frequency Rotary Traveling Wave OscillatorsabstractThis paper presents the design and implementation of low-swing, low-frequency rotary traveling wave oscillators (RT-WOs). Low-swing rotary clocks are designed to operate with low-swing D flip-flops. A methodology is proposed to perform dynamic frequency scaling for integer division ratios of 3 to n to generate low-frequency rotary clocks at target frequencies. SPICE-based experiments are performed on the ISPD'10 benchmark circuits operating at 100 MHz, 200 MHz, and 500 MHz. The low-swing, low-frequency resonant rotary clock based designs are compared against full-swing and low-swing traditional designs operating at the same frequency and voltage level. The results show that the low-swing rotary clock based ISPD'10 designs at 500 MHz consume 61% and 43% lower power on comparison to the traditional full-swing and low-swing designs, respectively, where the traditional designs are implemented with a PLL, and a bounded-skew tree. Ragh Kuttappa, Scott Lerner, Leo Filippini, Baris Taskin |
ISCAS | 1 |
| 2019 | Robust Low Power Clock Synchronization for Multi-Die SystemsabstractA novel clock generation and distribution network is proposed for multi-die architectures connected through an active silicon interposer. The proposed clock network generates and distributes a resonant clock through the active silicon interposer between dies, with each die served through resonant local clock trees. The proposed active silicon interposer rotary oscillator array (AI-ROA) serves to establish a unitary clock domain, providing constant phase and magnitude clock sources to the multiple die (i.e. multiple chiplets) in the package. Analysis is performed with multiple ARM CORTEX M0 cores per die of a homogeneous multi-die package architecture. Each M0 core of the multi-die package belongs to the unitary clock domain, designed with AI-ROA to operate at a frequency of 1 GHz. The multiple die are designed in the 28 nm technology node and the active interposer is designed in the 65 nm technology node. SPICE based simulations of post-layout models provides analysis and evaluation of the proposed architecture for performance metrics under process, voltage, and temperature variations. In particular, performance metrics are reported for 1) power consumption in comparison to PLL based architectures designed and synthesized with an industrial tool, 2) robustness against process variations, and 3) clock skew across the cores throughout the multiple die. Ragh Kuttappa, Baris Taskin, Scott Lerner, Vasil Pano, Ioannis Savidis |
ISLPED | 1 |
| 2019 | 3D NoCs with active interposer for multi-die systemsabstractAdvances in interconnect technologies for system-in-package manufacturing have re-introduced multi-chip module (MCM) architectures as an alternative to the current monolithic approach. MCMs or multi-die systems implement multiple smaller chiplets in a single package. These MCMs are connected through various package interconnect technologies, such as current industry solutions in AMD's Infinity Fabric, Intel's Foveros active interposer, and Marvell's Mochi Interconnect. Although MCMs improve manufacturing yields and are cost-effective, additional challenges on the Network-on-Chip (NoC) within a single chiplet and across multiple chiplets need to be addressed. These challenges include routing, scalability performance, and resource allocation. This work introduces a scalable MCM 3D interconnect infrastructure called "MCM-3D-NoC" with multiple 3D chiplets connected through an active interposer. System-level simulations of MCM-3D-NoC are performed to validate the proposed architecture and provide performance evaluation of network latency, throughput, and EDP. Vasil Pano, Ragh Kuttappa, Baris Taskin |
NOCS | 2 |
| 2018 | Low Frequency Rotary Traveling Wave OscillatorsabstractA methodology is presented in order to automate the selection of a design style for low frequency operation of ultra low power rotary traveling wave oscillators (RTWOs). The methodology systemically analyzes the power dissipation profiles of design styles with two priori known resonant frequency dividers, a static and a dynamic frequency divider. The proposed methodology is very effective in leading to a unique design style solution for systems with one frequency target. For systems with multiple frequency targets, the proposed methodology not only leads to a unique solution but also provides guidance on re-selection of frequency targets for improved performance. A 1.2 GHz RTWO is used to generate a target frequency of 135 MHz, while consuming 0.82× the power of an RTWO for that frequency using a static frequency divider. A dynamic frequency divider is used to generate two target frequencies of 137 MHz and 204 MHz using a master clock of 1.2 GHz, while consuming 0.75× and 0.8× the power respectively for that RTWO frequency. Ragh Kuttappa, Baris Taskin |
ISCAS | 1 |
| 2017 | Reconfigurable threshold logic gates using optoelectronic capacitorsabstractThis paper investigates the integration of optoelectronic devices with CMOS to implement reconfigurable threshold logic gates for Boolean functions. The weight of the optoelectronic device can be altered by changing the optical power which is used to reconfigure the threshold logic (TL) gate. These novel reconfigurable gates are called the optoelectronic capacitor based TL (OECTL) gates. The OECTL gates are designed for i) simplistic AND/NAND gates and OR/NOR gates with large fan-in and ii) linearly separable Boolean functions that can be reconfigured to other linearly separable Boolean functions, constrained in reconfiguration by the specifics of TL operation. SPICE simulations in 65nm bulk CMOS technology with a Verilog-A model for the optoelectronic capacitor demonstrate i) AND/NAND gates and OR/NOR gates are 2× faster as fan-in goes above 3 and consumes low power and ii) Boolean functions can be reconfigured with 0.58× smaller delay and 0.46× less power of standard CMOS design. Ragh Kuttappa, Lunal Khuon, Bahram Nabet, Baris Taskin |
DATE | 1 |
| 2017 | Stability of Rotary Traveling Wave Oscillators under process variations and NBTIabstractResonant rotary clocking is a low-power clocking technology for multi-phase clock generation in GHz frequency range. In this paper, Rotary Traveling Wave Oscillators (RTWOs) are analyzed under process variations and negative bias temperature instability (NBTI) at the 90nm technology node. The analysis is focused on 1) variations in the physical geometries of the rotary ring, 2) inter and intra-die transistor variations, 3) power supply fluctuation and 4) NBTI. Monte-Carlo based analysis are performed to study the effects of process variations and transistor aging on the operating frequency and power consumption of the rotary ring at a temperature of 110° C. SPICE based simulations show natural robustness against process variations, and NBTI. Ragh Kuttappa, Leo Filippini, Scott Lerner, Baris Taskin |
ISCAS | 1 |
| 2016 | Comparative analysis of robustness of spin transfer torque based look up tables under process variationsabstractSpin Transfer Torque (STT) switching realized using a Magnetic Tunnel Junction (MTJ) device has shown great potential for low power and non-volatile storage. A prime application of MTJs is in building non-volatile Look Up Tables (LUT) used in reconfigurable logic. Such LUTs use a hybrid integration of CMOS transistors and MTJ devices. This paper discusses the reliability of STT based LUTs under transistor and MTJ variations in nano-scale. The sources of process variations include both the CMOS device related variations and the MTJ variations. A key part of the STT based LUTs is the sense amplifier needed for reading out the MTJ state. We compare the voltage and current based sensing schemes in terms of the power, performance, and reliability metrics. Based on our simulation results in a 16nm CMOS, for the same total device area, the voltage mode sensing scheme offers 75% lower failure rates under threshold voltage (Vth) variations, 4.9X higher tolerance to MTJ resistance variations, 19% less delay, and 64% lower active power compared to the current sensing scheme. Ragh Kuttappa, Houman Homayoun, Hassan Salmani, Hamid Mahmoodi |
ISCAS | 1 |