Jaydeep P. Kulkarni

dblp:18/1691 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-0258-6776ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 31 · 5 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorSecurity and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 MARU: An ML-Based Framework for Area Estimation from FPGA Resource Usage
abstract
FPGA design evaluation faces significant challenges due to heterogeneous resource reporting across vendors and architectures, which hinders performance comparisons and complicates design space exploration. We present MARU (Machine-learning for Area estimation from Resource Usage), a framework that enables FPGA and ASIC area estimation directly from HLS utilization reports. MARU achieves a mean absolute percentage error of 1.5% and an R² of 0.997 in cross-FPGA, multi-circuit predictions, while reducing estimation time from 2.5 days to just 5 minutes per design. The framework enables three key advancements: 1) unified, area-based estimation across diverse FPGA designs, 2) accelerated HLS design space exploration through real-time prediction, and 3) ASIC migration analysis by estimating equivalent area for predictive cost modeling. MARU bridges the FPGA/ASIC methodology gap by providing both a comparative framework for academic research and a practical tool for industry-scale co-design decisions—whether optimizing FPGA implementations or evaluating ASIC transition feasibility through predictive area modeling.
Tarun Kholay, Anup Ashok Kedilaya, Aman Arora 0001, Jaydeep P. Kulkarni, Lizy Kurian John
FPGA4
2026 Metal Stack Exploration for Front- Versus Back- Side Clock and Signal Allocation for Advanced CMOS PPA Improvements
abstract
This study explores the potential of back-side contact (BSC) technology for enhancing power, performance, and area (PPA) in integrated circuits (ICs). Through metal stack exploration experiments on a 32-bit RISC-V core, the research demonstrates significant PPA gains by optimizing front- and back-side metal (BSM) stack and layer purpose allocation. Key findings include a 2.7% power reduction per added back-side layer for clock routing, a 5.5% frequency improvement with combined front- and back-side clock routing, and up to 14.71% power reduction or 10% frequency increase through metal stack optimization for low-power and high-performance targets, respectively. A similar study on an FPGA was also conducted by optimizing the layer stack configuration to achieve maximum frequency gains of 55.34% or 8.95% in total power consumption, respectively, compared to the baseline of an eight front-side-only layer stack configuration.
Sirish Oruganti, S. S. Teja Nibhanupudi, Anup Ashok Kedilaya, Xiuhao Zhang, Jaydeep P. Kulkarni
IEEE Trans. Very Large Scale Integr. Syst.6
2025 GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-centric Rasterization
abstract
D Gaussian Splatting (3DGS) has emerged as a promising real-time photorealistic radiance field rendering technique. Existing GPU and hardware accelerators face limitations due to insufficient parallelism in sequential rendering pipeline stages and the memory overhead associated with interim results. This paper presents GSAcc, a hardware accelerator co-designed with dataflow to render compressed 3DGS models on edge platforms efficiently. GSAcc enhances 3DGS rendering performance through several key innovations. First, it introduces Gaussian depth speculation, parallelizing preprocessing and sorting tasks. Second, GSAcc adopts a Gaussian-centric dataflow that interleaves preprocessing and rasterization, allowing all rendering steps to execute concurrently without storing intermediate results. Finally, it employs dedicated hardware acceleration to address sorting and rasterization bottlenecks within the optimized dataflow. We implemented and synthesized GSAcc using Intel16 PDK and evaluated its performance on real-world 3DGS scenes. Compared with desktop GPUs, GSAcc achieves up to $1.66 \times 10^{4} \mathrm{x}$ Power-Performance-Area (PPA) improvement as well as 48.7 x energy savings. Additionally, GSAcc outperforms the state-of-the-art hardware accelerator GSCore with up to 2.3x PPA improvement and 2.9x energy savings.
Mengtian Yang, Yipeng Wang 0017, Chieh-Pu Lo, Xiuhao Zhang, Sirish Oruganti, Jaydeep P. Kulkarni
DAC6
2025 SPARK: Sparsity Aware, Low Area, Energy-Efficient, Near-memory Architecture for Accelerating Linear Programming Problems
abstract
Integer Linear Programming (ILP) is an important mathematical approach for solving time-sensitive real-life optimization problems, including network routing, map routing, traffic scheduling, etc. However, the algorithms for solving ILPs are typically sparse and branch-intensive, and not CPU/GPU friendly. In the paper “What could a million cores do to solve Integer programs”, Koch et al. [40] presented data illustrating that Integer Linear Programming (ILP) applications take tens of hours of execution time even on the largest parallel computers. Long execution time is a problem because many real-life applications need a decision in seconds or minutes. The widely used ILP solvers, like Gurobi (optimized for CPUs), perform software-based optimizations to handle the inherent sparsity in ILPs but still do not meet decision threshold because of the limited throughput of CPUs. GPUs are suited for large-sized dot-product compute, however, GPU-based ILP solvers also do not meet decision thresholds as (i) GPU is not sparsity friendly and (ii) GPU incurs thread divergence for branching, resulting in under-utilization of streaming engines and periodic host-GPU interaction. We propose SPARK, a sparsity-aware, reuse-aware, energy-efficient, reconfigurable, near-cache ILP architecture that (i) re-configures the existing L1 cache present in CPUs to perform near-cache acceleration with easy integration into the baseline CPU pipeline with minimal area overhead ($\sim 1.4 \%$ of a CPU), (ii) performs near-cache sparsity detection and sparsity-aware compute, reducing the number of insignificant computations, and data movement energy overheads, (iii) leverages the computational patterns present in algorithms used for solving ILP to realize a reuse-aware architecture, and (iv) is applicable to solving sparse and dense ILPs and LPs (Linear Programs). We observe $15 x / 20 x$, and $152 x / 740 x$ performance/energy improvement over AMD’s Zen3 CPU, and Nvidia’s Tesla v100 GPU for sparse reallife ILPs in Mixed Integer Programming library (MIPLIB 2017). For sparse LPs (non-integer), SPARK achieves 7-17x/103-250x performance/energy improvement over CPU/ GPU indicating SPARK’s broad applicability.
Siddhartha Raman Sundara Raman, Lizy Kurian John, Jaydeep P. Kulkarni
HPCA3
2025 3D SRAM Disaggregation in Advanced CMOS Nodes using Hybrid Bonding Technology
abstract
This paper studies the potential of hybrid bonded Array-under-CMOS (AuC) technology to partition logic and high-performance L1 cache in advanced technology nodes. By decoupling the SRAM bitcells from the logic tier, we achieve independent optimization of both SRAM and logic devices, as well as back-end-of-line (BEOL) interconnects. Heterogeneous integration and BEOL aspect ratio optimization is implemented with different technology nodes to mitigate the delay penalty due to hybrid bond pad staggering. Addressing the performance degradation associated with scaled technology nodes, we investigate the impact of word-line (WL) and bit-line (BL) resistance on SRAM performance. Leveraging the flexibility of decoupled SRAM BEOL and within the AuC technology framework, we explore the sensitivity to WL and BL metal aspect ratios, comparing their performance against a 2D baseline. Our results demonstrate substantial performance improvements in AuC integration through two key approaches: (1) 5% enhancement via Back-End-of-Line (BEOL) optimization, and (2) 25% improvement enabled by heterogeneous integration, achieved by decoupling memory and logic tiers.
Bhawana Kumari, Anurag Swarnkar, Dawit Burusie Abdi, Fernando García-Redondo, James Myers, Julien Ryckaert, Jaydeep P. Kulkarni, Dwaipayan Biswas
ISCAS7
2025 3D IGZO Charge-Coupled Memory DTCO & STCO Analysis for Compute-near-Memory Applications
abstract
The demand for high-capacity and energy-efficient memory solutions has surged in the era of data-centric computing, particularly for Artificial Intelligence (AI) and Machine Learning (ML) workloads. This paper introduces a novel memory architecture leveraging Charge-Coupled Device (CCD) technology, engineered in a sequential-access block memory configuration, to enhance Compute-near-Memory (CnM) systems. We propose an optimized 3D IGZO CCD block memory as an on-chip weight buffer for high-capacity CnM systems. Our approach achieves 2.95−131.26× improvement in area efficiency and 1.32−4.33× improvement in energy efficiency compared to SRAM solutions.
Khakim Akhunov, Hyungrock Oh, Fernando García-Redondo, Yukai Chen, Arvind Sharma, Jiacong Sun, Sahan Gamage, Maarten Rosmeulen, Swaraj Bandhu Mahato, Rishabh Kishore, Subhali Subhechha, Jaydeep P. Kulkarni, Marian Verhelst, Dwaipayan Biswas, Marie Garcia Bardon, Wim Dehaene, Julien Ryckaert
ISCAS13
2025 Photonic Side-Channel Analyzer: Enabling Security-Aware Physical Design Methodology
abstract
74
Meizhi Wang, S. S. Teja Nibhanupudi, Elham Amini, Antonio Saavedra, Daniel Wasserman, Jean-Pierre Seifert, Jaydeep P. Kulkarni
ISPD9
2024 SACHI: A Stationarity-Aware, All-Digital, Near-Memory, Ising Architecture
abstract
Recently there have been efforts to solve difficult computation problems harnessing or drawing inspiration from nature. A prominent example is the use of Ising machines for solving NP-complete problems [1], [23]. Ising machines have evolved from quantum/optical annealers and oscillator-based designs [1] to the recent CMOS-based Von-Neumann [36]/in-memory designs [35]. While prior works have demonstrated the power of Ising machines to solve complex real-world problems, the state-of-the-art Ising accelerators are dedicated accelerators that are useful only for a class of problems, involve complex data converter circuits (ADCs/DACs), are unreliable compared to the rest of the CMOS SoC due to the use of process-variation sensitive/specific embedded memory technologies. In this paper, we present an all-digital Ising architecture realized using repurposing of L1 cache of a CPU. It relies on processing in-memory technology implemented in SRAM. SACHI solves the reliability problems of prior works such as BRIM, eliminates the need for ADCs/DACs, and provides Ising compute acceleration with minor hardware overhead over a CPU pipeline. The novelty of the proposed approach consists of (i) tightly coupled interfacing of the accelerator to the CPU, (ii) reuse/ repurposing of existing hardware to provide acceleration, (iii) ability to achieve higher parallelism than earlier Ising designs due to reuse-aware compute, and (iv) improved performance/energy for a wide variety of large-sized high precision real-life optimization problems using novel compute/mapping strategies. In comparison to BRIM, the proposed all-digital Ising accelerator achieves (i) 36x, 160x, 286x, 300x better performance, (ii) 72x, 79x, 80x, and 75x improved energy, (iii) reuse of 4x, 32x, 200x, and 4000x is observed for asset allocation, molecular dynamics, image segmentation, and traveling salesman respectively.
Siddhartha Raman Sundara Raman, Lizy Kurian John, Jaydeep P. Kulkarni
HPCA3
2024 NEM-GNN: DAC/ADC-less, Scalable, Reconfigurable, Graph and Sparsity-Aware Near-Memory Accelerator for Graph Neural Networks
abstract
Graph neural networks (GNNs) are of great interest in real-life applications such as citation networks and drug discovery owing to GNN’s ability to apply machine learning techniques on graphs. GNNs utilize a two-step approach to classify the nodes in a graph into pre-defined categories. The first step uses a combination kernel to perform data-intensive convolution operations with regular memory access patterns. The second step uses an aggregation kernel that operates on sparse data having irregular access patterns. These mixed data patterns render CPU/GPU-based compute energy-inefficient. Von Neumann based accelerators like AWB-GCN [ 7 ] suffer from increased data movement, as the data-intensive combination requires large data movement to/from memory to perform computations. ReFLIP [ 8 ] performs resistive random access memory based in-memory (PIM) compute to overcome data movement costs. However, ReFLIP suffers from increased area requirement due to dedicated accelerator arrangement, and reduced performance due to limited parallelism and energy due to fundamental issues in ReRAM-based compute. This article presents a scalable (non-exponential storage requirement), DAC/ADC-less PIM-based combination, with (i) early compute termination and (ii) pre-compute by reconfiguring SOC components. Graph and sparsity-aware near-memory aggregation using the proposed compute-as-soon-as-ready (CAR) broadcast approach improves performance and energy further. NEM-GNN achieves ∼80–230x, ∼80–300x, ∼850–1,134x, and ∼7–8x improvement over ReFLIP, in terms of performance, throughput, energy efficiency, and compute density.
Siddhartha Raman Sundara Raman, Lizy Kurian John, Jaydeep P. Kulkarni
ACM Trans. Archit. Code Optim.3
2023 Invited: Buried Power Rails and Back-side Power Grids: Prospects and Challenges
abstract
Buried power rails and back-side power grids are promising technology-scaling boosters for advanced CMOS technology nodes. System-level evaluation of these technologies shows tremendous promise from power-performance-area (PPA), IR drop, and dynamic voltage droop perspective. However, several process, device, and architectural challenges must be addressed to realize the full potential of this technology. This article reviews the advancements and challenges in successfully adopting buried power rail and back-side power grid technology.
S. S. Teja Nibhanupudi, Sirish Oruganti, Rahul Mathur, Meizhi Wang, Jaydeep P. Kulkarni
DAC6
2023 A 118 GOPS/mm23D eDRAM TensorCore Architecture for Large-scale Matrix Multiplication
abstract
The computational demands for recent large transformer- based language models and Neural Radiance Fields (NeRF) have rapidly increased, impacting applications like conversational AI and Mixed Reality (MR). Current accelerator architectures struggle to cope with the vast computational requirements, creating a gap with slowly growing hardware resources. This paper proposes repurposing memory components as high-density computational units, leveraging recent advancements in Back-End-Of-Line (BEOL) transistors and monolithic 3D integration techniques. An ultra-high density monolithic 3D eDRAM is presented as a reconfigurable matrix multiplication unit, co-designed with analog computation circuits, achieving energy efficiency up to 2.41 TOPS/W, performance up to 1.71 TOPS on bfloat16, and compute intensity up to 118 GOPS/mm2. A comprehensive multi-cube(core) architecture is also devised and optimized with bit stationary tensorcore dataflow. We evaluate the proposed architecture on state-of-the-art machine learning models: NeRF and LLaMa-7B, improving the computation density by up to 6.59x and 1.12x compared with GPU and state-of-the-art vector processor designs, respectively.
Mengtian Yang, Yipeng Wang 0017, Jaydeep P. Kulkarni
HiPC3
2023 CoMeFa: Deploying Compute-in-Memory on FPGAs for Deep Learning Acceleration
abstract
Block random access memories (BRAMs) are the storage houses of FPGAs, providing extensive on-chip memory bandwidth to the compute units implemented using logic blocks and digital signal processing slices. We propose modifying BRAMs to convert them to CoMeFa ( Co mpute-in- Me mory Blocks for F PG A s) random access memories (RAMs). These RAMs provide highly parallel compute-in-memory by combining computation and storage capabilities in one block. CoMeFa RAMs utilize the true dual-port nature of FPGA BRAMs and contain multiple configurable single-bit bit-serial processing elements. CoMeFa RAMs can be used to compute with any precision, which is extremely important for applications like deep learning (DL). Adding CoMeFa RAMs to FPGAs significantly increases their compute density while also reducing data movement. We explore and propose two architectures of these RAMs: CoMeFa-D (optimized for delay) and CoMeFa-A (optimized for area). Compared to existing proposals, CoMeFa RAMs do not require changing the underlying static RAM technology like simultaneously activating multiple wordlines on the same port, and are practical to implement. CoMeFa RAMs are especially suitable for parallel and compute-intensive applications like DL, but these versatile blocks find applications in diverse applications like signal processing and databases, among others. By augmenting an Intel Arria 10–like FPGA with CoMeFa-D (CoMeFa-A) RAMs at the cost of 3.8% (1.2%) area, and with algorithmic improvements and efficient mapping, we observe a geomean speedup of 2.55× (1.85×) across microbenchmarks from various applications and a geomean speedup of up to 2.5× across multiple deep neural networks. Replacing all or some BRAMs with CoMeFa RAMs in FPGAs can make them better accelerators of DL workloads.
Aman Arora 0001, Atharva Bhamburkar, Aatman Borda, Tanmay Anand, Rishabh Sehgal, Bagus Hanindhito, Pierre-Emmanuel Gaillardon, Jaydeep P. Kulkarni, Lizy Kurian John
ACM Trans. Reconfigurable Technol. Syst.8
2022 CoMeFa: Compute-in-Memory Blocks for FPGAs
abstract
Block RAMs (BRAMs) are the storage houses of FPGAs, providing extensive on-chip memory bandwidth to the compute units implemented using Logic Blocks (LBs) and Digital Signal Processing (DSP) slices. We propose modifying BRAMs to convert them to CoMeFa (Compute-In-Memory Blocks for FPGAs) RAMs. These RAMs provide highly-parallel compute-in-memory by combining computation and storage capabilities in one block. CoMeFa RAMs utilize the true dual port nature of FPGA BRAMs and contain multiple programmable single-bit bit-serial processing elements. CoMeFa RAMs can be used to compute in any precision, which is extremely important for evolving applications like Deep Learning. Adding CoMeFa RAMs to FPGAs significantly increases their compute density. We explore and propose two architectures of these RAMs: CoMeFa-D (optimized for delay) and CoMeFa-A (optimized for area). Compared to existing proposals, CoMeFa RAMs do not require changing the underlying SRAM technology like simultaneously activating multiple rows on the same port, and are practical to implement. CoMeFa RAMs are versatile blocks that find applications in numerous diverse parallel applications like Deep Learning, signal processing, databases, etc. By augmenting an Intel Arria-10-like FPGA with CoMeFa-D (CoMeFa-A) RAMs at the cost of 3.8% (1.2%) area, and with algorithmic improvements and efficient mapping, we observe a geomean speedup of 2.55x (1.85x), across several representative benchmarks. Replacing all or some BRAMs with CoMeFa RAMs in FPGAs can make them better accelerators of modern compute-intensive workloads.
Aman Arora 0001, Tanmay Anand, Aatman Borda, Rishabh Sehgal, Bagus Hanindhito, Jaydeep P. Kulkarni, Lizy Kurian John
FCCM6
2021 Compute-in-eDRAM with Backend Integrated Indium Gallium Zinc Oxide Transistors
abstract
With rapid growth in data intensive applications, there is an ever-increasing need for energy efficient machine learning/AI hardware accelerators. The performance and the energy efficiency of such accelerators are primarily limited due of massive amount of data movement between processing engines and the off-chip memory. This memory wall bottleneck can be mitigated by performing accelerator specific computations in the memory (CIM) array embedded with the rest of the logic blocks. Multiple embedded memory technologies are being explored to advance CIM designs. Among these, embedded Dynamic Random Access Memory (eDRAM) using backend of the line (BEOL) integrated C-Axis Aligned Crystalline (CAAC) Indium Gallium Zinc Oxide (IGZO) transistors is a promising candidate. IGZO transistor having extremely low leakage when used as an access transistor of the eDRAM bitcell can enable multi-level cell (MLC) eDRAM functionality. Moreover, higher bandwidth can be achieved by 3D stacking multiple layers of BEOL integrated IGZO devices in a monolithic manner improving the CIM performance. In this paper, we analyze various IGZO based eDRAM bitcell topologies and present an IGZO eDRAM CIM architecture. It supports 8-bit inputs/activations and 8-bit signed weights. 2-bit Flash Analog to Digital converter (ADC) is used for MLC weight bit read sensing. A representative neural network model using IGZO eDRAM and peripheral 8-b A/D converters based CIM design achieves 80% Top-1 inference accuracy for the CIFAR- 10 dataset, which is within 3% of ideal software accuracy.
Siddhartha Raman Sundara Raman, Jaydeep P. Kulkarni
ISCAS3
2021 A Systematic Evaluation of EM and Power Side-Channel Analysis Attacks on AES Implementations
abstract
The effectiveness of coarse- and fine-grained electromagnetic (EM) side-channel analysis (SCA) attacks, as well as power SCA attacks, are empirically evaluated on implementations of the Advanced Encryption Standard (AES) algorithm. Coarse-grained EM and power SCA attacks use a single sensor configuration to measure the aggregated EM emanation or power consumption for a large set of encryptions, and then analyze this set of signals to recover all encryption key bytes. In contrast, fine-grained EM SCA attacks first perform high-resolution scans with relatively small probes in multiple orientations to localize on-chip information leakage, and then use a specific probe configuration for each key byte to collect and analyze signals. The fine-grained EM SCA attacks are found to be up to >70× more effective than coarse-grained EM and power SCA attacks when extracting the key from 3 implementations of 128-bit AES. They are constrained, however, by the potentially prohibitive cost of the initial search to identify effective probe configurations. Search protocols, categorized according to the threat model, to reduce this one-time acquisition cost are presented and are found to require ~8–15× fewer measurements compared to an exhaustive search.
Vishnuvardhan V. Iyer, Meizhi Wang, Jaydeep P. Kulkarni, Ali E. Yilmaz
ISI3
2020 M2A2: Microscale Modular Assembled ASICs for High-Mix, Low-Volume, Heterogeneously Integrated Designs
abstract
With CMOS process technology scaling, the mask cost for fabricating nano-scale transistors, contacts, and interconnects has become prohibitively expensive, especially, for low volume designs. Moreover, higher transistor density has resulted in higher design complexity and large-sized die, which has led to an increase in the design cycle time and degradation in the process yield. These challenges are forcing low-volume application-specific integrated circuits (ASICs) toward highly suboptimal field-programmable gate arrays (FPGAs). In this article, we propose a new approach for designing and fabricating high-mix, low-volume heterogeneously integrated ASICs, referred to as Microscale Modular Assembled ASIC (M2A2), consisting of: 1) pick-and-place assembly of prefabricated blocks (PFBs) which utilizes the nano-precision placement capabilities developed in jet-and-flash imprint lithography (J-FIL) and 2) EDA design methodology utilizing unsupervised learning and graph-matching techniques. The EDA methodology leverages existing CAD tool infrastructure for easy adoption into the current EDA ecosystem. The proposed fabrication technology makes use of pick-and-place assembly technique to allow nano-precise assembly of PFBs. The PFBs can be fabricated in advanced process nodes and then knitted together on a wafer substrate. Custom-designed low-cost back-end metal layers can then be created/placed on top of the PFB knitted layer to realize a variety of high-mix, low-volume ASIC designs. M2A2 would allow more flexibility in front-end design by optimal PFB selection and knitting compared to the earlier proposed approaches such as structured ASICs (sASICs). In this article, the performance of M2A2-based designs are compared with different design technologies, such as baseline ASICs, FPGAs, and sASICs at 16 nm, 40 nm, and 130 nm CMOS process nodes. The post-PNR simulation results achieved over 15 IWLS benchmarks show that the proposed M2A2 designs achieve 27.11x -34.89x reduced power-delay-product (PDP) compared to FPGAs, and incur 1.69x -2.36x larger area compared to the baseline ASICs. The M2A2 designs achieve 15%-68.5% smaller area and 8.5%-52% higher performance compared to the sASIC methodologies. Moreover, the key fabrication steps in the proposed M2A2 technology are presented. The experimental fab results along with the proposed EDA flow simulations show promising results for the proposed M2A2 technology. Design tradeoffs and process challenges for large scale deployment of the M2A2 technology are discussed along with their mitigation strategies.
Aseem Sayal, Paras Ajay, Mark W. McDermott, S. V. Sreenivasan, Jaydeep P. Kulkarni
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2019 Editorial TVLSI Positioning - Continuing and Accelerating an Upward Trajectory
abstract
I. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5].
Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber
IEEE Trans. Very Large Scale Integr. Syst.27
2018 Soft-FET: phase transition material assisted soft switching field effect transistor for supply voltage droop mitigation
abstract
Phase Transition Material (PTM) assisted novel soft switching transistor architecture named "Soft-FET" is proposed for supply voltage droop mitigation. By utilizing the abrupt phase transition mechanism in PTMs, the proposed Soft-FET achieves soft switching of the gate input of a logic gate resulting in reduced peak switching current as well as steep current variations (di/dt). In addition, the Soft-FET incurs lower delay penalty across a wide voltage range compared to various baseline Complementary Metal Oxide Semiconductor (CMOS) logic gate variants for the same peak current. We perform a detailed PTM parameter optimization for optimum Soft-FET performance. Soft-FETs when used as power gates achieve ∼20mV lower supply droop and when used as an I/O buffer achieves 46% lower ground bounce with 8.8% improved energy efficiency.
Subrahmanya Teja, Jaydeep P. Kulkarni
DAC2
2018 Impact of Process Variation on Self-Reference Sensing Scheme and Adaptive Current Modulation for Robust STTRAM Sensing
abstract
Spin-Transfer-Torque RAM (STTRAM) is a promising technology for high-density on-chip cache due to low standby power and high speed. However, the process variation of the Magnetic Tunnel Junction (MTJ) and access transistor poses a serious challenge to sensing. Nondestructive sensing suffers from reference resistance variation, whereas destructive sensing suffers from failures due to unoptimized selection of data and reference currents. Furthermore, the sense speed is tightly coupled with the reference/data current requirement. In this work, we study the process variation effect on a self-reference sensing scheme to eliminate bit-to-bit process variation in MTJ resistance. Read current modulation is proposed to overcome the failures due to process variation. Simulation results reveal <0.01% failures at the cost of 9ns sense time and 190uW power consumption.
Seyedhamidreza Motaman, Swaroop Ghosh, Jaydeep P. Kulkarni
ACM J. Emerg. Technol. Comput. Syst.3
2017 Message from the program co-chairs
abstract
It is our great pleasure to welcome you to the 2017 installment of the IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED) in Taipei, Taiwan. This is the 22nd year of the conference, and we are continuing to push the boundaries of low power technology. This year's symposium continues its long tradition of being the premier forum for presentation of research results and industrial experience reports on leading-edge issues in low power design. ISLPED has always been unique in the sense that it brings together researchers and practitioners interested in various aspects of low power design at a single venue and provides them an opportunity to share their perspectives with each other.
Jaydeep P. Kulkarni, Thomas F. Wenisch
ISLPED1
2017 Editorial
abstract
As I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design.
Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber
IEEE Trans. Very Large Scale Integr. Syst.24
2015 A novel slope detection technique for robust STTRAM sensing
abstract
Spin-Torque-Transfer RAM (STTRAM) is a promising technology for high density on-chip cache due to low standby power and high speed. However, the process variation of magnetic tunnel junction (MTJ) and access transistor poses serious challenge to sensing. Nondestructive sensing suffers from reference resistance variation whereas destructive sensing suffers from failures due to unoptimized selection of data and reference currents. We propose a novel slope detection technique to exploit MTJ resistance switching from high to low state using low-overhead sample-and-hold circuit. The proposed sensing technique is destructive in nature and can be combined with double sampling for improved robustness. Simulation results reveal <;0.12% failure under process variation using single sampling (at 0.2% area overhead) and <;0.08% failures with double sampling (at 0.6% area overhead). The overall sense time is found to be 6.8ns.
Seyedhamidreza Motaman, Swaroop Ghosh, Jaydeep P. Kulkarni
ISLPED3
2013 Improving multi-core performance using mixed-cell cache architecture
abstract
Many enterprise and mobile systems must operate within strict power constraints. These systems dynamically trade off performance and power to maximize performance while keeping power within specified limits. In multi-core systems, maximizing the number of active cores within a strict power budget requires minimizing the power per core. Lowering core voltage dramatically reduces power, but compromises cache reliability. Mixed-cell cache architectures, where part of the cache is designed with larger, more robust cells, enable caches to operate reliably at low voltage while minimizing the added cost of larger cells. But mixed-cell caches suffer from poor low-voltage scalability since caches can only use robust cells at low voltage, sacrificing up to 75% of cache capacity. Such capacity reduction strains shared cache resources, leading to significant performance losses. In this paper, we propose a mixed-cell architecture that improves multi-core performance by allowing the use of both robust and non-robust cells. Our mechanisms store modified data only in robust lines by modifying the cache replacement policy and handling writes to non-robust lines. For a multi-core processor, our best mechanism improves performance by 17%, and reduces dynamic power in the L1 data cache by 50% over prior mixed-cell proposals.
Samira Manabi Khan, Alaa R. Alameldeen, Chris Wilkerson, Jaydeep P. Kulkarni, Daniel A. Jiménez
HPCA4
2012 Design for test and reliability in ultimate CMOS
abstract
This session brings together specialists from the DfT, DfY and DfR domains that will address key problems together with their solutions for the 14 nm node and beyond, dealing with extremely complex chips affected by high defect levels, unpredictable and heterogeneous timing behavior, circuit degradation over time, including extreme situations related with the ultimate CMOS nodes, where all processor nodes, routers and links of single-chip massively parallel tera-device processors could comprise timing faults (such as delay faults or clock skews); a large percentage of these parts are affected by catastrophic failures; all parts experience significant performance degradations over time; and new catastrophic failures occur at low MTBF.
Michael Nicolaidis, Lorena Anghel, Nacer-Eddine Zergainoh, Yervant Zorian, Tanay Karnik, Keith A. Bowman, James W. Tschanz, Shih-Lien Lu, Carlos Tokunaga, Arijit Raychowdhury, Muhammad M. Khellah, Jaydeep P. Kulkarni, Vivek De, Dimiter R. Avresky
DATE12
2012 Ultralow-Voltage Process-Variation-Tolerant Schmitt-Trigger-Based SRAM Design
abstract
We analyze Schmitt-Trigger (ST)-based differential-sensing static random access memory (SRAM) bitcells for ultralow-voltage operation. The ST-based SRAM bitcells address the fundamental conflicting design requirement of the read versus write operation of a conventional 6T bitcell. The ST operation gives better read-stability as well as better write-ability compared to the standard 6T bitcell. The proposed ST bitcells incorporate a built-in feedback mechanism, achieving process variation tolerance - a must for future nano-scaled technology nodes. A detailed comparison of different bitcells under iso-area condition shows that the ST-2 bitcell can operate at lower supply voltages. Measurement results on ten test-chips fabricated in 130-nm CMOS technology show that the proposed ST-2 bitcell gives 1.6× higher read static noise margin, 2× higher write-trip-point and 120-mV lower read-Vmincompared to the iso-area 6T bitcell.
Jaydeep P. Kulkarni, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2011 A Read-Disturb-Free, Differential Sensing 1R/1W Port, 8T Bitcell Array
abstract
We propose a read-disturb-free, 1-read/1-write port, 8-transistor (8T) bitcell utilizing differential sensing. The conflicting design requirement of read versus write operation in a conventional 6T SRAM bitcell is eliminated using separate read/write access transistors. A distributed read-access transistor shared across the bitcells of every row enables read-disturb-free differential sensing operation with eight transistors per bitcell. Write-access transistors are upsized to form a diffusion-notch-free layout which would result in improved manufacturability. 1R/1W port nature of the proposed 8T bitcell makes it an attractive choice for the high speed, dense register file (RF) designs. Bitcell failure measurements on 20 test-chips fabricated in 90-nm CMOS technology demonstrate that the proposed differential 8T bitcell shows 220 mV lower read-Vmin, 40 mV lower hold-Vmin, 25 mV higher weak-write voltage compared to the iso-area 6T bitcell at iso-performance. At 600 mV, the proposed 8T bitcell array operates up to 67.2 MHz.
Jaydeep P. Kulkarni, Ashish Goel, Patrick Ndai, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2010 Resilient design in scaled CMOS for energy efficiency
abstract
Traditional processors are designed to guarantee error-free operation under worst-case (1) device & interconnect parameter variations resulting from less than ideal manufacturing process control; (2) static & erratic defects; (3) operating environments such as temperature excursions and voltage droops; (4) critical path activation and path delay degradations due to multiple inputs switching simultaneously in gates containing transistor stacks, or signal coupling from neighboring lines in interconnect paths; (5) speed degradation over the operating lifetime due to transistor aging under voltage, temperature & current stress; (6) early-life failures due to latent defect accelerations; and (7) soft error due to cosmic rays and alpha particle impacts. The voltage-frequency settings for all processors are set based on these infrequently encountered worst-case considerations, even though under typical conditions voltage can be pushed down further or frequency increased without causing errors for most of the processors, thus limiting both energy efficiency and performance in scaled CMOS technologies.
James W. Tschanz, Keith A. Bowman, Muhammad M. Khellah, Chris Wilkerson, Bibiche M. Geuskens, Dinesh Somasekhar, Arijit Raychowdhury, Jaydeep P. Kulkarni, Carlos Tokunaga, Shih-Lien Lu, Tanay Karnik, Vivek De
ASP-DAC8
2010 Analysis of SRAM and eDRAM Cache Memories Under Spatial Temperature Variations
abstract
In scaled technologies, cache memories which are traditionally known as “cold” sections of the chip are expected to occupy a larger die area. Hence, different sections of a cache memory may experience different temperature profiles depending on their proximity to the active logic units such as the execution unit. In this paper, we performed thermal analysis of cache memories under the influence of hot-spots. In particular, 6-transistor (T) static random access memory (SRAM), 8-T SRAM, and embedded dynamic random access memory (eDRAM) cache memories were investigated. Thermal maps of the entire caches were generated using hierarchical compact thermal models while solving the leakage and temperature self-consistently. The 6-T and the 8-T SRAM bitcells were investigated in terms of stability, noise immunity, and performance under temperature variations for various technology nodes. The 3-T micro sense amplifier used in eDRAM cache memories was investigated for its robustness. Thermal-aware circuit design techniques were explored to improve cache stability under thermal gradients. Results show that, for all cache memories, spatial temperature variations have to be considered to achieve the optimal memory design.
Mesut Meterelliyoz, Jaydeep P. Kulkarni, Kaushik Roy 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2009 Device/circuit interactions at 22nm technology node
abstract
As transition is being made into 22nm node, technology considerations and device architectures suitable for such scaled technologies are being explored. To design circuits and systems at scaled nodes, we believe there is a need for technology aware circuit and system design methodology that considers device architecture, and technology challenges to achieve design optimality. In this paper, we discuss the challenges of device-circuit-system design at the 22 nm node and present techniques at different levels of design abstraction to meet these challenges. In particular, we discuss different device options for multi-gate FETs. Logic and memory design using multi-gate FETs is also considered. Finally, we briefly discuss process variation tolerant system design methodologies for such scaled technologies.
Kaushik Roy 0001, Jaydeep P. Kulkarni, Sumeet Kumar Gupta
DAC2
2008 Process variation tolerant SRAM array for ultra low voltage applications
abstract
In this work, we propose a Schmitt Trigger (ST) based differential sensing SRAM bitcell that can operate at ultra-low supply voltage. The proposed Schmitt Trigger SRAM cell addresses the fundamental conflicting design requirement of read versus write operation of a conventional 6T cell. Schmitt Trigger operation gives better read-stability and as well as better writeability compared to the standard 6T cell. The proposed ST bitcell incorporates a built-in feedback mechanism, achieving process variation tolerance - a must for future nano-scaled technology nodes. Measurements on 10 test-chips fabricated in 130nm technology show that the proposed Schmitt Trigger bitcell gives 58% higher read Static Noise Margin (SNM), 2X higher writetrip-point and 120mV lower read Vmin compared to the conventional 6T cell. The ST SRAM array is operational at 150mV of supply voltage.
Jaydeep P. Kulkarni, Keejong Kim, Sang Phill Park, Kaushik Roy 0001
DAC1
2008 Thermal analysis of 8-T SRAM for nano-scaled technologies
abstract
Different sections of a cache memory may experience different temperature profiles depending on their proximity to other active logic units such as the execution unit. In this paper, we perform thermal analysis of cache memories under the influence of hot-spots. In particular, 8-T SRAM bit cell is chosen because of its robust functionality at nano-scaled technologies. Thermal map of entire 8-T SRAM cache is generated using hierarchical compact thermal models while solving the leakage and temperature self consistently. The impact of spatial temperature variations on 8T-SRAM parameters such as local bitline (LBL) sensing delay, noise robustness and bitcell stability are evaluated for 45nm/32nm/22nm bulk CMOS technology nodes. The effectiveness of variable keeper sizing on LBL sensing delay is analyzed. It is predicted that at 22 nm node, the leakage induced temperature rise has severe effects on the 8-T SRAM characteristics.
Mesut Meterelliyoz, Jaydeep P. Kulkarni, Kaushik Roy 0001
ISLPED2
2007 A 160 mV, fully differential, robust schmitt trigger based sub-threshold SRAM
abstract
We propose a novel Schmitt Trigger (ST) based fully differential 10 transistor SRAM (Static Random Access Memory) bitcell suitable for sub-threshold operation. The proposed Schmitt trigger based bitcell achieves 1.56X higher read static noise margin (SNM) (VDD = 400mV) compared to the conventional 6T cell. The robust Schmitt trigger based memory cell exhibits built in process variation tolerance that gives tight SNM distribution across the process corners. It utilizes fully differential operation and hence does not require any architectural changes from the present 6T architecture. At iso-area and iso-read-failure probability the proposed memory bitcell operates at a lower (175mV) VDD with 18% reduction in leakage and 50% reduction in read/write power compared to the conventional 6T cell. Simulation results show that the proposed memory bitcell retains data at a supply voltage of 150mV. Functional SRAM with the proposed memory bitcell is demonstrated at 160mV in 0.13μm CMOS technology.
Jaydeep P. Kulkarni, Keejong Kim, Kaushik Roy 0001
ISLPED1