Niranjan Kulkarni

dblp:34/3710 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 1 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Integrated circuit design · 43% Energy-efficient computing · 30% Electronic design automation · 27%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Integrated circuit design › clocking
clock skew
0.922023
A New Approach to Clock Skewing for Area and Power Optimization of ASICs Using Differential Flipflops and Local Clocking · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
A Clock Skewing Strategy to Reduce Power and Area of ASIC Circuits · DAC 2017
Integrated circuit design
clocking
0.712023
A New Approach to Clock Skewing for Area and Power Optimization of ASICs Using Differential Flipflops and Local Clocking · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
Energy-efficient computing
low-power design
0.712023
A New Approach to Clock Skewing for Area and Power Optimization of ASICs Using Differential Flipflops and Local Clocking · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
Energy-efficient computing › low-power design
power and area optimization
0.712023
A New Approach to Clock Skewing for Area and Power Optimization of ASICs Using Differential Flipflops and Local Clocking · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
Electronic design automation › physical design
clock network synthesis
0.312017
A Clock Skewing Strategy to Reduce Power and Area of ASIC Circuits · DAC 2017
Integrated circuit design
low-power circuit design
0.312017
A Clock Skewing Strategy to Reduce Power and Area of ASIC Circuits · DAC 2017
Electronic design automation
physical design
0.312017
A Clock Skewing Strategy to Reduce Power and Area of ASIC Circuits · DAC 2017
Electronic design automation › logic synthesis
boolean function representation
0.112011
Identification of Threshold Functions and Synthesis of Threshold Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Electronic design automation › logic synthesis
decision diagrams
0.112011
Identification of Threshold Functions and Synthesis of Threshold Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Electronic design automation
logic synthesis
0.112011
Identification of Threshold Functions and Synthesis of Threshold Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Electronic design automation › logic synthesis › threshold logic synthesis
threshold function identification
0.112011
Identification of Threshold Functions and Synthesis of Threshold Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Electronic design automation › logic synthesis
threshold logic synthesis
0.112011
Identification of Threshold Functions and Synthesis of Threshold Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011

Methods — techniques the papers use, named apart from their topics

differential flipflops · 0.7deliberate clock skewing · 0.7source-target identification algorithm · 0.3differential flip-flop design · 0.3max literal factor tree · 0.1heuristic search · 0.1binary decision diagram · 0.1
YearPublicationVenuePosition
2023 A New Approach to Clock Skewing for Area and Power Optimization of ASICs Using Differential Flipflops and Local Clocking
abstract
A new design methodology for reducing the area and power of standard cell ASICs that uses a combination of differential flipflops and a method of deliberate clock-skewing, called local clocking (LC), is described. LC introduces clock skew without the use of extra buffers in the clock network. This is done by having some flipflops, called sources, generate clock signals for other flipflops, called targets. The method involves two key features: 1) the design of a new differential flipflop, referred to as KVFF, that is functionally identical to a double-latch edge-triggered$D$flipflop, but in addition, produces a completion signal that is a skewed version of its input clock, which is used to clock other flipflops and 2) an efficient algorithm that identifies the sources and targets involved in the new clocking scheme, with the objective of reducing area and power. These are reduced because deliberate skew introduces extra slack on the logic cones that feed the target flipflops, which is exploited by synthesis tools to reduce area and power. Furthermore, the area and power overhead of conventional methods of introducing skew, e.g., buffers, is eliminated. LC is shown to result in significant improvements in area, power, and wirelength for several, publicly available, benchmark circuits for 65 nm bulk CMOS and 28 nm FDSOI technologies. For 65 nm, the average improvement in area, power and wirelength were 27.7%, 13.4%, and 21.0%, respectively. For 28 nm FDSOI the average improvement in area, power, and wirelength were 20.0%, 10.5%, and 30.5%, respectively. In addition, this article demonstrates how LC can be used to eliminate hold time violations.
Ankit Wagle, Niranjan Kulkarni, Sarma B. K. Vrudhula
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 A Clock Skewing Strategy to Reduce Power and Area of ASIC Circuits
abstract
A new method for reducing power and area of standard cell ASICs is described. The method is based on deliberately introducing clock skew without the use of extra buffers in the clock network. This is done by having some flipflops, called sources, generate clock signals for other flipflops, called targets. The method involves two key features: (1) the design of new differential flipflop, referred to as KVFF, that is functionally identical to a master-slave edge-triggered D flipflop, but in addition, produces an completion signal that is a skewed version of its input clock, which is used to clock other flipflops; and (2) an efficient algorithm that identifies the sources and targets involved in the new clocking scheme, with the objective of reducing area and power. These are reduced because deliberate skew introduces extra slack on the logic cones that feed the target flipflops, which is exploited by synthesis tools to reduce area and power. In addition, the overhead of conventional methods of introducing skew, e.g. buffers, is eliminated. Using commercial tools, significant improvements in power and area are shown on placed and routed netlists of several circuits.
Niranjan Kulkarni, Aykut Dengi, Sarma B. K. Vrudhula
DAC1
2016 Reducing Power, Leakage, and Area of Standard-Cell ASICs Using Threshold Logic Flip-Flops
abstract
In this paper, we describe a new approach to reduce dynamic power, leakage, and area of application-specified integrated circuits, without sacrificing performance. The approach is based on a design of threshold logic gates (TLGs) and their seamless integration with conventional standard-cell design flow. We first describe a new robust, standard-cell library of configurable circuits for implementing threshold functions. Abstractly, the threshold gate behaves as a multi-input, single-output, edge-triggered flip-flop, which computes a threshold function of the inputs on the clock edge. The library consists of a small number of cells, each of which can compute a set of complex threshold functions, which would otherwise require a multilevel network. The function realized by a given threshold gate is determined by how signals are mapped to its inputs. We present a method for the assignment of signals to the inputs of a threshold gate to realize a given threshold function. Next, we present an algorithm that replaces a subset of flip-flops and portions of their logic cones in a conventional logic netlist, with threshold gates from the library. The resulting circuits, with both conventional and TLGs (called hybrid circuits), are placed and routed using commercial tools. We demonstrate significant reductions (using postlayout simulations) in power, leakage, and area of the hybrid circuits when compared with the conventional logic circuits, when both are operated at the maximum possible frequency of the conventional design.
Niranjan Kulkarni, Jae-sun Seo, Sarma B. K. Vrudhula
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Design of threshold logic gates using emerging devices
abstract
This article explores the use of threshold logic for reducing the power, delay, and/or area of digital logic circuits. We first describe the architecture of a differential threshold logic gate (TLG) using conventional MOSFETs. A TLG of a given number of inputs can be configured to realize a set of threshold functions by simply connecting the appropriate signals to its inputs. One characteristic of the proposed architecture for a TLG is the increased sensitivity to process variations (device mismatch) and noise. Problems due to device mismatch can be mitigated by proper cell design and optimization. The increased sensitivity to noise makes it difficult to scale the supply voltage of a TLG. We show a simple solution which involves integrating RRAMs within the TLG circuit, to achieve robust, low voltage and energy efficient operation. The third circuit implementation referred to as a spintronic threshold logic (STL) cell uses an STT-MTJ device as a intrinsic threshold logic gate. An STL cell is an very compact structure that can realize a large number of threshold functions, many of which would require a multilevel network of conventional CMOS logic gates.
Sarma B. K. Vrudhula, Niranjan Kulkarni
ISCAS2
2015 Fast and robust differential flipflops and their extension to multi-input threshold gates
abstract
In this paper, we describe two single input threshold gate (TLG) designs, which are functionally equivalent to differential flipflops. We present a detailed comparison of TLGs with two well established D-flipflop designs. The comparisons are done in both 65nm and 28nm commercial processes. We compare total delay, which is defined as the sum of setup delay and clock to output delay. We also show a comparison of tolerance against noise and process variation between the different designs. The two proposed designs are found to be as robust as an existing D-flipflop from both commercial standard cell libraries. Meanwhile, they are 33% and 25% faster in 65nm post-layout and 25% and 22% faster in 28nm post-layout compared to a commercial D-flipflop. The energy delay product of two proposed designs is 46% and 25% smaller in the 28nm process post layout simulation. Single input threshold gates can also be extended to multi-input threshold gates. We compare the total delay, power and leakage of a three-input threshold gate as well as a seven-input threshold gate with those of the equivalent (Complementary Metal Oxide Semiconductor) CMOS counterpart, which shows 52% 48% improvement in 65nm and 35% speed improvement in the 28nm process.
Niranjan Kulkarni, Joseph Davis, Sarma B. K. Vrudhula
ISCAS2
2014 A fast, energy efficient, field programmable threshold-logic array
abstract
Threshold-logic gates have long been known to result in more compact and faster circuits when compared to conventional AND/OR logic equivalents [1], However, threshold logic based design has not entered the mainstream design technology (neither custom ASIC nor FPGA) due to the lack of efficient and reliable gate implementations and the necessary infrastructure for automated synthesis and physical design. This paper is a step toward addressing this gap. We present the architecture of a novel programmable logic array, referred to as Field Programmable Threshold-Logic Array (FPTLA), in which the basic cells are differential mode threshold-logic gates (DTGs). Each individual DTG cell is a clock edge-triggered circuit that computes a threshold-logic function. A DTG can be programmed to implement different threshold logic functions by routing appropriate signals to their inputs. This reduces the number of SRAMs inside the logic blocks by about 60% compared to conventional CLBs, without adding any significant overhead in the routing infrastructure. Since a DTG is essentially a multi-input, edge-triggered flipflop that computes a threshold function, a network of DTGs forms a nano-pipelined circuit. The advantages of such a network are demonstrated on a set of deeply pipelined datapath circuits implemented on FPTLAs and conventional FPGAs using the well established FPGA design framework VTR (Verilog To Routing) and VPR (Versatile Place and Route) [2]. The results indicate that an FPTLA can achieve up to 2X improvement in delay for nearly the same energy and logic area compared to the conventional LUT based FPGA. Although differential mode circuits can potentially be more sensitive to process variations, FPTLAs can be made robust to such variations without sacrificing their improved energy efficiency and performance over FPGAs.
Niranjan Kulkarni, Sarma B. K. Vrudhula
FPT1
2014 Spintronic Threshold Logic Array (STLA) - A compact, low leakage, non-volatile gate array architecture
Nishant Nukala, Niranjan Kulkarni, Sarma B. K. Vrudhula
J. Parallel Distributed Comput.2
2012 Minimizing area and power of sequential CMOS circuits using threshold decomposition
abstract
This paper describes the design of a standard cell library of differential mode threshold gates, referred to as a Threshold Logic Latch or TLL, and new threshold function identification and decomposition methods to map a conventional logic network consisting of logic gates and flipflops, into a hybrid network that consists of both TLLs and conventional logic gates. After logic synthesis and physical design (placement and routing) using a commercial 65nm LP (low power) library, and commercial design tools, the hybrid circuits are shown to have up to 35% less dynamic power, about 50% less leakage power and around 37% less area when compared to the corresponding conventional design operated at the same (peak) frequency.
Niranjan Kulkarni, Nishant Nukala, Sarma B. K. Vrudhula
ICCAD1
2011 Identification of Threshold Functions and Synthesis of Threshold Networks
abstract
This paper presents a new and efficient heuristic procedure for determining whether or not a given Boolean function is a threshold function, when the Boolean function is given in the form of a decision diagram. The decision diagram based method is significantly different from earlier methods that are based on solving linear inequalities in Boolean variables that derived from truth tables. This method's success depends on the ordering of the variables in the binary decision diagram (BDD). An alternative data structure, and one that is more compact than a BDD, called a max literal factor tree (MLFT) is introduced. An MLFT is a particular type of factoring tree and was found to be more efficient than a BDD for identifying threshold functions. The threshold identification procedure is applied to the MCNC benchmark circuits to synthesize threshold gate networks.
Tejaswi Gowda, Sarma B. K. Vrudhula, Niranjan Kulkarni, Krzysztof S. Berezowski
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2009 A Novel Approach to Cell Formation
Radim Belohlávek, Niranjan Kulkarni, Vilém Vychodil
ICFCA2