EDBT 2026 Demo / reviewers in the wild / expert
Sunil P. Khatri
dblp:15/884
· DBLP profile ↗
154ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0001-7134-9929ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 143 · 4 first-author · 15 since 2021Software engineering, systems software and programming languages · 16 · 1 since 2021Theory of computation · 4Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 3Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Flash-Based Reconfigurable QCNN Inference Accelerator for Edge ApplicationsabstractThis paper presents a novel in-memory flash-based Quantized Convolutional Neural Network (QCNN) inference accelerator Integrated Circuit (IC), which achieves extremely high throughput and extremely low power, energy, and memory requirements. In this manner, our design significantly improves upon state-of-the-art inference accelerators operating at the edge. Our accelerator IC is designed as a field-programmable, reconfigurable dataflow architecture, which allows a wide variety of trained QCNN models to be programmed to our IC. Once programmed with a model, our IC performs inference withzeroaccess to off-chip memory. In our accelerator, the core computing structures are arrays of flash Field Effect Transistors (FETs) integrated in three dimensions (3D). We use 3D NOR flash stacks, thus achieving an area efficient design with the promise of further improvements as the depth of 3D NOR flash process technology advances. Our flash arrays perform dot product operations in the analog voltage domain, using an extremely low supply voltage (V DDlo, which is 100 mV in our design), resulting in a very high energy efficiency. The digital I/Os of our arrays are in the nominal supply voltage domain for our process (V DDhi, which is 0.8 V in our design), which allows them to be easily stored in registers without the need for level shifters. Neuron weights in our QCNN accelerator are stored in-memory in our flash arrays. These arrays perform dot product operations between a vector of all neuron inputs (in theV DDhidomain) and a vector of corresponding weights (stored in the flash devices as programmedVTHvalues), resulting in an output in theV DDlodomain. We show that for inference on ImageNet with a quantized AlexNet model, our IC achieves extremely low inference latency (28.7 ms at most), power consumption (5.82 mW at most), and energy consumption (4.35 µJ at most), making it especially suited for use in resource-constrained edge devices, such as portable medical devices used for disease diagnosis and ECG signal classification. We prove that our design is robust to Process, Voltage, and Temperature (PVT) variations, and report a drop in Top-1 (Top-5) ImageNet validation accuracy of at most 0.90% (1.07%) across PVT variations. We compare our QCNN accelerator against recent state-of-the-art in-memory neural network accelerators and report an energy efficiency that is at least 15.2µ higher and a parameter density that is at least 3.32µ higher than the best existing design. Kyler R. Scott, Sunil P. Khatri |
IEEE Trans. Computers | 2 |
| 2025 | A Novel Mixed-Signal Flash-based Finite Impulse Response (FFIR) Filter for IoT ApplicationsabstractIn this paper, we present a novel mixed-signal flash-based Finite Impulse Response (FFIR) filter architecture for IoT applications. The FFIR filter is scalable in that it can implement any filter with up to a provisioned maximum number of taps. Our FFIR filter utilizes flash transistors, a type of non-volatile memory (NVM) device, to perform analog computations in the current domain, achieving low power, energy, and area requirements. This design is well-suited for the Internet of Things (IoT) applications and other scenarios where resources are highly constrained. Our FFIR filter consists of several Flash-based Coefficient Multipliers (FCMs). The FIR coefficients of each FCM are stored in its constituent flash transistors, with the threshold voltage (Vt) of the flash transistors serving as a proxy for the filter coefficients. Furthermore, the impact of process or voltage variations is mitigated by precisely tuning the Vt of the flash transistors. The tuning of the Vt of the flash transistors can be performed either in the factory by the manufacturer (to negate process variations), or by the user in the field (to negate voltage variations or aging effects). We evaluate the tolerance of FFIR filters to manufacturing variations through Monte Carlo analysis, demonstrating robustness to process and VDD variations. Our FFIR design achieves a significant improvement over previous approaches. Compared to Digital FIR (DFIR) filters operating at the fastest frequency, we reduce the average of power, energy, and area by 4.05×, 1.95×, and 6.06×, while achieving an average peak signal-to-noise ratio (PSNR) of 38.04 dB and an average effective number of bits (ENOB) of 8.87 bits. In addition, we compare our FFIR filter with state-of-the-art Analog FIR (AFIR) filters as well. Our designs demonstrate significantly improved performance of at least 1.3×, 5.3×, and 18.5× in terms of energy per tap, area, and latency, respectively, when compared with the best among 4 recently published AFIR works. Cheng-Yen Lee, Sunil P. Khatri, Ali Ghrayeb |
ASP-DAC | 2 |
| 2024 | Self-Adaptive Physics Informed Neural Network for Paper Insulation Degree of Polymerization PredictionabstractThis paper proposes a self adaptive physics informed neural network (SAPINN) model to predict the degree of polymerization (DP) of oil-impregnated paper insulation to quantify the level of degradation and the remaining useful lifetime. The prediction is performed based on historical DP values and the corresponding prediction time step, which are used as input data points to the proposed model. The DP mathematical model is used to constrain the training phase of the AI-model through a weighted sum loss function. The weights of this loss function are adjusted for each epoch through a self-adaptive weighting method to determine the relative importance of the data component and the mathematical model throughout the training by defining these weights as trainable parameters. The trained model is then tested using different datasets which are not part of the training phase. The training and testing datasets are generated synthetically through an algorithm that considers the deviation from the ideal DP degradation curve and incorporates actual measurement noise. The performance of the proposed SAPINN is compared to the baseline PINN and NN (in the absence of physics) to highlight the importance of embedding the mathematical model and the self adaptation algorithm, and theses experiments demonstrate that SAPINN significantly enhances the DP prediction. Alamera Nouran Alquennah, Mohammad AlShaikh Saleh, Ali Ghrayeb, Haitham Abu-Rub, Shady S. Refaat, Mohammed Abdullah Al-Hajri, Sunil P. Khatri |
IECON | 7 |
| 2024 | An ASIC Accelerator for QNN With Variable Precision and Tunable Energy EfficiencyabstractThis paper presents TULIP, a new architecture for a variable precision Quantized Neural Network (QNN) inference. It is designed with the goal of maximizing energy efficiency per classification. TULIP is constructed by arranging a collection of unique processing elements (TULIP-PEs) in a single instruction multiple data (SIMD) fashion. Each TULIP-PE contains binary neurons that are interconnected using multiplexers. Each neuron also has a small dedicated local register connected to it. The binary neurons are implemented as standard cells and used for implementing threshold functions, i.e., an inner-product and thresholding operation on its binary inputs. The neurons can be reconfigured with a single change in the control signals to implement all the standard operations used in a QNN. This paper presents novel algorithms for implementing the operations of a QNN on the TULIP-PEs in the form of a schedule of threshold functions. TULIP was implemented as an ASIC in TSMC 40nm-LP technology. A QNN accelerator that employs a conventional MAC-based arithmetic processor was also implemented in the same technology to provide a fair comparison. The results show that TULIP is 30-50X more energy-efficient than an equivalent design, without any penalty in performance, area, or accuracy. Furthermore, TULIP achieves these improvements without using traditional techniques such as voltage scaling or approximate computing. Finally, the paper also demonstrates how the run-time trade-off between accuracy and energy efficiency is done on the TULIP architecture. Ankit Wagle, Gian Singh, Sunil P. Khatri, Sarma B. K. Vrudhula |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | A Mixed-Signal Quantized Neural Network Accelerator Using Flash TransistorsabstractThis paper presents a mixed-signal architecture for implementing Quantized Neural Networks (QNNs) using flash transistors, to achieve extremely high throughput with extremely low power, energy and memory requirements. Its low resource utilization makes our design especially suited for use in edge devices. The network weights are stored in-memory using flash transistors, and neurons perform operations in the analog current domain. Our design can be programmed with any QNN whose hyperparameters (the number of layers, filters, or filter size, etc) do not exceed the maximum provisioned. Once the flash devices are programmed with a trained model and our IC is given an input, our architecture performs inference with zero access to off-chip memory. We demonstrate the robustness of our design under current-mode non-linearities arising from process, voltage, and temperature (PVT) variations. We test validation accuracy on the ImageNet dataset, and show that our IC suffers only 0.71% and 0.92% reduction in classification accuracy for Top-1 and Top-5 outputs, respectively. Our implementation achieves between$2.1\times $and$125\times $better energy efficiency than previous NVM-based QNN accelerators. Our approach provides layer partitioning and neuron sharing options, which allow us to trade off latency, power, and area amongst each other. Kyler R. Scott, Cheng-Yen Lee, Sunil P. Khatri, Sarma B. K. Vrudhula |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | An Extremely Low-voltage Floating Gate Artificial NeuronabstractThis paper presents an artificial neuron design, based on a floating gate (flash) transistor array, that operates with an extremely low supply voltage domain$VDD_{lo}$(which we chose to be 100 mV). Neuron inputs and outputs are in another domain$VDD_{hi}$(which we chose to be 0.8 V). Since$VDD_{hi}$is the nominal supply voltage for our process, input and output data can be stored in registers without the need for level shifters. Neuron weights are stored in-memory in a flash transistor array, using a novel differential conductance encoding scheme. Our neuron performs operations in the analog voltage domain. We compare our design against a recent flash-based mixed-signal neuron design, which has reported the best results (in terms of power and energy, with insignificant error) thus far. We demonstrate that our neuron design beats the previous design in static power consumption$\mathbf{(52.9}\times$less), static energy consumption$\mathbf{(1.41}\times$less), and layout area (44% less). The goal of this work is to design a low-power, low-energy neuron which could be used to implement an entire Neural Network (NN). Our focus is on the neuron, and not the whole NN. Kyler R. Scott, Sunil P. Khatri |
ISCAS | 2 |
| 2023 | Scaled Population Division for Approximate ComputingabstractIn this paper we present an approximate division scheme for Scaled Population (SP) arithmetic, a technique that improves on the limitations of stochastic computing (SC). SP arithmetic circuits are designed (a) to perform all operations with a constant delay, and (b) they use scaling operations to help reduce errors compared to SC circuits. As part of this work, we also present a method to correlate two SP numbers with a constant delay. We compare our SP divider with SC dividers, as well as fixed-point dividers (in terms of area, power and delay). Our 512-bit SP divider has a delay (power) that is 0.08× (0.06x×) that of the equivalent fixed-point binary divider. Compared to a equivalent SC divider, our power-delay-product is 13× better. Kunal Bharathi, Sunil P. Khatri, Jiang Hu 0001 |
ISLPED | 2 |
| 2023 | A Digital Low Dropout (LDO) Voltage Regulator Using Pseudoflash TransistorsabstractIn this article, we present a pseudoflash-based digital low dropout (Digital LDO) voltage regulator. The novelty of our pseudoflash-based Digital LDO (PFD-LDO) voltage regulator lies in the fact that we use pseudoflash (or alternately, flash) transistor subarrays for voltage regulation. By changing the threshold voltage (and thereby, the ON resistance) of these transistors, we can use the same design to meet different regulator specifications. The threshold voltage can be programed either at the factory by the manufacturer or in the field by the user. This gives the manufacturer the ability to offer a family of LDO regulators with a single design, a significant economic advantage. In addition, aging effects and temperature variations are effectively erased since the threshold voltage of the pseudoflash (or flash) transistors can be tuned to a fine degree in the field. Similarly, process variations can be canceled after manufacturing in the factory. These advantages are absent in traditional LDO regulators. Our design uses two subarrays. A coarse subarray is used to reduce the recovery time and output voltage overshoot/undershoot, while a fine subarray regulates the output voltage, minimizing the output voltage ripple. Unlike state-of-the-art LDO regulators, our design can realize multiple specifications with the same circuit. For example, we demonstrate that the$V_{\mathrm{ out}}$of the proposed PFD-LDO regulator can range from 0.7 to$1.7 {V}$when the supply voltage$V_{\mathrm{ IN}}$ranges from 0.8 to$1.8 {V}$, using the same circuit design. Over this voltage range, the proposed PFD-LDO regulator achieves$V_{\mathrm{ shoot}}< 573 \text {mV}$,$t_{\mathrm{ rec}}< 0.50 \mu \text {s}$, and$V_{\mathrm{ ripple}}< 7.4 \text {mV}$when the$I_{\max }$ranges from 15 to 250 mA. Cheng-Yen Lee, Sunil P. Khatri |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | A Flash-based Current-mode IC to Realize Quantized Neural NetworksabstractThis paper presents a mixed-signal architecture for implementing Quantized Neural Networks (QNNs) using flash transistors to achieve extremely high throughput with extremely low power, energy and memory requirements. Its low resource consumption makes our design especially suited for use in edge devices. The network weights are stored in-memory using flash transistors, and nodes perform operations in the analog current domain. Our design can be programmed with any QNN whose hyperparameters (the number of layers, filters, or filter size, etc) do not exceed the maximum provisioned. Once the flash devices are programmed with a trained model and the IC is given an input, our architecture performs inference with zero access to off-chip memory. We demonstrate the robustness of our design under current-mode non-linearities arising from process and voltage variations. We test validation accuracy on the ImageNet dataset, and show that our IC suffers only 0.6% and 1.0% reduction in classification accuracy for Top-1 and Top-5 outputs, respectively. Our implementation results in a$\sim \boldsymbol {50}\times$reduction in latency and energy when compared to a recently published mixed-signal ASIC implementation, with similar power characteristics. Our approach provides layer partitioning and node sharing possibilities, which allow us to trade off latency, power, and area amongst each other. Kyler R. Scott, Cheng-Yen Lee, Sunil P. Khatri, Sarma B. K. Vrudhula |
DATE | 3 |
| 2022 | TD3lite: FPGA Acceleration of Reinforcement Learning with Structural and Representation OptimizationsabstractReinforcement learning (RL) is an effective and increasingly popular machine learning approach for optimization and decision-making. However, modern reinforcement learning techniques, such as deep Q-learning, often require neural network inference and training, and therefore are computationally expensive. For example, Twin-Delay Deep Deterministic Policy Gradient (TD3), a state-of-the-art RL technique, uses as many as 6 neural networks. In this work, we study the FPGA-based acceleration of TD3. To address the resource and computational overhead due to inference and training of the multiple neural networks of TD3, we propose TD3lite, an integrated approach consisting of a network sharing technique combined with bitwidth-optimized block floating-point arithmetic. TD3lite is evaluated on several robotic benchmarks with continuous state and action spaces. With only 5.7% learning performance degradation, TD3lite achieves 21 ×and 8 ×speedup compared to CPU and GPU implementations, respectively. Its energy efficiency is 26 ×of the GPU implementation. Moreover, it utilizes ~ 25 - 40% fewer FPGA resources compared to a conventional sinale-precision floating-point representation of TD3. Chan-Wei Hu, Jiang Hu 0001, Sunil P. Khatri |
FPL | 3 |
| 2022 | A Flash-based Digital to Analog Converter for Low Power ApplicationsabstractThis paper presents a novel technique for digital to analog conversion, which uses flash transistors embedded in a traditional current-steering digital to analog converter (DAC) architecture. Our design utilizes flash transistors as programmable, tunable current sources. Our DAC achieves extremely low latency, area, and power, making it especially suited for the internet of things (IoT) and other applications with a highly constrained resource budget. In addition, the use of flash transistors allows a user to cancel errors due to process/voltage variations and chip aging. This tuning can be performed in the fabrication facility, or in-field. We perform a Monte-Carlo analysis to demonstrate the performance of our DAC in a real-world context. Compared to recent DACs intended for use in IoT devices (or for applications that are highly resource constrained), we significantly improve error metrics. We reduce chip area by 4.3× compared with the smallest DAC intended for IoT. Additionally, we improve throughput by 55 ×and energy per conversion by 33×, compared with the fastest and the most energy efficient DAC intended for IoT, respectively. The proposed 12-bit DAC achieves a throughput of 100 MS/s, an energy per conversion of 36.6 fJ. Based on Monte Carlo analysis, we report a maximum INL (DNL) of 1.242 LSB (0.757 LSB), an average INL (DNL) of 0.286 LSB (0.088 LSB), an ENOB of 11.80 bits, and an SFDR of 82.89 dB. Kyler R. Scott, Sunil P. Khatri |
ICCD | 2 |
| 2022 | A Novel ASIC Design Flow Using Weight-Tunable Binary Neurons as Standard CellsabstractIn this paper, we describe a design of a mixed-signal circuit for an binary neuron (a.k.a perceptron, threshold logic gate) and a methodology for automatically embedding such cells in ASICs. The binary neuron, referred to as an FTL (flash threshold logic) uses floating gate or flash transistors whose threshold voltages serve as a proxy for the weights of the neuron. Algorithms for mapping the weights to the flash transistor threshold voltages are presented. The threshold voltages are determined to maximize both the robustness of the cell and its speed. The performance, power, and area of a single FTL cell are shown to be significantly smaller (79.4%), consume less power (61.6%), and operate faster (40.3%) compared to conventional CMOS logic equivalents. Also included are the architecture and the algorithms to program the flash devices of an FTL. The FTL cells are implemented as standard cells, and are designed to allow commercial synthesis and P&R tools to automatically use them in synthesis of ASICs. Substantial reductions in area and power without sacrificing performance are demonstrated on several ASIC benchmarks by the automatic embedding of FTL cells. The paper also demonstrates how FTL cells can be used for fixing timing errors after fabrication Ankit Wagle, Gian Singh, Sunil P. Khatri, Sarma B. K. Vrudhula |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | NIST-Lite: Randomness Testing of RNGs on an Energy-Constrained PlatformabstractRandom Number Generators (RNGs) are an essential part of many embedded applications and are used for security, encryption, and built-in test applications. The output of RNGs can be tested for randomness using the well-known NIST statistical test suite. Embedded applications using True Random Number Generators (TRNGs) need to test the randomness of their TRNGs periodically, because their randomness properties can drift over time. Using the full NIST test suite is unpracticed for this purpose, because the full NIST test suite is computationally intensive, and embedded systems (especially real-time systems) often have stringent constraints on the energy and runtime of the programs that are executed on them. In this paper, we propose novel algorithms to select the most effective subset of the NIST test suite, which works within specified runtime and energy budgets. To achieve this, we rank the NIST tests based on multiple metrics, including p-value/Time, p-value/Energy, p-value/Time2and p-value/Energy2. Based on the total runtime or energy constraint specified by the user, our algorithms proceed to choose a subset of the NIST tests using this rank order. We call this subset of NIST tests as NIST-Lite. Our algorithms also take into account the runtime and energy required to generate the random sequences required (on the same platform) by the NIST-Lite tests. We evaluate the effectiveness of our method against the full NIST test suite (referred to as NIST-Full) and also against a greedily chosen subset of the NIST test suite (referred to as NIST-Greedy). We explore different variants of NIST-Lite. On average, using the same input sequences, the p-value obtained for the 4 best variants of NIST-Lite is 2× and 7× better than the p-value of NIST-Full and NIST-Greedy respectively. NIST-Lite also achieves 158× (204×) runtime (energy) reduction compared to the NIST-Full. Further, we study the performance of NIST-Lite and NIST-Full for deterministic (non-random) input sequences. For such sequences, the pass rate of the NIST-Lite tests is within 16% of the pass rate of NIST-Full on the same sequences, indicating that our NIST-Lite tests have a similar diagnostic ability as NIST-Full. Cheng-Yen Lee, Kunal Bharathi, Joellen Lansford, Sunil P. Khatri |
ICCD | 4 |
| 2021 | CIDAN: Computing in DRAM with Artificial NeuronsabstractNumerous applications such as graph processing, cryptography, databases, bioinformatics, etc., involve the repeated evaluation of Boolean functions on large bit vectors. In-memory architectures which perform processing in memory (PIM) are tailored for such applications. This paper describes a different architecture for in-memory computation called CIDAN, that achieves a 3X improvement in performance and a 2X improvement in energy for a representative set of algorithms over the state-of-the-art in-memory architectures. CIDAN uses a new basic processing element called a TLPE, which comprises a threshold logic gate (TLG) (a.k.a artificial neuron or perceptron). The implementation of a TLG within a TLPE is equivalent to a multi-input, edge-triggered flipflop that computes a subset of threshold functions of its inputs. The specific threshold function is selected on each cycle by enabling/disabling a subset of the weights associated with the threshold function, by using logic signals. In addition to the TLG, a TLPE realizes some non-threshold functions by a sequence of TLG evaluations. An equivalent CMOS implementation of a TLPE requires a substantially higher area and power. CIDAN has an array of TLPE(s) that is integrated with a DRAM, to allow fast evaluation of any one of its set of functions on large bit vectors. Results of running several common in-memory applications in graph processing and cryptography are presented. Gian Singh, Ankit Wagle, Sarma B. K. Vrudhula, Sunil P. Khatri |
ICCD | 4 |
| 2021 | Hardware Acceleration of Hash Operations in Modern MicroprocessorsabstractModern microprocessors contain several special function units (SFUs) such as specialized arithmetic units, cryptographic processors, etc. In recent times, applications such as cloud computing, web-based search engines, and network applications are widely used, and place new demands on the microprocessor. Hashing is a key algorithm that is extensively used in such applications. Hashing can reduce the complexity of search and lookup from O(N) to O(N/n), where n bins are used. Hashing is typically performed in software. Thus, implementing a hardware-based hash unit on a modern microprocessor would potentially increase performance significantly. In this article, we propose a novel hardware hash unit (HU) design for use in modern microprocessors, at the microarchitecture level and at the circuit level. First, we present the design of the HU at the microarchitecture level. We simulate the HU to compare its performance with a software-based hash implementation. We demonstrate a significant speedup (up to 15×) for the HU. Furthermore, the performance scales elegantly with increasing database size and application diversity, without increasing the hardware cost. Second, we present the circuit design of the HU for use in modern microprocessors, using a 45nm technology. Our proposed hardware hash unit is based on the use of a content-addressable memory (CAM) to implement each bin of the hash table. We simulate the HU circuit and compare it with a traditional CAM design. We demonstrate an average power reduction of 5.48× using the HU over the traditional CAM. Also, we show that the HU can operate at a maximum frequency of 1.39 GHz (after accounting for process, voltage and temperature (PVT) variations and accounting for wiring parasitics). Furthermore, we present the delay, power and area trade-offs of the HU design with varying hash table sizes. Abbas A. Fairouz, Monther Abusultan, Viacheslav V. Fedorov, Sunil P. Khatri |
IEEE Trans. Computers | 4 |
| 2020 | Scaled Population Arithmetic for Efficient Stochastic ComputingabstractWe propose a new Scaled Population (SP) based arithmetic computation approach that achieves considerable improvements over existing stochastic computing (SC) techniques. First, SP arithmetic introduces scaling operations that significantly reduce the numerical errors as compared to SC. Experiments show accuracy improvements of a single multiplication and addition operation by 6.3× and 4.0×, respectively. Secondly, SP arithmetic erases the inherent serialization associated with stochastic computing, thereby significantly improves the computational delays. We design each of the operations of SP arithmetic to take O(1) gate delays, and eliminate the need of serially iterating over the bits of the population vector. Our SP approach improves the area, delay and power compared with conventional stochastic computing on an FPGA-based implementation. We also apply our SP scheme on a handwritten digit recognition application (MNIST), improving the recognition accuracy by 32.79% compared to SC. Sunil P. Khatri, Jiang Hu 0001, Frank Liu 0001 |
ASP-DAC | 2 |
| 2020 | Scaled Population Subtraction for Approximate ComputingabstractIn this paper we present Scaled Population Subtraction to fill a void in Scaled Population arithmetic. Scaled population (SP) arithmetic is a scheme that is inspired by stochastic computing (SC), a non-conventional approximate computing method that is well known for its simplicity, area efficiency and resilience to bit errors. SP arithmetic reduces the numerical errors compared to SC and also solves the serialization limitation of SC, since it is designed to have a O(1) gate delay. Previously, SP was limited to only addition and multiplication and did not have a way to perform subtraction. This paper introduces the basic SP subtraction idea, followed by a detailed study of several ways that the basic design can be improved to reduce the computational error. Our best SP design significantly improves the error compared to our basic SP subtraction idea (reducing it by 32.3%). We also study the trade-off between design complexity of the SP subtractor against output error. Also, our implementation of the SP subtractor exhibits an improved delay, power and area compared to fixed point realizations with the same size. Kunal Bharathi, Jiang Hu 0001, Sunil P. Khatri |
ICCD | 3 |
| 2020 | A Configurable BNN ASIC using a Network of Programmable Threshold Logic Standard CellsabstractThis paper presents Tulip, a new architecture for a binary neural network (BNN) that uses an optimal schedule for executing the operations of an arbitrary BNN. It was constructed with the goal of maximizing energy efficiency per classification. At the top-level, Tulip consists of a collection of unique processing elements (TULIP-PEs) that are organized in a SIMD fashion. Each Tulip- Peconsists of a small network of binary neurons, and a small amount of local memory per neuron. The unique aspect of the binary neuron is that it is implemented as a mixed-signal circuit that natively performs the inner-product and thresholding operation of an artificial binary neuron. Moreover, the binary neuron, which is implemented as a single CMOS standard cell, is reconfigurable, and with a change in a single parameter, can implement all standard operations involved in a BNN. We present novel algorithms for mapping arbitrary nodes of a BNN onto the TULIP-PEs. Tulip was implemented as an ASIC in TSMC 40nm-LP technology. To provide a fair comparison, a recently reported BNN that employs a conventional MAC-based arithmetic processor was also implemented in the same technology. The results show that Tulip is consistently 3X more energy-efficient than the conventional design, without any penalty in performance, area, or accuracy. Ankit Wagle, Sunil P. Khatri, Sarma B. K. Vrudhula |
ICCD | 2 |
| 2019 | A Memory-Efficient Markov Decision Process Computation Framework Using BDD-based Sampling RepresentationabstractAlthough Markov Decision Process (MDP) has wide applications in autonomous systems as a core model in Reinforcement Learning, a key bottleneck is the large memory utilization of the state transition probability matrices. This is particularly problematic for computational platforms with limited memory, or for Bayesian MDP, which requires dozens of such matrices. To mitigate this difficulty, we propose a highly memory-efficient representation for probability matrices using Binary Decision Diagram (BDD) based sampling, and develop a corresponding (Bayesian/classical) MDP solver on a CPU-GPU platform. Simulation results indicate our approach reduces memory by one and two orders of magnitude for Bayesian/classical MDP, respectively. Sunil P. Khatri, Jiang Hu 0001, Frank Liu 0001 |
DAC | 2 |
| 2019 | Threshold Logic in a FlashabstractThis paper describes a novel design of a threshold logic gate (a binary perceptron) and its implementation as a standard cell. This new cell structure, referred to as flash threshold logic (FTL), uses floating gate (flash) transistors to realize the weights associated with a threshold function. The threshold voltages of the flash transistors serve as proxy for the weights. An FTL cell can be equivalently viewed as a multi-input, edge-triggered flipflop which computes a threshold function on a clock edge. Consequently it can used in automatic synthesis of ASICs. The use of flash transistors in the FTL cell allows programming of the weights after fabrication, thereby preventing discovery of its function by a foundry or by reverse engineering. This paper focuses on the design and characteristics of the FTL cell. We present a novel method for programming the weights of an FTL cell for a specified threshold function using a modified perceptron learning algorithm. The algorithm is further extended to select weights to maximize the robustness of the design in the presence of process variations. The FTL circuit was designed in 40nm technology and simulations with layout-extracted parasitics included, demonstrate significant improvements in area (79.7%), power (61.1%), and performance (42.5%) when compared to the equivalent implementations of the same function in conventional static CMOS design. Weight selection targeting robustness is demonstrated using Monte Carlo simulations. The paper also shows how FTL cells can be used for fixing timing errors after fabrication. Ankit Wagle, Gian Singh, Sunil P. Khatri, Sarma B. K. Vrudhula |
ICCD | 4 |
| 2019 | Fast, Ring-Based Design of 3-D Stacked DRAMabstractAs computer memory increases in size and processors continue to get faster, the memory subsystem becomes a bottleneck to system performance. To mitigate the relatively slow dynamic random access memory (DRAM) chip speeds, a new generation of 3-D stacked DRAM is being developed, with lower power consumption and higher bandwidth. This paper proposes the use of 3-D ring-based data fabrics for fast data transfer between the chips in the 3-D stacked DRAM. The ring-based data fabric uses a fast standing wave oscillator to clock its transactions. With a fast clocking scheme and multiple channels sharing the same bus, more channels are utilized while significantly reducing the number of through-silicon vias. Our memory architecture using a ring-based scheme (MARS) can effectively trade off power, throughput, and latency to improve the system performance for different application spaces. Experimental results show that our ring-based data fabric can reduce read latencies and power consumption. MARS variants can deliver better latency (up to ~4χ), power (up to ~8χ), and performance per watt (up to ~8χ) over high bandwidth memory. We also compare our approach with Wide I/O, which is designed for power-constrained systems. MARS variants provide better latency (up to ~8χ) with similar performance per watt. Andrew J. Douglass, Sunil P. Khatri |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | A Homomorphic Encryption Scheme Based on Affine TransformsabstractAs more businesses and consumers move their information storage to the cloud, the need to protect sensitive data is higher than ever. Using encryption, data access can be restricted to only authorized users. However, with standard encryption schemes, modifying an encrypted file in the cloud requires a complete file download, decryption, modification, and upload. This is cumbersome and time-consuming. Recently, the concept of homomorphic computing has been proposed as a solution to this problem. Using homomorphic computation, operations may be performed directly on encrypted files without decryption, hence avoiding exposure of any sensitive user information in the cloud. This also conserves bandwidth and reduces processing time. In this paper, we present a homomorphic computation scheme that utilizes the affine cipher applied to the ASCII representation of data. To the best of the authors» knowledge, this is the first use of affine ciphers in homomorphic computing. Our scheme supports both string operations (encrypted string search and concatenation), as well as arithmetic operations (encrypted integer addition and subtraction). A design goal of our proposed homomorphism is that string data and integer data are treated identically, in order to enhance security. Kyle Loyka, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 3 |
| 2018 | A Plain-Text Incremental Compression (PIC) Technique with Fast Lookup AbilityabstractData compression is a key aspect of computing applications such as online search engines, cloud computing and big data computing. In recent times, with the increasing popularity of remote and cloud-based computation, compression is becoming more important. Reducing the size of a data object in this context would not only reduce the transfer time, but also the amount of data transferred. The key figures of merit of a data compression scheme are its compression ratio and its compression, decompression and search (lookup) speeds. Traditional compression techniques achieve high compression ratios, but require decompression before a lookup can be performed, thus increasing the lookup time. In this paper, we propose an incremental compression technique for plain-text data objects (PIC), that uses variable length encoding to compress data. The dictionary of possible words is sorted based on the statistical frequency of the words, and the words are encoded using the variable length code-words. Words that are not in the dictionary are handled as well. The driving motivation of our technique is to perform significantly faster lookups without the need to decompress the compressed data object. PIC also facilitates string operations (such as concatenation, insertion, deletion and lookup-and-replacement) on compressed text without the need of decompression. In this manner, these string operations are incrementally compressed, resulting in greatly improved efficiency. We implement our technique in C++, and compare PIC with industry standard tools like lz4, gzip and bzip2 in terms of compression ratio, lookup speed, and lookup-and-replace time. PIC is about 8.84×, 100.17× and 174.25× faster as compared to lz4, gzip and bzip2, respectively, when the data is looked-up, and restored into a compressed format. A lookup-and-replace operation on a PIC-compressed file is shown to be 2.83× faster than a plain-text file. We formally prove that our method does not produce any false positive or false negative results during a lookup, and also prove that PIC is amenable to file concatenation. We have also implemented a parallel version of PIC. Compression, decompression and search times are 2.24×, 1.51× and 5× faster when eight threads are used instead of one. We also demonstrate that the parallel version of PIC scales better compared to commonly available tools for parallel compression. PIC opens the possibility of performing all string manipulations on a PIC-compressed file, yielding about 2× compression over plain text, and 1-2 orders of magnitude faster operations than traditional compression schemes. Kunal Bharathi, Abbas A. Fairouz, Ahmad Al Kawam, Sunil P. Khatri |
ICCD | 5 |
| 2018 | Synchronization of Ring-Based Resonant Standing Wave Oscillators for 3D Clocking ApplicationsabstractRing-based Resonant Standing Wave Oscillators (RRSWOs) have been shown to be a powerful technique to perform low power clock generation and distribution for high-speed clocks. The resulting clock signals can achieve very high frequencies and exhibit low skew. Utilizing through-silicon-vias (TSVs), a RRSWO can form a 3D ring that distributes this high speed clock to all the die in a 3D IC stack. In this paper, in addition to presenting a RRSWO-based technique to generate and distribute a high frequency, low skew signal in a 3D IC, we also present a scheme to synchronize two such 3D ring-based resonant oscillators without the use of a Phase Locked Loop (PLL). This reduces cost, provides effective synchronization between the 3D ICs, and creates scalable systems with a low clock skew. Finally, we present the design of a robust bootstrap circuit for 3D RRSWOs. We validate out ideas using circuit simulations, including Monte Carlo simulations to quantify the resilience of our approach to variations. Andrew J. Douglass, Sunil P. Khatri |
ICCD | 2 |
| 2017 | Design of a Flash-based Circuit for Multi-valued LogicabstractFlash transistors serve as the technology of choice for implementing nonvolatile memory. Current flash memory densities are increasing to meet storage demands. One of the key features that has resulted in increasing flash memory densities is the ability to program a flash transistor to have multiple threshold voltages. This feature has recently been exploited to implement ternary-valued logic. However, this implementation exhibited increased delays, due to the lowered Vgs values which result from using multiple threshold voltages. In this work, we present a circuit implementation that uses flash transistors to implement multi-valued digital circuits. Monther Abusultan, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 2 |
| 2017 | Circuit Level Design of a Hardware Hash Unit for use in Modern MicroprocessorsabstractModern microprocessors contain several Special Function Units (SFUs) such as specialized arithmetic units, cryptographic processors, etc. In recent times, applications such as cloud computing, web-based search engines, and network applications are widely used, and place new demands on the microprocessor. Hashing is a key algorithm that is extensively used in such applications. Hashing is typically performed in software. Thus, implementing a hardware-based hash unit on a modern microprocessor would potentially increase performance significantly. In this paper, we present the circuit design for a hardware hash unit (HU) for modern microprocessors, using a 45nm technology. Our proposed hardware hash unit is based on the use of a CAM to implement each bin of the hash function. We simulate the HU circuit and compare it with a traditional CAM design. We demonstrate an average power reduction of 5.48x using the HU over the traditional CAM. Also, we show that the HU can operate at a maximum frequency of 1.39 GHz (after accounting for process, voltage and temperature (PVT) variations and accounting for wiring parasitics). Furthermore, we present the delay, power and area trade-offs of the HU design with varying hash table sizes. Abbas A. Fairouz, Monther Abusultan, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 3 |
| 2017 | A Robust C-element Design with Enhanced Metastability PerformanceabstractMetastability causes unpredictable behavior in circuits which can sometimes cause circuit failure. Although much work has been done to reduce the possibility of metastability in synchronous designs, metastability in asynchronous designs has not been given significant attention to date. For asynchronous designs, metastability resolution is assumed to be handled by the handshaking protocol. However, metastability might manifest (at the electrical level) in various asynchronous circuit elements for special timing cases.. Kinshuk Sharma, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 2 |
| 2017 | Fast, Ring-Based Design of 3D Stacked DRAMabstractAs computer memory increases in size and processors continue to get faster, the memory subsystem becomes an increasing bottleneck to system performance. To mitigate the relatively slow DRAM memory chip speeds, a new generation of 3D stacked DRAM is being developed, with lower power consumption and higher bandwidth. This paper proposes the use of 3D ring-based data fabrics for fast data transfer between these chips. The ring-based data fabric uses a fast standing wave oscillator to clock its transactions. With a fast clocking scheme, and multiple channels sharing the same bus, more channels are utilized while significantly reducing the number of through-silicon vias (TSVs). Experimental results show that our ring-based data fabric can reduce read latencies by almost 4X compared to traditional stacked memory chips. Variations of our scheme can also reduce power consumption compared to traditional memory stacks. Our Memory Architecture using a Ring-based Scheme (MARS) can effectively trade off power, throughput, and latency to improve system performance for different application spaces. We show that our MARS variants can deliver better latency (up to ~4X), power (up to ~8X), and performance per watt (up to ~4X) over HBM, when averaged over 11 SPEC CPU 2006 benchmarks. Other MARS variants provide higher throughput with similar power consumption compared to Wide I/O memory. Andrew J. Douglass, Sunil P. Khatri |
ICCD | 2 |
| 2017 | An FPGA-Based Coprocessor for Hash Unit AccelerationabstractIn recent times, applications like web-based search, antivirus scanners, cloud computing, social media applications, and network applications are extremely common. The hash table is a heavily used data structure in such applications. Modern microprocessors have several special function units (SFUs) such as a floating point unit, a memory management unit, and a cryptography unit. However, hashing is typically performed in software, which reduces the performance of such applications. In this paper, we propose an FPGA-based implementation of a hash unit (a hash function and a hash table) in an FPGA. The FPGA-based hash unit is implemented as a coprocessor for a CPU. The CPU and the FPGA communicate through a PCI Express (PCIe) interface. The hash table in our hash unit is implemented as a content-addressable memory (CAM), to enhance the speed of hash operations. The hash unit (HU) coprocessor is tested in the context of virus checking application, when the hashing operation only requires membership checks. Our HU can be used in other hashing applications as well; we use virus checking as a representative application. Hashing operations are performed in a batch on the FPGA, to provide better utilization of the PCIe bus. We demonstrate a significant performance of up to 7.3× for our FPGAbased hash unit implementation compared to a software-based hashing implementation. This speedup is for the entire virus checking application (not just the hash lookup portion of the virus checking application). Abbas A. Fairouz, Sunil P. Khatri |
ICCD | 2 |
| 2017 | A Survey of Software and Hardware Approaches to Performing Read Alignment in Next Generation SequencingabstractComputational genomics is an emerging field that is enabling us to reveal the origins of life and the genetic basis of diseases such as cancer. Next Generation Sequencing (NGS) technologies have unleashed a wealth of genomic information by producing immense amounts of raw data. Before any functional analysis can be applied to this data, read alignment is applied to find the genomic coordinates of the produced sequences. Alignment algorithms have evolved rapidly with the advancement in sequencing technology, striving to achieve biological accuracy at the expense of increasing space and time complexities. Hardware approaches have been proposed to accelerate the computational bottlenecks created by the alignment process. Although several hardware approaches have achieved remarkable speedups, most have overlooked important biological features, which have hampered their widespread adoption by the genomics community. In this paper, we provide a brief biological introduction to genomics and NGS. We discuss the most popular next generation read alignment tools and algorithms. Furthermore, we provide a comprehensive survey of the hardware implementations used to accelerate these algorithms. Ahmad Al Kawam, Sunil P. Khatri, Aniruddha Datta |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2016 | A GPU-based implementation of a sensor tasking methodology
Monther Abusultan, Suman Chakravorty, Sunil P. Khatri |
FUSION | 3 |
| 2016 | A flash-based digital circuit design flowabstractTraditionally, floating gate (flash) transistors have been used exclusively to implement non-volatile memory in its various forms. Recently, we showed that flash transistors can be used to implement digital circuits as well. In this paper, we present the details on the realization and characteristics of the block-level flash-based digital design. The current work describes the synthesis flow to decompose a circuit block into a network of interconnected FCs. The resulting network is characterized with respect to timing, power and energy, and the results are compared with a standard-cell based realization of the same block (obtained using commercial tools). We obtain significantly improved delay (0.59×), power (0.35×) and cell area (0.60×) compared to a traditional CMOS standard-cell based approach, when averaged over 12 standard benchmarks. It is generally rare that a circuit methodology yields results that are better than existing commercial standard-cell based flows in terms of delay, area, power and energy, and in this sense, we submit that our results are significant. Additional benefits of a flash-based digital design is that it allows for precision speed binning in the factory, and also enables in-field re-programmability (we note that our flash-based design is not an FPGA, but rather an ASIC style design) to counteract the speed degradation of a design due to aging. These benefits arise from the fact that the threshold voltage of flash devices can be controlled with precision. Monther Abusultan, Sunil P. Khatri |
ICCAD | 2 |
| 2016 | Implementing low power digital circuits using flash devicesabstractFloating gate (flash) transistors are used exclusively for memory applications today. These applications include SD cards of various form factors, USB flash drives and SSDs. This paper presents the first approach to use flash transistors to implement binary-valued digital circuits. Since the threshold voltage of flash devices can be modified at a fine granularity during programming, several advantages are obtained by our approach. For one, speed binning at the factory can be controlled with precision. Secondly, an IC can be re-programmed in the field, to negate effects such as aging, which has been a significant problem in recent times, particularly for mission-critical applications. We present the circuit topology that we use in our flash-based digital circuit approach, and, through circuit simulations, show that our approach yields significantly improved delay (0.84×), power (0.35×), energy (0.30×) and area (0.54×) characteristics compared to a traditional CMOS standard cell based approach, when averaged over 20 randomly generated designs. Note that we used the same operating voltage (1V) for both design styles. Our proposed circuit design style is not an FPGA, because it uses hardwired interconnect. Rather, our design approach is a method to design ASIC or custom/semi-custom digital circuits. Monther Abusultan, Sunil P. Khatri |
ICCD | 2 |
| 2016 | Exploring static and dynamic flash-based FPGA design topologiesabstractField programmable gate arrays (FPGAs) are the implementation platform of choice when it comes to design flexibility. However, SRAM-based FPGAs suffer from high power consumption, prolonged boot delays (due to the volatility of the configuration bits), and a significant area overhead (due to the use of 5T SRAM cells for the configuration bits). Floating gate (flash) based FPGAs can avert these problems. This paper presents a study of flash-based FPGA designs (both static and dynamic), and presents the tradeoff of delay, power dissipation and energy consumption of the various designs. Our work differs from previously proposed flash-based FPGAs, since we embed the flash transistors (which store the configuration bits) directly within the logic and interconnect fabrics. We also present a detailed description of how the programming of the configuration bits is accomplished. Our delay and power estimates are derived from circuit level simulations. Our proposed static flash-based LUT structure yields 10% faster operation, 12% lower dynamic power dissipation, 21% lower energy consumption and 29% lower static power dissipation compared to a traditional SRAM-based LUT. We also show that, for high performance applications, a dynamic flash-based LUT can achieve further performance improvements (32% lower delay) with higher energy consumption (37% higher) compared to an SRAM-based LUT. We also show that a flash-based interconnect structure provides 89% lower delay and 71% lower overall power consumption compared to the traditional interconnect structure used in SRAM-based FPGAs. Monther Abusultan, Sunil P. Khatri |
ICCD | 2 |
| 2016 | A novel hardware hash unit design for modern microprocessorsabstractHistorically, microprocessor instructions were designed in order to obtain high performance on integer and floating point computations. Today's applications, however, demand high performance for cloud computing, web-based search engines, network applications, and social media tasks. Such software applications involve an extensive use of hashing in their computation. Hashing can reduce the complexity of search and lookup from O(n) to O(n/k), where k bins are used. In modern microprocessors hashing is done in software. In this paper, we propose a novel hardware hash unit design for use in modern microprocessors. We present the design of the Hash Unit (HU) at the micro-architecture level. We simulate the new HU to compare its performance with a software-based hash implementation. We demonstrate a significant speed-up (up to 12×) for the HU. Furthermore, the performance scales elegantly with increasing database size and application diversity, without increasing the hardware cost. Abbas A. Fairouz, Monther Abusultan, Sunil P. Khatri |
ICCD | 3 |
| 2016 | FTCAM: An Area-Efficient Flash-Based Ternary CAM DesignabstractThis paper presents a Ternary Content-addressable Memory (TCAM) design which is based on the use of floating-gate (flash) transistors. TCAMs are extensively used in high speed IP networking, and are commonly found in routers in the internet core. Traditional TCAM ICs are built using CMOS devices, and a single TCAM cell utilizes 17 transistors. In contrast, our TCAM cell utilizes only two flash transistors, thereby significantly reducing circuit area. We cover the chip-level architecture of the TCAM IC briefly, focusing mainly on the TCAM block which does fast parallel IP routing table lookup. Our flash-based TCAM (FTCAM) block is simulated in SPICE, and we show that it has a significantly lowered area compared to a CMOS based TCAM block, with a speed that can meet current ($\sim$400 Gb/s) data rates that are found in the internet core. Viacheslav V. Fedorov, Monther Abusultan, Sunil P. Khatri |
IEEE Trans. Computers | 3 |
| 2015 | Delay, Power and Energy Tradeoffs in Deep Voltage-scaled FPGAsabstractIn this paper, we present a circuit-level analysis of deep voltage-scaled FPGAs, which operate from full supply to sub-threshold voltages. The logic as well as the interconnect of the FPGA are modeled at the circuit level, and their relative contribution to the delay, power and energy of the FPGA are studied by means of circuit simulations. Three representative designs are studied to explore these design trade-offs. We conclude that the energy and delay-minimal FPGA design is one in which both the interconnect and logic are curtailed from scaling below a fixed voltage (about 550mV in our experiments). If power is a more important design factor (at the cost of delay), it is beneficial to operate both the logic and interconnect between 300mV and 800mV. Monther Abusultan, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 2 |
| 2015 | An Efficient Approach to Sample On-Chip Power SuppliesabstractIn recent years, post-silicon debugging has become a significantly difficult exercise due to the increase in the size of the electrical state of the IC being debugged, coupled with the limited fraction of this state that is visible to the debug engineer. As the number of transistors increases, the number of possible electrical states increases exponentially, while the amount of information that can be accessed grows at a much slower rate. This difficulty is compounded by the outsourcing of IP blocks, which creates more black boxes that the debug engineer must work around. As a result, when an IC fails tracking down the cause of the failure becomes a monumental task, and debugging becomes more art than science. One source of errors in a test circuit is the fluctuation of the power supplies during a single clock cycle. These supply variations can increase or decrease the speed of a circuit and lead to errors such as hold time violations and setup time violations. This paper presents a circuit that samples precisely the power supply multiple times in a clock cycle, allowing the debug engineer to quantify the variations in the supply over a clock cycle. With this information, a better understanding of the electrical state of the test chip is made possible. The circuit presented in this paper can sample the supply voltage with a quantization of 0.291mV, and the output is linear with an R2 value of 0.9987. Luke Murray, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 2 |
| 2015 | Exploring the viability of stochastic computingabstractRecently, stochastic circuits have received significant attention from academia. Stochastic circuits claim to have a reduced energy consumption at the cost of accuracy and delay. In this paper, we explore the power, delay, energy and area of a stochastic circuit (a stochastic multiplier in particular), and compare these metrics with those of a regular multiplier, implemented using the Sum Of Products (SOP) approach. The SOP based multiplier is implemented both using a Kogge-Stone Adder, as well as a Ripple-Carry adder. Our results show that when the stochastic number generator (SNG) and counter are included in the stochastic multiplier (SM), even for 3 bits, the SM consumes more energy to finish one multiplication than an SOP based regular binary multiplier (RM), and this energy consumption grows exponentially as the number of bits increases. If we only consider the stochastic multiplier cell (SMC, which is simply a 2-input AND gate) and ignore the energy of the SNG and counter, the SMC has a better energy consumption for multiplications up to 12 bits. However, even for 3 bits, the SM (or the SMC) is slower by over 5x compared to the regular multiplier, and this delay increases exponentially as the number of bits increases. The area of the SM (including the area of the SNG and counter) is smaller for multipliers with more than 6 bits. Joao Marcos de Aguiar, Sunil P. Khatri |
ICCD | 2 |
| 2014 | Look-up Table Design for Deep Sub-threshold through Full-Supply OperationabstractField programmable gate arrays (FPGAs) are the implementation platform of choice when it comes to design flexibility. However, the high power consumption of FPGAs (which arises due to their flexible structure), make them less appealing for extreme low power applications. In this paper, we present a design of an FPGA lookup table (LUT), with the goal of seamless operation over a wide band of supply voltages. The same LUT design has the ability to operate at sub-threshold voltage when low power is required, and at higher voltages whenever faster performance is required. The results show that operating the LUT in sub-threshold mode yields a (~80×) lower power and a (~4×) lower energy than full supply voltage operation, for a 6-input LUT implemented in a 22nm predictive technology. The key drawback of sub-threshold operation is its susceptibility to process, temperature, and supply voltage (PVT) variations. This paper also presents the design and experimental results for a closed-loop adaptive body biasing mechanism to dynamically cancel global (spacial) as well as local (random) PVT variations. For the same 22nm technology, we demonstrate that the closed-loop adaptive body biasing circuits can allow the FPGA LUT to operate over an operating frequency range that spans an order of magnitude (40 MHz to 1300 MHz). We also show that the closed-loop adaptive body biasing circuits can cancel delay variations due to supply voltage changes, and reduce the effect of process variations on setup and hold times by 1.8× and 2.9× respectively. The dynamic body biasing circuits incur a 3.49% area overhead when designed to each drive a cluster of 25 LUTs. Monther Abusultan, Sunil P. Khatri |
FCCM | 2 |
| 2014 | FPGA LUT design for wide-band dynamic voltage and frequency scaled operation (abstract only)abstractField programmable gate arrays (FPGAs) are the implementation platform of choice when it comes to design flexibility. However, the high power consumption of FPGAs (which arises due to their flexible structure), make them less appealing for extreme low power applications. In this paper, we present a design of an FPGA look-up table (LUT), with the goal of seamless operation over a wide band of supply voltages. The same LUT design has the ability to operate at sub-threshold voltage when low power is required, and at higher voltages whenever faster performance is required. The results show that operating the LUT in sub-threshold mode yields a (~80x) lower power and (~4x) lower energy than full supply voltage operation, for a 6-input LUT implemented in a 22nm predictive technology. The key drawback of sub-threshold operation is its susceptibility to process, temperature, and supply voltage (PVT) variations. This paper also presents the design and experimental results for a closed-loop adaptive body biasing mechanism to dynamically cancel these PVT variations. For the same 22nm technology, we demonstrate that the closed-loop adaptive body biasing circuits can allow the FPGA to operate over an operating frequency range that spans an order of magnitude (40 MHz to 1300 MHz). We also show that the closed-loop adaptive body biasing circuits can cancel delay variations due to supply voltage changes, and reduce the effect of process variations on setup and hold times by 1.8x and 2.9x respectively. Monther Abusultan, Sunil P. Khatri |
FPGA | 2 |
| 2014 | A comparison of FinFET based FPGA LUT designsabstractThe FinFET device has gained much traction in recent VLSI designs. In the FinFET device, the conduction channel is vertical, unlike a traditional bulk MOSFET, in which the conduction channel is planar. This yields several benefits, and as a consequence, it is expected that most VLSI designs will utilize FinFETs from the 20nm node and beyond. Despite the fact that several research papers have reported FinFET based circuit and layout realizations for popular circuit blocks, there has been no reported work on the use of FinFETs for Field Programmable Gate Array (FPGA) designs. The key circuit in the FPGA that enables programmability is the n-input Look-up Table (LUT). An n-input LUT can implement any logic function of up to n inputs. In this paper, we present an evaluation of several FPGA LUT designs. We compare these designs from a performance (delay, power, energy) as well as an area perspective. Comparisons are conducted with respect to a bulk based LUT as well. Our results demonstrate that all the FinFET based LUTs exhibit better delays and energy than the bulk based LUT. Based on our comparisons, we have two winning candidate LUTs, one for high performance designs (3X faster than a bulk based LUT) and another for low energy, area constrained designs (83% energy and 58% area compared to a bulk based LUT). Monther Abusultan, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 2 |
| 2014 | An asynchronous Network-on-Chip router with low standby powerabstractThe Network-on-Chip (NoC) paradigm is now widely used to interconnect the processing elements (PEs) in a chip multi-processor (CMP). It has been reported that the NoC consumes about a third of the total power consumption of the multi-core processor. To address this, asynchronous NoC routers have been proposed, to eliminate the clocking power associated with the NoC implementation, which is typically a large fraction of the NoC power consumption. In this work, we present a technique to reduce the standby power of a state-of-the-art asynchronous NoC router. In our approach, the router is put in a known input state when idle, and each gate in the unmodified router is replaced by a logically equivalent gate whose supply pin is connected to a PMOS device with a high threshold voltage in case its output in the idle state was 0. On the other hand, if the output of the unmodified gate in the idle state was 1, it is replaced by a logically equivalent gate whose ground terminal is connected to a NMOS device with a high threshold voltage. Our router is inserted in a NoC, and verified logically for correct routing functionality. We also simulated it at the circuit level using a 45nm fabrication technology, and show that it has a low wake-up time from sleep, and a minimal steady-state routing delay (13%) and area (23%) overhead, and a 8.1× lower standby power, when compared to an unmodified asynchronous NoC router, which was also implemented. Our leakage improvement is achieved in part by using a novel method to control the leakage of the inverter chain used to drive the sleep signal, something which that is not possible with traditional leakage reduction techniques. Amr Elshennawy, Sunil P. Khatri |
ICCD | 2 |
| 2014 | An area-efficient Ternary CAM design using floating gate transistorsabstractThis paper presents a Ternary Content-addressable Memory (TCAM) design which is based on the use of floating-gate (flash) transistors. TCAMs are extensively used in high speed IP networking, and are commonly found in routers in the internet core. Traditional TCAM ICs are built using CMOS devices, and a single TCAM cell utilizes 17 transistors. In contrast, our TCAM cell utilizes only 2 flash transistors, thereby significantly reducing circuit area. We cover the chip-level architecture of the TCAM IC briefly, focusing mainly on the TCAM block which does fast parallel IP routing table lookup. Our flash based TCAM block is simulated in SPICE, and we show that it has a significantly lowered area compared to a CMOS based TCAM block, with a speed that can meet current (~400 Gb/s) data rates that are found in the internet core. Viacheslav V. Fedorov, Monther Abusultan, Sunil P. Khatri |
ICCD | 3 |
| 2013 | Crosstalk avoidance codes for 3D VLSIabstractIn 3D VLSI, through-silicon vias (TSVs) are relatively large, and closely spaced. This results in a situation in which noise on one or more TSVs may deteriorate the delay and signal integrity of neighboring TSVs. In this paper, we first quantify the parasitics in contemporary TSVs, and then come up with a classification of crosstalk sequences as 0C, 1C, … 8C sequences. Next, we present inductive approaches to quantify the exact overhead for 8C, 6C and 4C crosstalk avoidance codes (CACs) for a 3×n mesh arrangement of TSVs. These overheads for different CACs for a 3×n mesh arrangement of TSVs are used to calculate the lower bounds on the corresponding overheads for an n×n mesh arrangements of TSVs. We also discuss an efficient way to implement the coding and decoding (CODEC) circuitry for limiting the maximum crosstalk to 6C. Our experimental results show that for a TSV mesh arrangement driven by inverters implemented in a 22nm technology, the coding based approaches yields improvements which are in line with the theoretical predictions. Sunil P. Khatri |
DATE | 2 |
| 2013 | Exploring topologies for source-synchronous ring-based network-on-chipabstractThe mesh interconnection network has been preferred by the Network-on-Chip (NoC) community due to its simple implementation, high bandwidth and overall scalability. Most existing mesh-based NoC designs operate the mesh at the same or lower clock speed as the processing elements (PEs). Recently, a new source synchronous ring-based NoC architecture has been proposed, which runs significantly faster than the PEs and offers a significantly higher bandwidth and lower communication latency. The authors implement the NoC topology as a mesh of rings, which occupies the same area as that of a mesh. In this work, we evaluate two alternate source synchronous ring-based NoC topologies called the ring of stars (ROS) and the spine with rings (SWR), which occupy a much lower area, and are able to provide better performance in terms of communication latency compared to a state of the art mesh. In our proposed topologies, the clock and the data NoC are routed in parallel, yielding a fast, synchronous, robust design. Our design allows the PEs to extract a low jitter clock from the high speed ring clock by division. The area and performance of these ring-based NoC topologies is quantified. Experimental results on synthetic traffic show that the new ring-based NoC designs can provide significantly lower latency (upto 4.6×) compared to a state of the art mesh. The proposed floorplan-friendly topologies use fewer buffers (upto 50% less) and lower wire length (upto 64.3% lower) compared to the mesh. Depending on the performance and the area desired, a NoC designer can select among the topologies presented. Ayan Mandal, Sunil P. Khatri, Rabi N. Mahapatra |
DATE | 2 |
| 2013 | GPU implementation of a scalable non-linear congruential generator for cryptography applicationsabstractFast key generation algorithms which can generate random sequences of varying atomic lengths and throughput are important for secure data communication. In this paper, we present a non-linear congruential method for generating high quality random numbers at flexible throughput rates of upto 66 Gbps, on a GPU platform. Each random number can have up to 4096 key bits. The method can be easily extended for implementation on hardware platforms like FPGAs and ASICs as well. Our key generator is comprised of N Linear Congruential Generators (LCGs) running in parallel; we have chosen N=4096 for the GPU implementation. The outputs of the LCGs are combined using N encoded majority functions. The encoded majority function used for any bit is changed in every generation iteration. In our GPU implementation of the Non-Linear Congruential Generator (GPU-NLCG), it is possible to alter the LCG functions on the fly by changing the primes periodically without interrupting the generation. Our GPU-NLCG can be used for high speed cryptographic key generation for rates up to 66 Gbps and can be easily integrated into multi-threaded applications in cryptography and Monte Carlo methods. The GPU-NLCG passes the NIST, Diehard and Dieharder battery of tests of randomness, which ensure the quality of our ciphers. Aditya Belsare, Steve Liu, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 3 |
| 2013 | Architecture and 3D device simulation of a PIN diode-based Gamma radiation detectorabstractIn this paper, we present a new IC-based gamma radiation detector. We report 3D simulation results for the PIN diode structure which is used in this detector, along with a discussion of the architecture of the readout electronics for this detector. Gamma detection traditionally consists of three parts -- i) a scintillator, ii) a detector unit and iii) the readout electronics. Our PIN diode based detector uses a traditional scintillator and a bialkali photocathode to convert the scintillated light into an electron current. This current is detected by the PIN diodes in the detector, which is fused to the photocathode. This structure has many advantages. Sensitivity is enhanced by placing the photocathode and detector IC on all six faces of the NaI scintillating crystal. Our detector is small, has a low-dead-time and it consumes very low power. The sensitive area of the chip is 75%, which is higher than any existing solid-state gamma detector. We have verified the operation of each component of the system by performing circuit-level simulations. The detector is implemented in a low-cost 0.5 um process. The optimal parameters for the PIN diode detector are explored through 3D device simulations. Our results show that a detector sensitivity better than 20 keV is easily achievable with a PIN diode of dimensions 2.5 um X 2.5 um. Amr Elshennawy, Craig M. Marianno, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 3 |
| 2013 | A source-synchronous Htree-based network-on-chipabstractMost existing Network-on-Chip (NoC) designs operate at the same or lower clock speed as the processing elements (PEs). Recently, a new source-synchronous ring-based NoC architecture has been proposed, which runs significantly faster than the PEs and offers a significantly higher bandwidth and lower communication latency. However, the ring-based design assumes a separate clock distribution scheme for the NoC and the PEs, and uses a standard mesh topology for the NoC. In this work, we present a source synchronous ring-based NoC, laid out in an H-tree topology, with each data link being routed parallel to a clock ring. The clock is generated and distributed by multiple standing wave oscillator (SWO) rings, which are also laid out in an H-tree topology. Our design allows the PEs to extract a low jitter clock directly from the high speed ring-based SWO clock by division. Moreover, since the PEs are synchronous with the ring clock, they do not need synchronizers while communicating with the NoC. We also show that by recursively duplicating links in the H-tree based source synchronous NoC (Hnoc), we can obtain new hybrid NoC structures. In the limit, this recursive duplication causes the H-tree based NoC to morph into the meshbased source synchronous NoC (Mnoc). The performance of each such intermediate hybrid NoC structure is quantified in terms of area, link utilization and contention free latency. We also enhance the performance of the hybrid NoCs by widening congested links, and quantify the tradeoffs. Experimental results show that the hybrid NoC designs can provide significantly lower latency (upto 5× lower) and are able to sustain a higher injection rate (upto 6.8× higher) compared to a state of the art mesh. Moreover, these hybrid NoC designs use fewer buffers (upto 19.4% less) and lower wire length (upto 19.7% lower) compared to a mesh. Based on the performance and the area tradeoffs, an NoC designer can select any hybrid NoC structure among the presented. Ayan Mandal, Sunil P. Khatri, Rabi N. Mahapatra |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | Noise-based algorithms for functional equivalence and tautology checkingabstractIn this paper, we present noise-based algorithms for functional equivalence and tautology checking using noise-based logic (NBL). A key property of NBL is that literals are represented by independent noise sources, from which we can construct noise-based cubes, and superpositions of such noise-based cubes, to create a noise-based Boolean function on a single wire. In our algorithms, the Boolean sum-of-products (SOP) formula is expressed in NBL as a superposition of its minterms. This noise-based representation of the SOP can then be compared with that of another SOP formula for equivalence checking (or with the noise-based formula representing tautology, for tautology checking) using a single operation. We validate our approach using software simulation. Pey-Chang Kent Lin, Sunil P. Khatri |
ICCD | 2 |
| 2013 | A low-jitter phase-locked resonant clock generation and distribution schemeabstractClock distribution networks have traditionally been optimized to minimize end-to-end delay of the distribution network. However, since most digital ICs have an on-chip PLL, a more relevant design goal is to minimize cycle-to-cycle jitter. In this paper, we present a novel low-jitter phase-locked clock generation and distribution methodology which uses resonant standing wave oscillators (SWOs). In contrast to traveling wave oscillator rings (TWOs or “rotary” clocks), our SWO achieves the same phase at every point in the ring, making it amenable to a synchronous design methodology. The standing wave oscillator is controlled by coarse as well as fine tuning. Coarse tuning is achieved by varying the ring inductance, while fine tuning is accomplished by varying the ring capacitance. Clock distribution is done by routing the resonant ring chip-wide in a “comb” like manner. Experimental results demonstrate that the cycle-to-cycle jitter and skew of our approach is dramatically lower than existing schemes, while the power consumption is significantly lower as well. These benefits occur due to the resonant nature of our SWO-based clock generation and distribution approach. Ayan Mandal, Kalyana C. Bollapalli, Nikhil Jayakumar, Sunil P. Khatri, Rabi N. Mahapatra |
ICCD | 4 |
| 2012 | Application of logic synthesis to the understanding and cure of genetic diseasesabstractIn the quest to understand and cure genetic diseases such as cancer, the fundamental approach being taken is undergoing a gradual change. It is becoming more acceptable to view these diseases as an engineering problem, and systems engineering approaches are becoming more accepted as a means to tackle genetic diseases. In this light, we believe that logic synthesis techniques can play a very important role. Several techniques from the field of logic synthesis can be adapted to assist in the arguably huge effort of modeling and controlling such diseases. The set of genes that control a particular genetic disease can be modeled as a Finite State Machine (FSM) called the Gene Regulatory Network (GRN). Important problems include (i) inferring the GRN from observed gene expression data from patients and (ii) assuming that such a GRN exists, determining the "best" set of drugs so that the disease is "maximally" cured. In this paper, we report initial results on the application of logic synthesis techniques that we have developed to address both these problems. In the first technique, we present Boolean Satisfiability (SAT) based approaches to infer the logical support of each gene that regulates melanoma, using gene expression data from patients of the disease. From the output of such a tool, biologists can construct targeted experiments to understand the logic functions that regulate a particular gene. The second technique assumes that the GRN is known, and uses a weighted partial Max-SAT formulation to find the set of drugs with the least side-effects, that steer the GRN state towards one that is closest to that of a healthy individual, in the context of colon cancer. Our group is currently exploring the application of several other logic techniques to a variety of related problems in this domain. Pey-Chang Kent Lin, Sunil P. Khatri |
DAC | 2 |
| 2012 | Boolean satisfiability using noise based logicabstractNoise-based Logic (NBL) is a probabilistic logic system which can be used to simultaneously apply a superposition of arbitrarily many input vectors to a SAT instance. Using this property, we can determine whether an instance is SAT in a single operation. A satisfying solution can be found by iteratively performing SAT checks up to n times, where n is the number of variables in the SAT instance. In this paper, we formulate NBL-based SAT, and discuss its scalability. The NBL-based SAT engine has been simulated in software for validation purposes, although the focus of the paper is on the theory of NBL-based SAT. Pey-Chang Kent Lin, Ayan Mandal, Sunil P. Khatri |
DAC | 3 |
| 2012 | A fast, source-synchronous ring-based network-on-chip designabstractMost network-on-chip (NoC) architectures are based on a mesh-based interconnection structure. In this paper, we present a new NoC architecture, which relies on source synchronous data transfer over a ring. The source synchronous ring data is clocked by a resonant clock, which operates significantly faster than individual processors that are served by the ring. This allows us to significantly improve the cross section bandwidth and the latency of the NoC. We have validated the design using a 22 nm predictive process. Compared to the state-of-the-art mesh based NoC, our scheme achieves a 4.5× better bandwidth, 7.4× better contention free latency with 11% lower area and 35% lower power. Ayan Mandal, Sunil P. Khatri, Rabi N. Mahapatra |
DATE | 2 |
| 2012 | Alleviating NBTI-induced failure in off-chip output driversabstractNegative Bias Temperature Instability (NBTI) causes the threshold voltage of PMOS devices to degrade with time, resulting in a reduced lifetime of a CMOS IC. In this paper, we present an approach to mitigate the degradation due to NBTI for off-chip output drivers. Our approach is based on forcibly inducing relaxation in the individual fingers of the output driver (which is typically implemented in a multi-fingered fashion). The individual fingers are relaxed in a round-robin manner, such that at any given time, k out of n fingers of the driver are being relaxed. Our results show that the proposed approach significantly extends the lifetime of the output driver. Bhavitavya Bhadviya, Ayan Mandal, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 3 |
| 2012 | An efficient arithmetic Sum-of-Product (SOP) based multiplication approach for FIR filters and DFTabstractIn this paper, we present an arithmetic Sum-of-Product (SOP) based approach to implement an efficient Discrete Fourier Transform (DFT) as well as an FIR filter circuit. Our SOP based DFT engine uses an improved column compression algorithm, and also handles the sign of the input efficiently. The partial products of the computation are compressed down to 2 operands, which are then added using a single hybrid adder (which is comprised of a ripple carry adder for the early-arriving lower-order bits, a Kogge-Stone adder for the slower middle bits, and a carry-select adder for the early-arriving higher order bits). The DFT can also be cast as an instance of the Multiple Constant Multiplication (MCM) problem. We compare our SOP-based DFT implementation with the RAG-n approach, the best in-class existing implementation for the MCM problem. RAG-n utilizes a cascade of adders, and attempts to heuristically minimize the number of adders by sharing them across different computations of the DFT. We implemented both approaches using a 45 nm cell library, and demonstrate that our approach yields a faster DFT engine (by about 12-13%), with a small (about 5%) area penalty and a significantly better algorithmic runtime. Our approach is able to complete for DFT problems with a much higher bit precision than the RAG-n approach. The approach of our paper is generalized to implement digital filters as well, and we demonstrate that our approach realizes FIR filters with hard-to-implement coefficients with a 4× speedup and 1.4× area penalty compared to two recent adder-cascade based approaches [1]. Ayan Mandal, Sunil P. Khatri |
ICCD | 3 |
| 2012 | Architectural simulations of a fast, source-synchronous ring-based Network-on-Chip designabstractRecently, a new source-synchronous ring-based NoC architecture has been proposed, which runs significantly faster than the PEs and offers a high bandwidth and low contention free latency. Architectural simulations show that the original ring-based NoC design suffers from deadlock. In this paper, we explore the architectural aspects of the fast ring-based NoC after redesigning the routers used in the previous authors' work to avoid deadlock. Architectural results obtained on synthetic traffic demonstrate that the modified ring-based NoC has up to 3.5× lower latency and up to 2.9× higher maximum sustained injection rate compared with a state of the art mesh-based NoC. Ayan Mandal, Sunil P. Khatri, Rabi N. Mahapatra |
ICCD | 2 |
| 2012 | Timing aware partitioning for multi-FPGA based logic simulation using top-down selective hierarchy flatteningabstractIn order to accelerate logic simulation, it is highly beneficial to simulate the circuit design on FPGA hardware. This is often referred to as emulation, and we use the terms simulation and emulation interchangeably in this paper. However, limited hardware on FPGAs prevents large designs from being implemented on a single FPGA. Hence there is a need to partition the design and simulate it on a multi-FPGA platform. In contrast to existing FPGA-based post-synthesis partitioning approaches which first completely flatten the circuit and then possibly perform bottom-up clustering, we perform a selective top-down flattening and thereby avoid the potential netlist blowup. This also allows us to preserve the design hierarchy to guide the partitioning and to make subsequent debugging easier. Our approach analyzes the hierarchical design and selectively flattens instances using two metrics based on slack. The resulting partially flattened netlist is converted to a hypergraph, partitioned using hMetis, and reconverted back to a plurality of FPGA netlists, one for each FPGA of the FPGA-based accelerated logic simulation platform. We compare our approach with a partitioning approach that operates on a completely flattened netlist. Static timing analysis was performed for both approaches, and over 15 large examples from the OpenCores project, our approach yields a 52% logic simulation speedup and about 0.74× runtime for the entire flow, compared to the completely flat approach. The entire tool chain of our approach is automated in an end-to-end flow from hierarchy extraction, selective flattening, partitioning, and netlist reconstruction. Compared to an existing method which also performs slack-based partitioning of a hierarchical netlist, we obtain a 35% simulation speedup. Our method scales very well, yielding a significantly better simulation speedup and runtime improvement for larger examples. Subramanian Poothamkurissi Swaminathan, Pey-Chang Kent Lin, Sunil P. Khatri |
ICCD | 3 |
| 2012 | On Optimal and Achievable Fix-Free CodesabstractFix-free codes are prefix condition codes which can also be decoded in the reverse direction. They have attracted attention from several communities and are used in video standards. Two variations (with additional constraints) have also been considered for joint source-channel coding: 1) “symmetric” codes, which require the codewords to be palindromes; 2) codes with distance constraints on pairs of codewords. Approaches to determine the existence of a fix-free code with a given set of codeword lengths, for each of the three variations, are proposed. These appear to involve the first use of Boolean satisfiability (SAT) for the existence and design of source codes. Branch-and-bound algorithms to find the collection of optimal codes for asymmetric and symmetric fix-free codes are described. The first bound for the performance of optimal symmetric binary fix-free codes is provided. An earlier conjecture on optimal symmetric binary fix-free codes is proven and related results for asymmetric fix-free codes are presented. A variation of the 3/4 conjecture for fix-free codes is introduced. A key idea in the first conjecture and the new one is a definition of how one sequence of nondecreasing natural numbers dominates another. Serap A. Savari, S. M. Hossein Tabatabaei Yazdi, Navid Abedini, Sunil P. Khatri |
IEEE Trans. Inf. Theory | 4 |
| 2011 | A novel cryptographic key exchange scheme using resistorsabstractRecently, a secure key exchange technique was developed, in which both communicators (Alice and Bob) randomly select between two known resistors. By measuring the resulting thermal noise on a shared wire, they can each determine the resistor chosen by their counterpart, while the eavesdropper (Eve) cannot determine this. By repeating this transaction, they can create a common secure key, one bit a time. Although theoretically elegant, this approach is difficult to realize in practice. In this paper, we present a practical realization of a secure key exchange technique, intended for use over the Ethernet. Our approach is inspired by the above scheme with significant differences. In our approach, Alice and Bob utilize programmable resistors and exchange their resistance values securely. Our technique has been implemented in a hardware FPGA based platform, and was found to be able to exchange 4 secure bits per transaction over a 100ft CAT5 cable. Pey-Chang Kent Lin, Alex Ivanov, Bradley Johnson, Sunil P. Khatri |
ICCD | 4 |
| 2010 | Implementing digital logic with sinusoidal suppliesabstractIn this paper, a new type of combinational logic circuit realization is presented. Logic values are implemented as sinusoidal signals. Sinusoidal signals of the same frequency are phase shifted by ¿ to destructively interfere with each other, and represent the logic 0 and 1 values of Boolean Logic. These properties of sinusoids can be used to identify a signal without ambiguity. Thus, representing logic values as sinusoidal signals yields a realizable system of logic. The paper presents a logic gate family that can operate using the sinusoidal signals for logic 0 and logic 1 values. Due to orthogonality of sinusoid signals with different frequencies, multiple sinusoids could be transmitted on a single wire. This provides a natural way of implementing multilevel logic. Signals traveling long distances could take advantage of this fact and can share interconnect lines. Recent research in circuit design has made it possible to harvest sinusoidal signals of the same frequency and 180° phase difference from a single resonant clock ring in a distributed manner. Other advantage of such a logic family is its immunity from external additive noise. The experiments in this paper indicate that this paradigm, when used to implement binary valued logic, yields an improvement in switching (dynamic) power. Kalyana C. Bollapalli, Sunil P. Khatri, Laszlo B. Kish |
DATE | 2 |
| 2010 | A SAT-Based Scheme to Determine Optimal Fix-Free CodesabstractFix-free or reversible-variable-length codes are prefix condition codes which can also be decoded in the reverse direction. They have attracted attention from several communities and are used in video standards. Two variations of fix-free codes (with additional constraints) have also been considered for joint source-channel coding: 1) "symmetric" fix-free codes, which require the codewords to be palindromes; 2) fix-free codes with distance constraints on pairs of codewords. We propose a new approach to determine the existence of a fix-free code with a given set of codeword lengths, for each of the three variations of the problem. We also describe a branch-and-bound algorithm to find the collection of optimal codes for asymmetric and symmetric fix-free codes. Navid Abedini, Sunil P. Khatri, Serap A. Savari |
DCC | 2 |
| 2010 | Boolean satisfiability on a graphics processorabstractBoolean Satisfiability (SAT) is a core NP-complete problem. Several heuristic software and hardware approaches have been proposed to solve this problem. In this paper we present a Boolean satisfiablity approach with a new GPU-enhanced variable ordering heuristic. Our results demonstrate that over several satisfiable and unsatisfiable benchmarks, our technique (MESP) performs better than MiniSAT. We show a 2.35× speedup on average, over 68 from the SAT Race (2008) competition. Kanupriya Gulati, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | VLSI implementation of a non-linear feedback shift register for high-speed cryptography applicationsabstractFor secure high data-rate communications, fast key generation algorithms are crucial. In this paper, we present a VLSI implementation of a Non-Linear Feedback Shift Register (NLFSR) for cryptography applications. Unlike existing cryptographic key generation techniques, our NLFSR generates multiple (64 in our implementation) key bits in each clock cycle. This enables its use in secure, high speed communications. Our NLFSR is implemented using a plurality (3 in our implementation) of LFSRs. The outputs of 64 bits from each LFSR are combined using 64 encoded majority functions, where the majority function used for any bit is changed at every clock cycle. We demonstrate that our NLFSR can generate keys which may be used for OC-768 optical fiber communication, which operates at 40 Gbps. The keys from our NLFSR pass all the tests in the NIST suite, which is a defacto benchmark used in industry to evaluate the quality of ciphers. Pey-Chang Kent Lin, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | Efficient arithmetic sum-of-product (SOP) based Multiple Constant Multiplication (MCM) for FFTabstractIn this paper, we present an arithmetic sum-of-products (SOP) based realization of the general Multiple Constant Multiplication (MCM) algorithm. We also propose an enhanced SOP based algorithm, which uses Partial Max-SAT (PMSAT) to further optimize the SOP. The enhanced algorithm attempts to reduce the number of rows (partial products) of the SOP, by i) shifting coefficients to realize other coefficients when possible, ii) exploring multiple implementations of each coefficient using a Minimal Signed Digit (MSD) format and iii) exploiting the mutual exclusiveness within certain groups of partial products. Hardware implementations of the Fast Fourier Transform (FFT) algorithm require the incoming data to be multiplied by one of several constant coefficients. We test/validate it for FFT, which is an important problem. We compare our SOP-based architectures with the best existing implementation of MCM for FFT (which utilizes a cascade of adders), and show that our approaches show a significant improvement in area and delay. Our architecture was synthesized using 65nm technology libraries. Vinay Karkala, Joseph Wanstrath, Travis Lacour, Sunil P. Khatri |
ICCAD | 4 |
| 2010 | An efficient pulse flip-flop based launch-on-shift scan cellabstractAt-speed testing is essential for VLSI ICs implemented in nanometer technologies, operating at high clock speeds. Traditional scan based methodologies can be used for at-speed testing using a transition delay fault model. There are two common techniques to launch the transition-launch-on-shift (LOS) and launch-on-capture (LOC). LOS gives better fault coverage than LOC, but the main drawback of LOS is its requirement of a global at-speed scan enable (SE) signal that needs to be distributed across the IC. In this paper, we propose a pulsed flip-flop based LOS scan cell (PUFLOS cell) and a fast local scan enable generation circuit Our pulsed flip-flop based scan cell has 23.2% lower power dissipation and 27.3% better timing than a conventional muxed D-flip-flop based LOS scan cell. The layout area of our PUFLOS cell is 21% smaller than conventional LOS scan cell. Monte Carlo simulations demonstrate that our design is more robust to process variations than the conventional scan cell. Sunil P. Khatri |
ISCAS | 2 |
| 2010 | Fault Table Computation on GPUs
Kanupriya Gulati, Sunil P. Khatri |
J. Electron. Test. | 2 |
| 2010 | A Simultaneous Input Vector Control and Circuit Modification Technique to Reduce Leakage with Zero Delay PenaltyabstractLeakage power currently comprises a large fraction of the total power consumption of an IC. Techniques to minimize leakage have been researched widely. However, most approaches to reducing leakage have an associated performance penalty. In this article, we present an approach which minimizes leakage by simultaneously modifying the circuit while deriving the input vector that minimizes leakage. In our approach, we selectively modify a gate so that its output (in sleep mode) is in a state which helps minimize the leakage of other gates in its transitive fanout. Gate replacement is performed in a slack-aware manner, to minimize the resulting delay penalty. One of the major advantages of our technique is that we achieve a significant reduction in leakage without increasing the delay of the circuit. Nikhil Jayakumar, Sunil P. Khatri |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2009 | Efficient analytical determination of the SEU-induced pulse shapeabstractSingle event upsets (SEUs) have become problematic for both combinational and sequential circuits in the deep sub-micron era due to device scaling, lowered supply voltages and higher operating frequencies. To design radiation tolerant circuits efficiently, techniques are required to analyze the effects of a radiation particle strike on a circuit early in the design flow, and hence evaluate the circuit's resilience to SEU events. For an accurate estimation of the SEU tolerance of a circuit, it is important to consider the effects of electrical masking. This is typically done by performing circuit simulations, which are slow. In this paper, we present an analytical model for the determination of the shape of radiation-induced voltage glitches in combinational circuits. The output of our approach can be propagated to the primary outputs of the circuit using existing tools, thereby modeling the effects of electrical masking. This enables an accurate and quick evaluation of the SEU robustness of a circuit. Experimental results demonstrate that our model is very accurate, with a very low root mean square percentage error in the estimation of the shape of the voltage glitch of (4.5%) compared to SPICE. Our model gains its accuracy by using a non-linear model for the load current of the gate, and by considering the effect of taubetaon the radiation induced voltage glitch. Our analytical model is very fast (275 times faster than SPICE) and accurate, and can therefore be easily incorporated in a design flow to estimate the SEU tolerance of circuits early in the design process. Rajesh Garg, Sunil P. Khatri |
ASP-DAC | 2 |
| 2009 | Fast circuit simulation on graphics processing unitsabstractSPICE based circuit simulation is a traditional workhorse in the VLSI design process. Given the pivotal role of SPICE in the IC design flow, there has been significant interest in accelerating SPICE. Since a large fraction (on average 75%) of the SPICE runtime is spent in evaluating transistor model equations, a significant speedup can be availed if these evaluations are accelerated. This paper reports on our early efforts to accelerate transistor model evaluations using a Graphics Processing Unit (GPU). We have integrated this accelerator with a commercial fast SPICE tool. Our experiments demonstrate that significant speedups (2.36times on average) can be obtained. The asymptotic speedup that can be obtained is about 4times. We demonstrate that with circuits consisting of as few as about 1000 transistors, speedups in the neighborhood of this asymptotic value can be obtained. By utilizing the recently announced (but not currently available) quad GPU systems, this speedup could be enhanced further, especially for larger designs. Kanupriya Gulati, John F. Croix, Sunil P. Khatri, Rahm Shastry |
ASP-DAC | 3 |
| 2009 | Accelerating statistical static timing analysis using graphics processing unitsabstractIn this paper, we explore the implementation of Monte Carlo based statistical static timing analysis (SSTA) on a graphics processing unit (GPU). SSTA via Monte Carlo simulations is a computationally expensive, but important step required to achieve design timing closure. It provides an accurate estimate of delay variations and their impact on design yield. The large number of threads that can be computed in parallel on a GPU suggests a natural fit for the problem of Monte Carlo based SSTA to the GPU platform. Our implementation performs multiple delay simulations at a single gate in parallel. A parallel implementation of the Mersenne Twister pseudo-random number generator on the GPU, followed by box-Muller transformations (also implemented on the GPU) is used for generating gate delay numbers from a normal distribution. The mu and sigma of the pin-to-output delay distributions for all inputs and for every gate, are obtained using a memory lookup, which benefits from the large memory bandwidth of the GPU. Threads which execute in parallel have no data/control dependencies on each other. All threads compute identical instructions, but on different data, as required by the single instruction multiple data (SIMD) programming semantics of the GPU. Our approach is implemented on a NVIDIA GeForce GTX 8800 GPU card. Our results indicate that our approach can obtain an average speedup of about 260times as compared to a serial CPU implementation. With the recently announced quad 8800 GPU cards, we estimate that our approach would attain a speedup of over 785times. The correctness of the Monte Carlo based SSTA implemented on a GPU has been verified by comparing its results with a CPU based implementation. Kanupriya Gulati, Sunil P. Khatri |
ASP-DAC | 2 |
| 2009 | Closed-loop modeling of power and temperature profiles of FPGAsabstractIn recent times, the contribution of leakage power to the total power consumption of a chip has been increasing at an alarming rate. Leakage power is expected to exceed dynamic power in newer process technologies. Since leakage exhibits an exponential increase with temperature, it is possible that the high leakage of an IC causes a temperature increase, which in turn causes an increase in leakage, and so on, until the IC fails due to overheating. At the very least, this may cause the temperature and power consumption of the IC to be poorly estimated by traditional thermal or power modeling techniques. We developed a framework to model this situation in an FPGA context. Our CAD framework accurately models the total power consumption of the design at a given temperature, finds the thermal profile of the IC under this power consumption, and then uses this new thermal information to update the power consumption. This is iterated until the temperature of the IC converges, or until the temperatures on the die exceed a safe value. The iterations are very fast, due to the use of accurate and compact mathematical macromodels for leakage and temperature computation in the inner loop. We have exhaustively verified the fidelity of all our leakage macromodels. They estimate the leakage, at any temperature, to within 3% of the values generated by SPICE, while providing greater than four orders of magnitude speedup over explicit SPICE runs. Our experiments show that this model helps avoid an incorrect estimation of chip temperature and total power consumption, and also helps detect the increase in device temperature beyond a safe value. The average (maximum) error of our temperature estimates has been found to be within 1% (2.5%) compared to a full-chip 3D temperature modeling tool. Kanupriya Gulati, Sunil P. Khatri, Peng Li 0001 |
FPGA | 2 |
| 2009 | Low power and high performance sram design using bank-based selective forward body biasabstractLeakage power consumption is a large fraction of the total power consumption in contemporary VLSI designs. Since memories occupy a large portion of the total area of many high-performance ICs, it is crucial to reduce the leakage energy of memories. This problem is particularly aggravated for memories implemented in the 45nm technology node, since these processes exhibit significantly higher leakage power. For these memories, leakage is a significant problem not only from a power point of view, but also from a performance degradation standpoint. In this paper, we quantify this problem and provide a solution, using a 512KByte SRAM implemented in a 45nm bulk process as a design example. We show that implementing the SRAM as a monolithic memory results in increased delay as well as power. We illustrate a methodology to optimally reduce leakage power and improve performance in memories by splitting the memory array into word line groups (WLGs) which are selectively forward body biased when accessed. We present a derivation of optimal number of WLGs and the forward body bias voltage value, and show that our approach results in a 9:2% access time reduction, and a 53:4% reduction in power during a read operation. Our approach also achieves an 18% reduction in power during a write operation and a 69% leakage power improvement. The area overhead of our scheme is 7:2% compared to a monolithic memory. Kalyana C. Bollapalli, Rajesh Garg, Kanupriya Gulati, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 4 |
| 2009 | Robust window-based multi-node technology-independent logic minimizationabstractMulti-node optimization using Boolean relations is a powerful approach for network minimization. In this paper, we present an algorithm to perform Boolean relation-based multi-node optimization using a robust, fast and memory efficient algorithm. In particular, we simultaneously optimize two nodes at a time. The robustness of our approach arises from the use of a window-based technique for computing these Boolean relations. Secondly, we perform early quantification during the computation, keeping memory utilization low. Finally, we employ smart heuristics for selecting the node pair to be optimized simultaneously. Jeff L. Cobb, Kanupriya Gulati, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 3 |
| 2009 | Introduction to GPU programming for EDAabstractAdvances in GPU technology have propelled the GPU into arenas far afield from the traditional, isolated roles they have previously played. With hundreds of processing units in a single GPU, substantial speedups can be achieved by harnessing their power to augment the performance of the traditional single- or multi-core CPU on certain compute-intensive applications. However, utilizing the GPU requires both a change in the programmer's traditional algorithmic model as well as a judicious selection of algorithm being used for the problem. This paper reviews the GPU architecture and the tools available to utilize this valuable resource. It also provides insight into the type of problem best suited for the GPU as well as programming styles required to fully harness the power of the GPU. We present examples of specific EDA algorithms that can benefit from GPU acceleration, using both the CUDA and OpenCL environments. John F. Croix, Sunil P. Khatri |
ICCAD | 2 |
| 2009 | On-chip bidirectional wiring for heavily pipelined systems using network codingabstractIn this paper, we describe a low-area, reduced-power on-chip point-to-point bidirectional communication scheme for heavily pipelined systems. When data needs to be transmitted bidirectionally between two on-chip locations, the traditional approach resorts to either using two unidirectional wires, or to using a single wire (with a unidirectional transfer at any given time instant). In contrast, our bidirectional communication scheme allows data to be transmitted simultaneously between two on-chip locations, with a single wire performing the bidirectional data transfer. Our approach borrows ideas from the emerging area of network coding (in the field of communication). By utilizing coding units (which also serve the purpose of buffering the signals) along the wire between the two endpoints, we are able to achieve the same throughput as a traditional approach, while reducing the total area utilization by about 49.8% (thereby reducing routing congestion), and the total power consumption by about 11.5%. The area and power results include the contribution of routing wires, coding units, drivers, the clock distribution network and the required reset wire. Our bidirectional communication approach is ideally suited for heavily pipelined data intensive systems. Kalyana C. Bollapalli, Rajesh Garg, Kanupriya Gulati, Sunil P. Khatri |
ICCD | 4 |
| 2009 | 3D simulation and analysis of the radiation tolerance of voltage scaled digital circuitabstractIn recent times, dynamic supply voltage scaling (DVS) has been extensively employed to minimize the power and energy of VLSI systems. Also, sub-threshold circuits are becoming more popular. At the same time, the reliability of VLSI systems has become a major concern under Single Event Upsets (SEUs). SEUs are very problematic even for circuits operating at nominal voltages. With the increasing demand for low power reliable systems, it is therefore necessary to harden DVS and sub-threshold circuits efficiently. In this paper, we perform 3D simulations of radiation particle strikes in an inverter implemented using DVS and sub-threshold design. We analyze the sensitivity of the inverter to radiation particle strikes by varying the inverter size, the inverter load, the supply voltage (VDD) and the energy of the radiation particles. From these 3D simulations, we make several observations which are important to consider during radiation hardening of DVS and sub-threshold circuits. Based on these observations, we propose several guidelines for radiation hardening of DVS and sub-threshold circuit designs. These guidelines suggest that the traditional radiation hardening approaches need to be revisited for DVS and sub-threshold designs. We also propose a charge collection model for DVS circuits. Our model can accurately estimate (with an average error of 6.3%) the charge collected at the output of a gate for different supply voltages and different gate sizes for medium and high energy particle strikes. The parameters of our charge collection model can be included in SPICE model cards of transistors, to improve the accuracy of SPICE based radiation simulations for DVS circuits. Rajesh Garg, Sunil P. Khatri |
ICCD | 2 |
| 2009 | A PLL design based on a standing wave resonant oscillatorabstractIn this paper, we present a new continuously variable high frequency standing wave oscillator, and demonstrate its use in generating the phase locked clock signal of a digital IC. The ring based standing wave resonant oscillator is implemented with a plurality of wires connected in a mobius configuration, with a cross coupled inverter pair connected across the wires. The oscillation frequency can be modulated by two means. Coarse modification is achieved by altering the number of wires in the ring that participate in the oscillation, by driving a digital word to a set of passgates which are connected to each wire in the ring. Fine tuning of the oscillation frequency is achieved by varying the body bias voltage of both the PMOS transistors in the cross coupled inverter pair which sustains the oscillations in the resonant ring. We have validated our PLL design in a 90 nm process technology. 3D parasitic RLCs for our oscillator simulations were extracted, with skin effect accounted for. Our PLL has been implemented to provide a frequency locking range from ~6 GHz to ~9 GHz, with a center frequency of 7.5 GHz. The oscillator alone consumes about 25 mW of power, and the complete PLL consumes a power of 28.5 mW. The observed jitter of the PLL is 2.56%. Vinay Karkala, Kalyana C. Bollapalli, Rajesh Garg, Sunil P. Khatri |
ICCD | 4 |
| 2009 | A robust pulsed flip-flop and its use in enhanced scan designabstractDelay faults are frequently encountered in nanometer technologies. Therefore, it is critical to detect these faults during factory test. Testing for a delay fault requires the application of a pair of test vectors in an at-speed manner. To maximize the delay fault detection capability, it is desired that the vectors in this pair are independent. Independent vector pairs cannot always be applied to a circuit implemented with standard scan design approaches. However, this can be achieved by using enhanced scan flip-flops, which store two bits of data. This paper has two contributions. First, we develop a pulsed flip-flop (PFF) design. Second, we present an enhanced scan flipflop design, based on our PFF circuit. We have compared the performance of our pulse based flip-flop with recently published pulse based flip-flop designs, as well as a traditional master-slave D flip-flop. Our PFF shows significant improvements in power and timing compared to the other designs. Our pulse based enhanced scan flip-flop (PESFF) has 13% lower power dissipation and 26% better timing than a conventional D flipflop based enhanced scan flip-flop (DESFF). The layout area of our PESFF is 5.2% smaller than the DESFF. Monte Carlo simulations demonstrate that our design is more robust to process variations than the DESFF. Kalyana C. Bollapalli, Rajesh Garg, Tarun Soni, Sunil P. Khatri |
ICCD | 5 |
| 2009 | A radiation tolerant Phase Locked Loop design for digital electronicsabstractWith decreasing feature sizes, lowered supply voltages and increasing operating frequencies, the radiation tolerance of digital circuits is becoming an increasingly important problem. Many radiation hardening techniques have been presented in the literature for combinational as well as sequential logic. However, the radiation tolerance of clock generation circuitry has received scant attention to date. Recently, it has been shown that in the deep submicron regime, the clock network contributes significantly to the chip level Soft Error Rate (SER). The on-chip Phase Locked Loop (PLL) is particularly vulnerable to radiation strikes. In this paper, we present a radiation hardened PLL design. Each of the components of this design - the voltage controlled oscillator (VCO), the phase frequency detector (PFD) and the loop filter are designed in a radiation tolerant manner. Whenever possible, the circuit elements used in our PLL exploit the fact that if a gate is implemented using only PMOS (NMOS) transistors then a radiation particle strike can result only in a logic 0 to 1 (1 to 0) flip. By separating the PMOS and NMOS devices, and splitting the gate output into two signals, extreme high levels of radiation tolerance are obtained. Our PLL is tested for radiation immunity for critical charge values up to 250fC. Our results demonstrate that over a large number of radiation strikes on a number of sensitive nodes in our design, the worst case jitter is just 18%. In the worst case, our PLL returns to the locked state in 16 cycles of the VCO clock, after a radiation strike. Vinay Karkala, Rajesh Garg, Tanuj Jindal, Sunil P. Khatri |
ICCD | 5 |
| 2009 | Sorting Binary Numbers in Hardware - A Novel Algorithm and its ImplementationabstractThis paper describes a novel algorithm for sorting binary numbers in hardware, along with a custom VLSI hardware design for the same. For sorting n, k-bit binary numbers, our proposed algorithm takes O(n + 2 k) time. Sorting is performed by assigning relative ranks to the input numbers. A rank matrix of size n times n is used to store ranks. Each row of the rank matrix corresponds to one of the n numbers, and it stores a single non-zero entry. The position of this entry represents the relative rank of the corresponding number. In the worst case, our algorithm requires n+2 k clock cycles for assigning the final ranks. We start with a condition in which each number has an identical rank. In each of clock cycle, ranks are iteratively updated until the final ranks are determined after n+2 k clock cycles. The proposed algorithm is implemented in a 65 nm process, using a custom design approach to obtain a fast circuit. Our design is significantly faster than the fastest reported hardware sorting engine, with area performance which is superior for larger numbers. Srikanth Alaparthi, Kanupriya Gulati, Sunil P. Khatri |
ISCAS | 3 |
| 2009 | FPGA-based hardware acceleration for Boolean satisfiabilityabstractWe present an FPGA-based hardware solution to the Boolean satisfiability (SAT) problem, with the main goals of scalability and speedup. In our approach the traversal of the implication graph as well as conflict clause generation are performed in hardware, in parallel. The experimental results and their analysis, along with the performance models are discussed. We show that an order of magnitude improvement in runtime can be obtained over MiniSAT (the best-in-class software based approach) by using a Virtex-4 (XC4VFX140) FPGA device. The resulting system can handle instances with as many as 10K variables and 280K clauses. Kanupriya Gulati, Suganth Paul, Sunil P. Khatri, Srinivas Patil, Abhijit Jas |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2009 | Efficient On-Chip Crosstalk Avoidance CODEC DesignabstractInterconnect delay has become a limiting factor for circuit performance in deep sub-micrometer designs. As the crosstalk in an on-chip bus is highly dependent on the data patterns transmitted on the bus, different crosstalk avoidance coding schemes have been proposed to boost the bus speed and/or reduce the overall energy consumption. Despite the availability of the codes, no systematic mapping of datawords to codewords has been proposed for CODEC design. This is mainly due to the nonlinear nature of the crosstalk avoidance codes (CAC). The lack of practical CODEC construction schemes has hampered the use of such codes in practical designs. This work presents guidelines for the CODEC design of the “forbidden pattern free crosstalk avoidance code” (FPF-CAC). We analyze the properties of the FPF-CAC and show that mathematically, a mapping scheme exists based on the representation of numbers in the Fibonacci numeral system. Our first proposed CODEC design offers a near-optimal area overhead performance. An improved version of the CODEC is then presented, which achieves theoretical optimal performance. We also investigate the implementation details of the CODECs, including design complexity and the speed. Optimization schemes are provided to reduce the size of the CODEC and improve its speed. Chunjie Duan, Victor H. Cordero Calle, Sunil P. Khatri |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2009 | Circuit-Level Design Approaches for Radiation-Hard Digital ElectronicsabstractIn this paper, we present a novel circuit design approach for radiation hardened digital electronics. Our approach is based on the use of shadow gates, whose task it is to protect the primary gate in case it is struck by a heavy cosmic ion. We locally duplicate the gate to be protected, and connect a pair of diode-connected transistors (or diodes) between the outputs of the original and shadow gates. These transistors turn on when the voltages of the two gates deviate during a radiation strike. Our experiments show that at the level of a single gate, our circuit structure has a delay overhead about 1.76% on average, and an area overhead of 277%. At the circuit level, however, we do not need to protect all gates. We present a methodology to selectively protect specific gates of the circuit in a manner that guarantees radiation tolerance for the entire circuit. With this methodology, we demonstrate that at the circuit level, the average delay overhead is about 3% and the average placed-and-routed area overhead is 28%, compared to an unprotected circuit (for delay mapped designs). We also propose an improved circuit protection algorithm to reduce the area overhead associated with our approach. With this approach for circuit protection, the area and delay overheads are further lowered. Rajesh Garg, Nikhil Jayakumar, Sunil P. Khatri, Gwan S. Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2009 | A Fast Hardware Approach for Approximate, Efficient Logarithm and Antilogarithm ComputationsabstractThe realization of functions such as log() and antilog() in hardware is of considerable relevance, due to their importance in several computing applications. In this paper, we present an approach to compute log() and antilog() in hardware. Our approach is based on a table lookup, followed by an interpolation step. The interpolation step is implemented in combinational logic, in a field-programmable gate array (FPGA), resulting in an area-efficient, fast design. The novelty of our approach lies in the fact that we perform interpolation efficiently, without the need to perform multiplication or division, and our method performs both the log() and antilog() operation using the same hardware architecture. We compare our work with existing methods, and show that our approach results in significantly lower memory resource utilization, for the same approximation errors. Also our method scales very well with an increase in the required accuracy, compared to existing techniques. Suganth Paul, Nikhil Jayakumar, Sunil P. Khatri |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | Forbidden transition free crosstalk avoidance CODEC designabstractIn this work, we present a CODEC design for the forbidden transition free crosstalk avoidance code. Our mapping and coding scheme is based on the Fibonacci numeral system and the mathematical analysis shows that all numbers can be represented by FTF vectors in the Fibonacci numeral system (FNS). The proposed CODEC design is highly efficient, modular and can be easily combined with a bus partitioning technique. We also investigate the implementation issues and our experimental results show that the proposed CODEC complexity is orders of magnitude better compared to the brute force implementation. Compared to the best existing approaches, we achieve a 17% improvement in logic complexity. A high speed design can be achieved through pipelining. Chunjie Duan, Chengyu Zhu, Sunil P. Khatri |
DAC | 3 |
| 2008 | A fast, analytical estimator for the SEU-induced pulse width in combinational designsabstractSingle event upsets (SEUs) are becoming increasingly problematic for both combinational and sequential circuits with device scaling, lower supply voltages and higher operating frequencies. To design radiation tolerant circuits efficiently, techniques are required to analyze the effects of a particle strike on a circuit early in the design flow and also to evaluate the circuit's resilience to SEU events. In this paper, we present an analytical model for SEU induced transients in combinational circuits. The pulse width of the voltage glitch due to an SEU event is a good measure of SEU robustness and our model efficiently computes it for any combinational gate. The experimental results demonstrate that our model is very accurate with a very low pulse width estimation error of 4% compared to SPICE. Our model gains its accuracy by using a non-linear transistor current model, and by considering the effect of τβ of the radiation induced current pulse. Our analytical model is very fast and accurate, and can therefore be easily incorporated in a design flow to implement SEU tolerant circuits. Rajesh Garg, Charu Nagpal, Sunil P. Khatri |
DAC | 3 |
| 2008 | Towards acceleration of fault simulation using graphics processing unitsabstractIn this paper, we explore the implementation of fault simulation on a Graphics Processing Unit (GPU). In particular, we implement a fault simulator that exploits thread level parallelism. Fault simulation is inherently parallelizable, and the large number of threads that can be computed in parallel on a GPU results in a natural fit for the problem of fault simulation. Our implementation fault-simulates all the gates in a particular level of a circuit, including good and faulty circuit simulations, for all patterns, in parallel. Since GPUs have an extremely large memory bandwidth, we implement each of our fault simulation threads (which execute in parallel with no data dependencies) using memory lookup. Fault injection is also done along with gate evaluation, with each thread using a different fault injection mask. All threads compute identical instructions, but on different data, as required by the Single Instruction Multiple Data (SIMD) programming semantics of the GPU. Our results, implemented on a NVIDIA GeForce GTX 8800 GPU card, indicate that our approach is on average 35 x faster when compared to a commercial fault simulation engine. With the recently announced Tesla GPU servers housing up to eight GPUs, our approach would be potentially 238× faster. The correctness of the GPU based fault simulator has been verified by comparing its result with a CPU based fault simulator. Kanupriya Gulati, Sunil P. Khatri |
DAC | 2 |
| 2008 | Clock Distribution Scheme using Coplanar Transmission LinesabstractThe current work describes a new standing wave oscillator scheme aimed for clock propagation on coplanar transmission lines on a silicon die. The design is aimed for clock signaling in the gigahertz range (we are able to achieve clock rates of 8 GHz and above). The clock is transported as an oscillatory wave on a pair of conductors. An oscillatory standing wave is formed across a transmission line loop, which is connected beginning-to-end through a Mobius configuration. A single cross coupled inverter pair is required to maintain oscillation across the ring. The design is aimed to achieve low skew, low power and extreme high frequency global clock situations. The energy recycling nature of a standing wave along a transmission line allows us to keep very high frequencies oscillations along a conductor with almost no power consumption at all. A special wide input range driver was designed to convert the differential signals on the coplanar transmission lines into a square clock pulse for standard clock sinks. The design uses CMOS 90 nm BSim3v model cards for all simulations, with the transmission lines implemented on Metal8. Victor H. Cordero Calle, Sunil P. Khatri |
DATE | 2 |
| 2008 | Energy Efficient and High Speed On-Chip Ternary BusabstractWe propose two crosstalk reducing coding schemes using ternary busses. In addition to low power consumption and reduced delay, our schemes offer other advantages over binary coding schemes such as zero area overhead and simple, regular and fast codec design. Chunjie Duan, Sunil P. Khatri |
DATE | 2 |
| 2008 | A Single-supply True Voltage Level ShifterabstractWhen a signal traverses on-chip voltage domains, a level shifter is required. Inverters can handle a high to low voltage shift with minimal leakage. For a low to high voltage level translation, inverters tend to consume a large amount of leakage power, and hence special circuits have been proposed for this type of translation. This paper reports a novel single-supply "true" (in the sense that it can handle a low to high, or high to low voltage level conversion) voltage level shifter, which can handle low-to-high and high-to-low voltage translation. Such a requirement arises in many modern ICs or systems-on-chip (SoCs). The use of single supply voltage reduces circuit complexity by eliminating the need for routing both supply voltages. The proposed circuit was extensively simulated in a 90 nm technology using SPICE. Simulation results demonstrate that the level shifter is able to perform voltage level shifting with low leakage for both low to high, as well as high to low voltage level translation. We have validated the correct operation of the proposed level shifter under process and temperature variations as well. Rajesh Garg, Gagandeep Mallarapu, Sunil P. Khatri |
DATE | 3 |
| 2008 | A Delay-efficient Radiation-hard Digital Design Approach Using CWSP ElementsabstractIn this paper, we present a radiation-hardened digital design approach. This approach is based on the use of code word state preserving (CWSP) elements at each flip-flop of the design, and leaving the rest of the design unaltered. The CWSP element provides 100% SET protection for glitch widths up to min{Dmin/2, (Dmax- Delta)/2}, where Dminand Dmaxare the minimum and maximum circuit delay respectively and Delta is an extra delay associated with our SET protection circuit. The CWSP circuit has two inputs - the latch output signal and the same signal delayed by a quantity delta. In case an SET error is detected, then the current computation is repeated, using the correct output, which is generated later in the same clock period by the CWSP element. Unlike previous approaches, we use the CWSP element in a secondary path and the CWSP logic is designed to minimally impact the critical delay path of the design. The delay penalty of our approach (averaged over several designs) is less than 1%. Thus our technique is applicable for high-speed designs, where the additional delay associated with SET protection must be kept at a minimum. Charu Nagpal, Rajesh Garg, Sunil P. Khatri |
DATE | 3 |
| 2008 | A lithography-friendly structured ASIC design approachabstractIntegrated circuit manufacturing costs are increasing with decreasing device feature sizes, due to significant increases in mask costs. At the same time, systematic processing variations due to optical proximity effects are also increasing, making it harder to predict the circuit behavior with fidelity. Therefore, there is a need to implement designs using regular circuit structures. In this paper, we present a new structured ASIC approach which utilizes an array of 2-input NAND gates. Our NAND2 array based circuit implementation reduces manufacturing costs, and design turn-around times because different designs can share the same masks up to the poly layer. The regular layout structure of our NAND2 array also helps in reducing systematic variations. We compared the performance of our NAND2 array with the ASIC approach by implementing several benchmark circuits using both methods. The experimental results demonstrate that on average, our approach has a delay penalty of 40%, an area penalty of 12%, and a power increase of 7%, compared to an ASIC design approach. This is better than the previously reported structured ASIC approaches. We also performed lithographical simulations of the poly and metal masks of the designs implemented using our approach as well as the ASIC design approach. These lithographical simulation results demonstrate that our approach has lower errors on the poly and the Metal1 layers by 7% and 24% respectively, compared to the ASIC approach. Salman Gopalani, Rajesh Garg, Sunil P. Khatri, Mosong Cheng |
ACM Great Lakes Symposium on VLSI | 3 |
| 2008 | Improving FPGA routability using network codingabstractWith current technology trends, FPGA routing is an important problem, since routing in FPGAs contributes significantly to delay and resource utilization, as compared to the logic portion of FPGAs. In this paper we improve the FPGA routing characteristics by applying the technique of network coding. Kanupriya Gulati, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 2 |
| 2008 | Pipelined network of PLA based circuit designabstractIn this paper, we present a pipelined Network of PLA based circuit design approach. Our approach can be used to realize an arbitrary logic circuit with an extremely high throughput and low latency. Using logic synthesis tools to decompose a logic circuit into this framework, and appropriately inserting "stutter" blocks to balance the logical depth of all paths in the decomposed circuit, we come up with a pipelined network of PLA netlist. We have demonstrated the effectiveness of the approach via SPICE simulations and layout generation experiments. Throughput, latency, and area are compared with competing approaches, demonstrating the power of this design style. We show that our approach has a 75% better throughput than the asynchronous micropipelining based technique, and a latency which is 63% that of the asynchronous scheme. Both techniques were implemented in a super-threshold fashion. We have also conducted Monte Carlo experiments to validate the approach under variations. Suganth Paul, Rajesh Garg, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 3 |
| 2008 | A robust, fast pulsed flip-flop designabstractHigh Speed VLSI design utilizes heavy pipelining, resulting in a large number of \nip-\nops in the circuit. Hence there is a strong motivation to design fast, low power and area e-cient ip-\nops. In this paper, we present a pulsed \nip-\nop design based on a novel pulse generator circuit. Our design achieves signicantly improved speed when compared to re-cent pulsed \nip-\nop design, as well as a traditional master-slave D \nip-\nop. Monte Carlo simulations demonstrate that our design is signicantly more robust to variations than the other \nip-\nops. Our design consumes low power as well. Also we have performed the layout of our design and shown that our layout area is smaller than a traditional D \nip-\nop. Arunprasad Venkatraman, Rajesh Garg, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 3 |
| 2008 | A novel, highly SEU tolerant digital circuit design approachabstractIn this paper, we present a new radiation tolerant CMOS standard cell library, and demonstrate its effectiveness in implementing radiation hardened digital circuits. We exploit the fact that if a gate is implemented using only PMOS (NMOS) transistors then a radiation particle strike can result only in logic a 0 to 1 (1 to 0) flip. Based on this observation, we derive our radiation hardened gates from regular static CMOS gates. In particular, we separate the PMOS and NMOS devices, and split the gate output into two signals. One of these outputs of our radiation tolerant gate is generated using PMOS transistors, and it drives other PMOS transistors (only) in its fanout. Similarly, the other output (generated from NMOS transistors) drives only other NMOS transistors in its fanout. Now, if a radiation particle strikes one of the outputs of the radiation tolerant gate, then the gates in the fanout enter a high-impedance state, and hence preserve their output values. Our radiation hardened gates exhibit an extremely high degree of SEU tolerance, which is validated at the circuit level. Using these gates, we also implement circuit level hardening based on logical masking, to selectively harden those gates in a circuit which contribute most to the soft error failure of the circuit. The gates with a low probability of logical masking are replaced by SEU tolerant gates from our new library, such that the digital design achieves a 90% soft error rate reduction. Experimental results demonstrate that this reduction is achieved with a modest layout area and delay penalty of 62% and 29% respectively, for area mapped designs. In contrast with existing approaches, our approach results in SEU immunity for extremely large critical charge values (>650fC). Rajesh Garg, Sunil P. Khatri |
ICCD | 2 |
| 2008 | Modeling dynamic stability of SRAMS in the presence of single event upsets (SEUs)abstractSRAM yield is very important from an economics viewpoint, because of the extensive use of memory in modern processors and SOCs. Therefore, SRAM stability analysis tools have become essential. SRAM stability analysis based on static noise margin (SNM) often results in pessimistic designs because SNM cannot capture the transient behavior of the noise. Therefore, to improve accuracy, dynamic stability analysis is required. The model presented in this paper performs dynamic stability analysis of an SRAM cell in the presence of an SEU event. The experimental results demonstrate that our model is very accurate, with a critical charge estimation error of 2.5% compared to HSPICE. The run-time of our model is also significantly lower (1200× lower) than the HSPICE run-time. Thus, our model enables the SRAM designer to quickly and accurately analyze stability during the design phase. Rajesh Garg, Peng Li 0001, Sunil P. Khatri |
ISCAS | 3 |
| 2008 | A probabilistic method to determine the minimum leakage vector for combinational designs in the presence of random PVT variations
Kanupriya Gulati, Nikhil Jayakumar, Sunil P. Khatri, D. M. H. Walker |
Integr. | 3 |
| 2008 | Resource sharing among mutually exclusive sum-of-product blocks for area reductionabstractIn state-of-the-art digital designs, arithmetic blocks consume a major portion of the total area of the IC. The arithmetic sum-of-product (SOP) is the most widely used arithmetic block. Some of the examples of SOP are adder, subtractor, multiplier, multiply-accumulator (MAC), squarer, chain-of-adders, incrementor, decrementor, etc. In this article, we introduce a novel, area-efficient architecture to share different SOP blocks which are used in a mutually exclusive manner. We implement the core functions of the largest SOP only once and reuse different parts of the core subblocks for all other SOP operations with the help of multiplexers. This architecture can be used in the nontiming-critical paths of the design, to save significant amounts of area. Our experimental data shows that the proposed sharing-based architecture results in about 37% area savings compared to the results obtained from a commercially available best-in-class datapath synthesis tool. In addition, our proposed shared implementation consumes about 18% less power. These improvements were verified on placed-and-routed designs as well. Sabyasachi Das, Sunil P. Khatri |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2008 | SAT-based ATPG using multilevel compatible don't-caresabstractIn a typical IC design flow, circuits are optimized using multilevel don't cares. The computed don't cares are discarded before Technology Mapping or Automatic Test Pattern Generation (ATPG). In this paper, we present two combinational ATPG algorithms for combinational designs. These algorithms utilize the multilevel don't cares that are computed for the design during technology independent logic optimization. They are based on Boolean Satisfiability (SAT), and utilize the single stuck-at fault model. Both algorithms make use of the Compatible Observability Don't Cares (CODCs) associated with nodes of the circuit, to speed up the ATPG process. For large circuits, both algorithms make use of approximate CODCs (ACODCs), which we can compute efficiently. Our first technique speeds up fault propagation by modifying the active clauses in the transitive fanout (TFO) of the fault site. In our second technique, we define new j - active variables for specific nodes in the transitive fanin (TFI) of the fault site. Using these j-active variables we write additional clauses to speed up fault justification. Experimental results demonstrate that the combination of these techniques (when using CODCs) results in an average reduction of 45% in ATPG runtimes. When ACODCs are used, a speed-up of about 30% is obtained in the ATPG run-times for large designs. We compare our method against a commercial structural ATPG tool as well. Our method is slower for small designs, but for large designs, we obtain a 31% average speedup over the commercial tool. Nikhil Saluja, Kanupriya Gulati, Sunil P. Khatri |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2008 | A Novel Hybrid Parallel-Prefix Adder Architecture With Efficient Timing-Area CharacteristicabstractTwo-operand binary addition is the most widely used arithmetic operation in modern datapath designs. To improve the efficiency of this operation, it is desirable to use an adder with good performance and area tradeoff characteristics. This paper presents an efficient carry-lookahead adder architecture based on the parallel-prefix computation graph. In our proposed method, we define the notion of triple-carry-operator, which computes the generate and propagate signals for a merged block which combines three adjacent blocks. We use this in conjunction with the classic approach of the carry-operator to compute the generate and propagate signals for a merged block combining two adjacent blocks. The timing-driven nature of the proposed design reduces the depth of the adder. In addition, we use a ripple-carry type of structure in the nontiming critical portion of the parallel-prefix computation network. These techniques help produce a good timing-area tradeoff characteristic. The experimental results indicate that our proposed adder is significantly faster than the popular Brent-Kung adder with some area overhead. On the adder hand, the proposed adder also shows marginally faster performance than the fast Kogge-Stone adder with significant area savings. Sabyasachi Das, Sunil P. Khatri |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Dynamically De-Skewable Clock Distribution MethodologyabstractIn a typical clock distribution scheme, a central clock signal is distributed to several sites on the integrated circuit (IC). Local regenerators at these sites buffer the clock signal for the logic in regions close to the regenerator. Minimizing the skew between the clocks at these regeneration sites is critical. In recent times, this is becoming harder due to increasing intra-die processing variations. In this paper, we describe a novel technique to distribute a clock signal from a central location to several sites on a VLSI IC. Our technique uses a buffered H-tree and includes circuitry to dynamically remove any skew that may result due to intra-die processing variations. While existing approaches to deskewing a clock tree have utilized several phase detection circuits (number of phase detectors dependent on the number of clock regenerators), our method requires only one phase detector. Also, in our approach, the resolution of the phase detector is inconsequential unlike existing techniques. Our deskewing technique can be applied dynamically, either at boot time or periodically during the operation of the IC. Using a six-level H-tree clock distribution network with process variations deliberately included, we demonstrate that our technique can reduce skews as high as 300 ps down to just 3 ps. We compare our clock tree with traditional buffered and unbuffered H-tree networks. Arjun Kapoor, Nikhil Jayakumar, Sunil P. Khatri |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | An algorithm to minimize leakage through simultaneous input vector control and circuit modificationabstractLeakage power currently comprises a large fraction of the total power consumption of an IC. Techniques to minimize leakage have been researched widely. In this paper, the authors present an approach which minimizes leakage by simultaneously modifying the circuit while deriving the input vector that minimizes leakage. In this approach, the authors selectively modify a gate so that its output (in sleep mode) is in a state which helps minimize the leakage of other gates in its transitive fanout. Gate replacement is performed in a slack-aware manner, to minimize the resulting delay penalty Nikhil Jayakumar, Sunil P. Khatri |
DATE | 2 |
| 2007 | Toggle Equivalence Preserving (TEP) Logic OptimizationabstractWe describe a procedure (called the TEP procedure) that, given a multi-output circuit M, builds another multi-output circuit M* that is toggle equivalent to M. The TEP procedure can be used in the following two scenarios. First, since for single- output circuits toggle equivalence means functional equivalence, the TEP procedure can be used in "regular" logic synthesis. Second, the TEP procedure enables a powerful synthesis method called LS_TE (Logic Synthesis preserving Toggle Equivalence). Given a circuit N and its partitioning into subcircuits NiLS_TE builds an optimized circuit N* by replacing subcircuits Niwith their toggle equivalent counterparts Ni. The replacement of Niwith N*iis done by the TEP procedure. We give results of optimizing single-output circuits by the TEP procedure and some preliminary results of using the TEP procedure in LS_TE. These results show the promise of the TEP procedure and LS_TE. Eugene Goldberg, Kanupriya Gulati, Sunil P. Khatri |
DSD | 3 |
| 2007 | A Structured ASIC Design Approach Using Pass Transistor LogicabstractIn this paper, we describe a structured ASIC design methodology which utilizes a regular, pre-fabricated array of pass transistor logic based if-then-else (ITE) cells as the building block for the circuit. Given a logic netlist, we first construct reduced order binary decision diagrams (ROBDDs) for the circuit in a partitioned manner, thereby allowing the approach to handle large designs. We place the ITE cells corresponding to the ROBDD nodes in a manner that minimizes crossings in the ROBDD graph. Our placement also effectively 'folds' the ITE cells of different variables into a single row, so as to obtain a layout with a more uniform distribution of ITE cells along each physical row of ITE cells. The design methodology has been demonstrated to implement sequential as well as combinational designs, by customizing the lowest 4 METAL layers along with their associated VIA layers. A low area and delay overhead is achieved, in comparison with an ASIC approach. In particular, the average delay (area) overhead is 1.5 times (3.41 times) for combinational designs and 2 times (6 times) for sequential designs Kanupriya Gulati, Nikhil Jayakumar, Sunil P. Khatri |
ISCAS | 3 |
| 2007 | A methodology for interconnect dimension determinationabstractThe determination of metal wire dimensions and inter-layer dielectric thicknesses has become increasingly important in recent times, as wire delays have begun to dominate transistor delays. In this paper, we propose metrics to guide the determination of these dimensions, along with results which describe the efficacy of these metrics. The first metric, which we refer to as cross-bar bandwidth(CBB), compares different wiring configurations in terms ofthe resulting bandwidth they can support in a square of unitsize. The second metric, called power-adjusted cross-bar bandwidth (PCBB), measures the bandwidth of a square, while accounting for the power consumed by the interconnect in the square. The application of these metrics suggests that the traditional approach (of fabricating interconnect of the finest pitch possible) may be sub-optimal with respect to either metric that we present. We have conducted our experiments for layers METAL1 through METAL4, although the approach is general and can be applied to the problem of determining interconnect dimensions for any layer. Without considering driver resistance, our approach yields up to 16% improvement in CBB and 19% improvement in PCBB. When drivers are modeled, the improvements are even greater. Jeff L. Cobb, Rajesh Garg, Sunil P. Khatri |
ISPD | 3 |
| 2007 | A Predictably Low-Leakage ASIC Design StyleabstractIn this paper, we describe a new low-leakage standard cell based application-specific integrated circuit (ASIC) design methodology. This design is based on the use of modified standard cells, designed to reduce leakage currents (by almost two orders of magnitude) in standby mode and also allow precise estimation of leakage current. For each cell in a standard cell library, two low-leakage variants of the cell are designed. If the inputs of a cell during the standby mode of operation are such that the output has a high value, we minimize the leakage in the pull-down network, and similarly we minimize leakage in the pull-up network if the output has a low value. In this manner, two low-leakage variants of each standard cell are obtained. While technology mapping a circuit, we determine the particular variant to utilize in each instance, so as to minimize leakage of the final mapped design. We have performed experiments to compare placed-and-routed area, leakage and delays of this new methodology against Multithreshold CMOS (MTCMOS) and a regular standard cell based design style. The results show that our new methodology (which we call the "HL" methodology) has better speed and area characteristics than MTCMOS implementations. The leakage current for HL designs can be dramatically lower than the worst-case leakage of MTCMOS based designs, and two orders of magnitude lower than the leakage of traditional standard cells. An ASIC design implemented in MTCMOS would require the use of separate power and ground supplies for latches and combinational logic, while our methodology does away with such a requirement. Another advantage of our methodology is that the leakage is precisely estimable, in contrast with MTCMOS. Our primary contribution in this paper is a new low leakage design style for static CMOS designs. In addition, we also discuss techniques to reduce leakage in dynamic (domino logic) designs Nikhil Jayakumar, Sunil P. Khatri |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Controlling inductive cross-talk and power in off-chip buses using CODECsabstractThe parasitic inductances within IC packaging cause supply bounce as well as glitches on the signal pins, significantly limiting the frequency of high-speed inter-chip communication. Also, off-chip communication contributes a large fraction of the total system power. Until recently, the parasitic inductance problem was addressed by aggressive package design, which is expensive. In this work we present a technique to encode the off-chip data transmission to i) limit bounce on the supplies, ii) reduce glitching caused by inductive signal coupling from neighboring signals, iii) limit the edge degradation of signals due to mutually inducted voltages from neighboring switching signals and iv) control the total power consumption of the I/O logic. All these factors are modeled in a unified mathematical framework. Our experimental results show that the proposed encoding based techniques result in reduced supply bounce and signal glitching due to inductive cross-talk, closely matching the theoretical predictions. Also, we show that the bus size overhead is reasonable even after stringent power reduction constraints are imposed. We demonstrate that the overall bandwidth of a bus actually increases by 100% over an unencoded bus, using our technique with inductive constraints only (even after accounting for the encoding overhead). When the power constraints were added (to limit the power to 20% of worst case switching power) in addition to the inductive constraints, the bandwidth was again 100% improved over the unencoded bus. The asymptotic bus size overhead depends on how stringent the user-specified power and inductive cross-talk parameters are. We have validated our approach by simulating it in an ASIC setting as well as prototyping and testing it in an FPGA environment. Brock J. LaMeres, Kanupriya Gulati, Sunil P. Khatri |
ASP-DAC | 3 |
| 2006 | A design approach for radiation-hard digital electronicsabstractsunilkhatri at tamu.edu In this paper, we present a novel circuit design approach for radiation hardened digital electronics. Our approach is based on the use of shadow gates, whose task it is to protect the primary gate in case it is struck by a heavy cosmic ion. We locally duplicate the gate to be protected, and connect a pair of transistors (or diodes) between the outputs of the original and shadow gates. These transistors turn on when the voltages of the two gates deviate during a radiation strike. Our experiments show that at the level of a single gate, our circuit structure has a delay overhead of about 4% on average, and an area overhead of over 100%. At the circuit level, however, we do not need to protect all gates. We present a methodology to selectively protect specific gates of the circuit in a manner that guarantees radiation tolerance for the entire circuit. With this methodology, we demonstrate that at the circuit level, the delay overhead is about 4 % and the placed-and-routed area overhead is 30%, compared to an unprotected circuit (for delay mapped designs) Rajesh Garg, Nikhil Jayakumar, Sunil P. Khatri, Gwan S. Choi |
DAC | 3 |
| 2006 | A PLA based asynchronous micropipelining approach for subthreshold circuit designabstractPower consumption is a dominant issue in contemporary circuit design. Sub-threshold circuit design is an appealing means to dramatically reduce this power consumption. However, sub-threshold designs suffer from the drawback of being significantly slower than traditional designs. To reduce the speed gap between sub-threshold and traditional designs, we propose a sub-threshold circuit design approach based on asynchronous micropipelining of a levelized network of PLAs. We describe the handshaking protocol, circuit design and logic synthesis issues in this context. Our preliminary results demonstrate that by using our approach, a design can be sped up by about 7x, with an area penalty of 47%. Further, our approach yields an energy improvement of about 4x, compared to a traditional network of PLA design. Our approach is quite general, and can be applied to traditional circuits as well. Nikhil Jayakumar, Rajesh Garg, Bruce Gamache, Sunil P. Khatri |
DAC | 4 |
| 2006 | Bus stuttering: an encoding technique to reduce inductive noise in off-chip data transmissionabstractSimultaneous switching noise due to inductance in VLSI packaging is a significant limitation to system performance. The inductive parasitics within IC packaging causes bounce on the power supply pins in addition to glitches and rise-time degradation on the signal pins. These factors bound the maximum performance of off-chip busses, which limits overall system performance. Until recently, the parasitic inductance problem was addressed by aggressive package design which attempts to decrease the total inductance in the package interconnect. In this work we present an encoding technique for off-chip data transmission to limit bounce on the supplies and reduce inductive signal coupling. This is accomplished by inserting intermediate (henceforth called "stutter") states in the data, transmission to bind the maximum number of signals that switch simultaneously, thereby limiting the overall inductive noise. Bus stuttering is cheaper than expensive package design since it increases the bus performance without changing the package. We demonstrate that bus stuttering can bound the maximum amount of inductive noise, which results in increased bus performance even after accounting for the encoding overhead. Our results show that the performance of an encoded bus can be increased up to 225% over using un-encoded data. In addition, synthesis results of the encoder in a TSMC 0.13mum process show that the encoder size and delay are negligible in a modern VLSI design Brock J. LaMeres, Sunil P. Khatri |
DATE | 2 |
| 2006 | Resource and delay efficient matrix multiplication using newer FPGA devicesabstractMatrix multiplication is a fundamental building block for many applications including image processing, coding, and digital signal processing. This paper presents a delay and resource efficient methodology for implementing integer and floating point matrix multiplication using FPGAs. We present a scalable architecture that provides a significant reduction in total computation time and resource utilization over previous solutions. The improvements of our method are attributed to a new method to compute partial products in parallel, utilizing the new features of modern FPGAs. The implementation of our algorithm for various matrix dimensions using Xilinx FPGAs is also described. When compared with the best reported previous method, our approach achieves an improvement in the parallelization of 60% for 64-bit floating point computations. Scott J. Campbell, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 2 |
| 2006 | A design flow to optimize circuit delay by using standard cells and PLAsabstractThis paper presents a design flow that optimizes a standard cell based circuit for performance by implementing critical paths in a Programmable Logic Array (PLA). Given a standard-cell based circuit as input, our approach iteratively extracts critical paths from this circuit, which are then implemented using a PLA circuit. PLAs are a good candidate for such an approach, since they exhibit a gradual increase in delay as additional vectors are added. In subsequent iterations, these critical paths are treated as don't cares, allowing the standard cell based design to be simplified after each iteration. The final design consists of a portion which is implemented using a PLA, and another portion which is implemented using standard cells. We demonstrate that on average, our approach can achieve about 22.5% improvement in the SPICE based delay of a design, along with a placed-and-routed area improvement of 11%. Rajesh Garg, Kanupriya Gulati, Nikhil Jayakumar, Sunil P. Khatri |
ACM Great Lakes Symposium on VLSI | 6 |
| 2006 | Implementation of MOSFET based capacitors for digital applicationsabstractTo compensate for the growing effect of process, temperature and voltage variations in digital ICs, several dynamic approaches have been proposed. These approaches require the use of capacitors to offset the effects of variations. Although MOSFET based capacitors are a natural choice, such capacitances vary significantly depending on the applied voltage. In this paper, we propose techniques to make the capacitance of MOSFET based capacitors relatively constant over the applied voltage. We study approaches that are based on gate as well as diffusion capacitors. We study the trade-off (in terms of capacitance variation, requirement of a bias voltage and area efficiency) of all the proposed schemes. Simulation results are presented for two process technologies. We show that by using our techniques the capacitance variation across applied voltage can be reduced from 75% down to 3% for a 70nm process technology. Sunil P. Khatri, Takis Zourntos |
ACM Great Lakes Symposium on VLSI | 2 |
| 2006 | A High-Speed Fully-Programmable VLSI Decoder for Regular LDPC CodesabstractThis paper presents a VLSI implementation of a Low-Density Parity Check (LDPC) decoder that achieves 2.4 Gbps throughput yet permits real-time configuration of (1) rate, (2) code length, and (3) the parity equations. This decoder can be programmed in the field, much like an FPGA. We describe the architectural, circuit-level and layout-level details of our implementation. Our design can handle variable rate codes of length up to 1024, and is implemented in a 0.1 μm VLSI fabrication process. Our design has a die size of 12mm by 8mm and a power consumption of 7W. This implementation can extended to handle longer codes in a partially parallel manner, and allow for on-the-fly modification of the code. Euncheol Kim, Nikhil Jayakumar, Pankaj Bhagawat, Anand Selvarathinam, Gwan S. Choi, Sunil P. Khatri |
ICASSP (3) | 6 |
| 2006 | Network coding for routability improvement in VLSIabstractWith the standard approach for establishing multicast connections over a network, network nodes are utilized to forward and duplicate the packets received over the incoming links. Recently, there has been a significant interest in a novel paradigm of network coding. Network coding generalizes the traditional routing approach by allowing the network nodes to generate new packets by performing algebraic operations on packets received over the incoming links. It has been shown that network coding can increase the throughput of multicast communication. In this paper, we explore the benefits of network coding for improving the routing characteristics of VLSI designs. We demonstrate that when data has to be routed across the IC, it is often beneficial to perform network coding. Initial results demonstrate that network coding can result in a healthy reduction in wire length, wire area, interconnect power as well as the active area associated with the interconnects. This comes at a small delay penalty. Nikhil Jayakumar, Sunil P. Khatri, Kanupriya Gulati, Alexander Sprintson |
ICCAD | 2 |
| 2006 | On the Improvement of Statistical Static Timing AnalysisabstractAs the minimum feature sizes of VLSI fabrication processes continue to shrink, the impact of process variations is becoming increasingly significant. This has prompted research into extending traditional static timing analysis so that it can be performed statistically. However, statistical static timing analysis (SSTA) tends to be quite pessimistic. In this paper we present a sensitizable statistical timing analysis (StatSense) technique to overcome the pessimism of SSTA. Our StatSense approach implicitly eliminates false paths, and also uses different delay distributions for different input transitions for any gate. These features enable our StatSense approach to perform less conservative timing analysis than the SSTA approach. Our results show that on average, the worst case (mu + 3sigma) circuit delay reported by StatSense is about 20% lower than that reported by SSTA. Rajesh Garg, Nikhil Jayakumar, Sunil P. Khatri |
ICCD | 3 |
| 2006 | CMOS Comparators for High-Speed and Low-Power ApplicationsabstractIn this paper, we present two designs for CMOS comparators: one which is targeted for high-speed applications and another for low-power applications. Additionally, we present hierarchical pipelined comparators which can be optimized for delay, area, or power consumption by using either design in different stages. Simulation results for our fastest hierarchical 64-bit comparator with a 1.2 V 100 nm process demonstrate a worst-case delay of 440 ps. To enable a fair comparison with previously reported approaches, we also simulated our designs with a 5.0 V AMIS 0.5 mum process as well. For this experiment, the fastest design has a latency of 1.33 ns, which represents a 37% speed improvement over the best previously reported approach to date (which was implemented in a 0.5 mum process). Eric Menendez, Dumezie Maduike, Rajesh Garg, Sunil P. Khatri |
ICCD | 4 |
| 2006 | An Efficient, Scalable Hardware Engine for Boolean SATisfiabilityabstractBoolean Satisfiability (SAT) is a core NP-complete problem in logic synthesis. Several heuristic software and hardware approaches have been proposed to solve this problem. In this paper, we present a hardware solution to the SAT problem. We propose a custom IC to implement our approach, in which the traversal of the implication graph as well conflict clause generation are performed in hardware, in parallel. In our approach, clause literals are stored in specially designed cells. Clauses are implemented in banks, in a manner that allows clauses of variable width to be accommodated in these banks. To maximize the utilization of these banks, we initially partition the SAT problem. Our design is flexible in that it can implement various Boolean Constraint Propagation (BCP) engines on the same die, at the same time, allowing the user to switch BCP engines dynamically. Our solution has significantly larger capacity than existing hardware SAT solvers, and is scalable in the sense that several ICs can be used to simultaneously operate on the same SAT instance, effectively increasing capacity further. Our area and performance figures are derived from layout and SPICE (using extracted parasitics) estimates. Additionally, the approach presented in this paper have been functionally validated in Verilog. Preliminary results demonstrate that our approach can accommodate instances with approximately 63K clauses on a single IC of size 1.5cmx 1.5cm. The approach re suits in over 4 orders of magnitude speed improvement over BCP based software SAT approaches (2-3 orders of magnitude over other hardware SAT approaches). The capacity of our approach is significantly higher than most hardware based approaches. Mandar Waghmode, Kanupriya Gulati, Sunil P. Khatri, Weiping Shi |
ICCD | 3 |
| 2006 | Memory-based crosstalk canceling CODECs for on-chip busesabstractIn recent times, the ratio of the cross-coupling capacitance between adjacent on-chip wires on the same metal layer to the total capacitance of any wire is becoming quite large. As a consequence, signal wires exhibit a significant delay variation and noise immunity problems. This problem is aggravated for long on-chip buses. In this paper, we develop memory-based crosstalk canceling CODECs for on-chip buses. We describe an reduced ordered binary decision diagram (ROBDD) based methodology to accurately compute the bus area overhead of the CODECs. We report the asymptotic overhead for CODECs which cancel three kinds of crosstalk patterns, and demonstrate that the bus size overheads are lower than the corresponding overheads for a memoryless CODEC. This results in a reduced overall area utilization for memory-based crosstalk canceling CODECs, compared to their memoryless counterparts. We also demonstrate that the use of these crosstalk canceling CODECs enables a user to speed up a bus by a factor of over 6/spl times/. Further, by using our techniques, a user may trade off the speed gain against the attendant bus size overhead. Chunjie Duan, Kanupriya Gulati, Sunil P. Khatri |
ISCAS | 3 |
| 2006 | Computing during supply voltage switching in DVS enabled real-time processorsabstractIn recent times, much attention has been devoted to power optimization for real-time systems, while guaranteeing that such systems meet their hard (or soft) scheduling deadlines. To reduce power, different tasks in such systems may be run at different power supply voltages, in order to maximally utilize slack in the schedule. However, prior approaches have ignored the practical aspects of switching the power supply. In a typical IC, the VDD net is highly capacitive, and as a result, its voltage cannot be changed instantaneously. In traditional approaches, the assumption is that this net switches instantaneously, which in effect makes it essential to include the VDD switching time in the worst-case execution time (WCET) of a process (adding pessimism to the WCET value). In our approach, we precisely model the switching of the VDD net, and allow the system to continue computations while VDD is being switched. The effect on the delay of tasks during this transition is precisely modeled. This allows a designer to obtain more realistic estimates of the WCET of a process, reducing the pessimism inherent in real-time system scheduling. Our approach can be implemented as a simple look-up table in a real-time scheduler. Our experimental results show that our model is highly accurate, with an error of <0.2% compared to SPICE simulations. Chunjie Duan, Sunil P. Khatri |
ISCAS | 2 |
| 2006 | Generalized buffering of PTL logic stages using Boolean divisionabstractPass transistor logic (PTL) is a well known approach for implementing digital circuits. In order to handle larger designs, and also to ensure that the total number of series devices in the resulting circuit is bounded, partitioned reduced ordered binary decision diagrams (ROBDDs) can be used to generate the PTL circuit. The output signals of each partitioned block typically needs to be buffered. In this paper, we present a methodology to perform generalized buffering of the outputs of PTL blocks. By performing the Boolean division of each PTL block using different gates in a library, we select the gate that results in the largest reduction in the height of the PTL block. In this manner, these gates serve the function of buffering the outputs of the PTL blocks, while also reducing the height and delay of the PTL block. Over a number of examples, we demonstrate that our approach results in a 26% reduction in circuit delay and number of MUXes required, with a modest improvement in circuit area, compared to a traditional buffered PTL implementation of the circuit Rajesh Garg, Sunil P. Khatri |
ISCAS | 2 |
| 2006 | A probabilistic method to determine the minimum leakage vector for combinational designsabstract"Parking" a circuit in a minimum leakage state during its standby mode of operation is one of the techniques of reducing leakage power consumption in a circuit. However, the problem of finding this minimum leakage state is NP-hard. In this paper, we present a heuristic approach to determine the input vector which minimizes leakage for a combinational design. Our approach utilizes approximate signal probabilities of internal nodes to aid in finding the minimum leakage vector. We use a probabilistic heuristic to select the next gate to be processed, as well as to select the best state of the selected gate. A fast SAT solver is employed to ensure the consistency of the assignments that are made in this process. Experimental results indicate that our method has very low run-times, with excellent accuracy, compared to existing approaches Kanupriya Gulati, Nikhil Jayakumar, Sunil P. Khatri |
ISCAS | 3 |
| 2006 | Efficient don't care computation for hierarchical designsabstractIn this paper, we describe a BDD-based hierarchical don't care computation algorithm. In contrast to traditional don't care computation techniques, our method retains the hierarchy in the design netlist during the don't care computation. Although this may reduce some of the flexibility inherent in the optimization process, it allows our technique to handle large designs. Our method computes don't cares at input and output interfaces of different modules in the hierarchy by an image computation process. In case an exact image cannot be computed, our method computes the largest approximate image. Once the don't cares at the input and output interfaces are computed, the hierarchical instances are optimized separately using a traditional optimization flow. Experimental results demonstrate that our technique can achieve a 36% reduction in literal count for large hierarchical designs, with reasonable runtimes. Our method can complete for several examples in which flattened optimization fails. Kanupriya Gulati, M. Lovell, Sunil P. Khatri |
ISCAS | 3 |
| 2005 | A dynamic voltage scaling algorithm for energy reduction in hard real-time systemsabstractAs the quantity and functional complexity of battery powered portable devices continues to rise, energy efficient design of such devices has become increasingly important. Many real-time scheduling algorithms have been developed recently to reduce energy consumption in hard real-time embedded systems that use dynamic voltage scaling (DVS) capable processors. This paper explores an algorithm that seeks to reduce energy consumption by considering tasks in tandem, with the intuition that what may be a good frequency for one task, may be much worse for another. In particular, our algorithm considers pairs of tasks, and optimizes them simultaneously so that their total energy consumption is minimized while all deadlines are met. Experimental results demonstrate that our method is able to effectively improve on the results of look-ahead EDF, one of the best energy-aware schedulers, especially for task sets with moderate utilization, and "harmonious" task periodicity. Van R. Culver, Sunil P. Khatri |
ASP-DAC | 2 |
| 2005 | A self-adjusting scheme to determine the optimum RBB by monitoring leakage currentsabstractReverse body biasing (RBB) is often used to reduce the leakage power of a device. However, recent research has shown that if this applied RBB is too high, the leakage power can actually increase due to the contribution of Band-to-Band Tunneling (BTBT) currents. Hence, there exists an optimal RBB value at which the leakage is minimum. This optimum point can vary with temperature and process variations. In this paper we show that it is desirable to operate at the optimal RBB point which minimizes total leakage. We present a scheme that monitors the total leakage current (the sum of the sub-threshold, BTBT and gate leakage) of an IC with a representative leaking device and, using this monitored value, automatically finds the optimum RBB value across temperature and process corners, using a self-adjusting circuit. Our approach has a modest placed-and-routed area utilization, and a low power consumption. Nikhil Jayakumar, Sandeep Dhar, Sunil P. Khatri |
DAC | 3 |
| 2005 | A variation tolerant subthreshold design approachabstractDue to their extreme low power consumption, sub-threshold design approaches are appealing for a widening class of applications which demand low power consumption and can tolerate larger circuit delays. However, sub-threshold circuits are extremely sensitive to variations in supply, temperature and processing factors. In this paper, we present a sub-threshold design methodology which dynamically self-adjusts for inter and intra-die process, supply voltage and temperature (PVT) variations. This adjustment is achieved by performing bulk voltage adjustments in a closed-loop fashion, using a charge pump and a phase-detector. Nikhil Jayakumar, Sunil P. Khatri |
DAC | 2 |
| 2005 | Encoding-Based Minimization of Inductive Cross-Talk for Off-Chip Data TransmissionabstractInductive cross-talk within IC packaging is becoming a significant bottleneck in high-speed inter-chip communication. The parasitic inductance within IC packaging causes bounce on the power supply pins in addition to glitches and rise-time degradation on the signal pins. Until recently, the parasitic inductance problem was addressed by aggressive package design. We present a technique to encode the off-chip data transmission to limit bounce on the supplies and reduce inductive signal coupling due to transitions on neighboring signal lines. Both these performance limiting factors are modeled in a common mathematical framework. Our experimental results show that the proposed encoding based techniques result in reduced supply bounce and signal degradation due to inductive cross-talk, closely matching the theoretical predictions. We demonstrate that the overall bandwidth of a bus actually increases by 85% using our technique, even after accounting for the encoding overhead. The asymptotic bus size overhead is between 30% and 50%, depending on how stringent the user-specified inductive cross-talk parameters are. Brock J. LaMeres, Sunil P. Khatri |
DATE | 2 |
| 2005 | A Boolean satisfiability based solution to the routing and wavelength assignment problem in optical telecommunication networksabstractWavelength division multiplexing (WDM) effectively multiplies the bandwidth of an optic fiber by transmitting data over several different wavelengths on the same fiber. WDM is widely used to handle the ever-increasing demand for bandwidth in fiber optic telecommunication networks. Routing and wavelength assignment (RWA) is a critical problem to be addressed in WDM optical telecommunication networks. The goal of RWA is to maximize throughput by optimally and simultaneously assigning routes and wavelengths for a given pattern of routing or connection requests. In this paper, we present a novel technique to solve the static RWA problem using Boolean satisfiability (SAT). After formulating the RWA problem as a SAT instance, we utilize a very efficient SAT solver to find a solution. We report results for networks with and without wavelength translation capabilities in the nodes. In both cases we obtain an assignment in significantly less than one second (which is 3-4 orders of magnitude faster than existing approaches) for a set of benchmark problems. Our technique can handle arbitrary network topologies, and, due to its efficiency, can be extended to handle dynamic RWA instances, in which the network topology, link capacities and connection requests are time-varying. John Valavi, Nikhil Saluja, Sunil P. Khatri |
ICC | 3 |
| 2005 | Practical techniques to reduce skew and its variations in buffered clock networksabstractClock skew is becoming increasingly difficult to control due to variations. Link based non-tree clock distribution is a cost-effective technique for reducing clock skew variations. However, previous works based on this technique were limited to unbuffered clock networks and neglected spatial correlations in the experimental validation. In this work, we overcome these shortcomings and make the link based non-tree approach feasible for realistic designs. The short circuit risk and multi-driver delay issues in buffered non-tree clock networks are investigated. Our approach is validated with SPICE based Monte Carlo simulations, considering spatial correlations among variations. The experimental results show that our approach can reduce the maximal skew by 47%, improve the skew yield from 15% to 73% on average with a decrease on the total wire and buffer capacitance. Ganesh Venkataraman, Nikhil Jayakumar, Jiang Hu 0001, Peng Li 0001, Sunil P. Khatri, Anand Rajaram, Patrick McGuinness, Charles J. Alpert |
ICCAD | 5 |
| 2005 | X-Routing using Two Manhattan Route InstancesabstractIn deep sub-micron (DSM) technologies, wire delays comprise a dominant fraction of the total delay of a design. As a consequence, routing techniques which reduce the total wire length of a design are highly relevant to such technologies. One such approach which holds promise is that of non-Manhattan routing (or X routing). In this paper, we describe a technique to perform non-Manhattan routing by combining the results of two related Manhattan routing instances. The first is a regular, unrotated routing instance. The second routing instance is derived from the first by rotating the coordinate system by 45/spl deg/. Both instances are routed on the same pair of metal layers. By selectively combining the results of the two instances, we obtain a final routing result that contains non-Manhattan wire segments. Our approach utilizes a powerful Floyd-Warshall based engine to combine the results of the two instances. We demonstrate that our router produces highly efficient results, reducing the total wire length by an average of about 20% (31%) over the unrotated (rotated) results, with a via-count decrease of between 4% (43%). Seraj Ahmad, Nikhil Jayakumar, Vijay Balasubramanian, Edward Hursey, Sunil P. Khatri, Rabi N. Mahapatra |
ICCD | 5 |
| 2005 | Minimum Energy Near-threshold Network of PLA based DesignabstractIn recent times, there has been a significant growth in applications for battery powered portable electronics, as well as low power sensor networks. While sub-threshold circuit design approaches can reduce the power consumption significantly, a design operating at sub-threshold voltages is not necessarily optimal in terms of energy consumption. In this paper, we describe a technique to find the energy optimum VDD value for a design, and show that for minimum energy consumption, the circuit should be operated at VDD values which are above the NMOS threshold voltage value. We study this problem in the context of designing a circuit using a network of dynamic NOR-NOR PLAs. Nikhil Jayakumar, Sunil P. Khatri |
ICCD | 2 |
| 2005 | Broadband Impedance Matching for Inductive Interconnect in VLSI PackagesabstractNoise induced by impedance discontinuities from VLSI packaging is one of the leading challenges facing system level designers in the next decade. The performance of IC cores far exceeds that of current packaging technology. The risetimes of IC signals require that the interconnect of the package be treated as transmission lines. As a result, impedance discontinuities in the package cause reflections which may result in intermittent switching of digital signals and edge time degradation, both of which limit system performance. The major cause of the impedance discontinuity in the package is the high inductance of the wire bond interconnects. To compensate for this problem, capacitance can be placed near the wire bond to reduce its effective impedance over a given frequency range. This paper presents the application of this impedance matching technique for use in broadband digital signals that are prevalent in modern VLSI designs. Both static and dynamic compensation approaches are presented. The static compensator places pre-defined capacitances on the package and on the IC to surround the wire bond inductance. The dynamic compensator places a switchable capacitance on the IC that can be programmed to a desired value, thereby enabling the designer to overcome design and manufacturing variations in the Mire bond. Both techniques presented are shown to bound the reflections of the wire bond to less than 5% (down from 20% for an uncompensated structure) for wire bonds up to 5mm in length, and for frequencies up to 3GHz. In addition, both circuits utilizes less area than a typical wire bond pad, making them ideal for placement directly beneath the wire bond pads. Brock J. LaMeres, Sunil P. Khatri |
ICCD | 2 |
| 2005 | An algebraic decision diagram (ADD) based technique to find leakage histograms of combinational designsabstractIn this paper, we present an Algebraic Decision Diagram (ADD) based approach to determine and implicitly represent the leakage value for all input vectors of a combinational circuit. In its exact form, our technique can compute the leakage value of each input vector. To broaden the applicability of our technique, we present an approximate version of our algorithm as well. The approximation is done by limiting the total number of discriminant nodes in any ADD. Previous sleep vector computation techniques can find either the maximum or minimum sleep vector. Our technique computes the leakages for all vectors, storing them implicitly in an ADD structure. We experimentally demonstrate that these approximate techniques produce results which have reasonable errors. We also show that limiting the number of discriminants to a value between 12 and 16 is practical, allowing for good accuracy and lowered memory utilization Kanupriya Gulati, Nikhil Jayakumar, Sunil P. Khatri |
ISLPED | 3 |
| 2005 | Efficient SAT-based combinational ATPG using multi-level don't-caresabstractIn this paper, we present two combinational ATPG algorithms for combinational designs. These algorithms utilize the multi-level don't cares that are computed for the design during technology independent logic optimization. They are based on Boolean satisfiability (SAT), and utilize the single stuck-at fault model. Both algorithms make use of the compatible observability don't cares (CODCs) associated with nodes of the circuit, to speed up the ATPG process. For large circuits, both algorithms make use of approximate CODCs (ACODCs), which we can compute efficiently. Our first technique speeds up fault propagation by modifying the active clauses in the transitive fanout (TFO) of the fault site. In our second technique, we define new j-active variables for specific nodes in the transitive fanin (TFI) of the fault site. Using these j-active variables we write additional clauses to speed up fault justification. Experimental results demonstrate that the combination of these techniques (when using CODCs) results in an average reduction of 45% in ATPG run-times. When ACODCs are used, a speed-up of about 30% is obtained in the ATPG run-times for large designs. We compared our method against a commercial structural ATPG tool as well. Our method was slower for small designs, but for large designs, we obtained a 31% average speedup over the commercial tool Nikhil Saluja, Sunil P. Khatri |
ITC | 2 |
| 2004 | A robust algorithm for approximate compatible observability don't care (CODC) computationabstractCompatible Observability Don't Cares (CODCs) are a powerful means to express the flexibility present at a node in a multi-level logic network. Despite their elegance, the applicability of CODCs has been hampered by their computational complexity. The CODC computation for a network involves several image computations, which require the construction of global BDDs of the circuit nodes. The size of BDDs of circuit nodes is unpredictable, and as a result, the CODC computation is not robust. In practice, CODCs cannot be computed for large circuits due to this limitation.In this paper, we present an algorithm to compute approximate CODCs (ACODCs). This algorithm allows us to compute compatible don't cares for significantly larger designs. Our ACODC algorithm is scalable in the sense that the user may trade off time and memory against the accuracy of the ACODCs computed. The ACODC is computed by considering a subnetwork rooted at the node of interest, up to a certain topological depth, and performing its don't care computation. We prove that the ACODC is an approximation of its CODC.We have proved the soundness of the approach, and performed extensive experiments to explore the trade-off between memory utilization, speed and accuracy. We show that even for small topological depths, the ACODC computation gives very good results. Our experiments demonstrate that our algorithm can compute ACODCs for circuits whose CODC computation has not been demonstrated to date. Also, for a set of benchmark circuits whose CODC computation yields an average 28% reduction in literals after optimization, our ACODC computation yields an average 22% literal reduction. Our algorithm has runtimes which are about 25x and memory utilization which is 33x better that of the CODC computation of SIS. Nikhil Saluja, Sunil P. Khatri |
DAC | 2 |
| 2004 | Exploiting Crosstalk to Speed up On-Chip BuseabstractIn modern VLSI processes, the cross-coupling capacitance between adjacent neighboring wires on the same metal layer is a very large fraction of the total wire capacitance. This leads to problems of delay variation due to crosstalk and reduced noise immunity, arguably one of the biggest obstacles in the design ICs in recent times. This problem is particularly severe in long on-chip buses, since bus signals are routed at minimum pitch for long distances. In this work, we propose to solve this problem by the use of crosstalk canceling CODECs. We only utilize memoryless CODECs, to reduce the logical complexity and enhance the robustness of our techniques. Bus data patterns can be classified (as 4/spl middot/C, 3/spl middot/C, 2/spl middot/C, 1/spl middot/C or 0/spl middot/C patterns) based on the maximum amount of crosstalk that they can exhibit. Crosstalk avoidance CODECs which eliminate 4/spl middot/C and 3/spl middot/C patterns have been reported. In this paper, we describe crosstalk avoidance techniques which eliminate 2/spl middot/C and 1/spl middot/C patterns. We describe an analytical methodology to accurately characterize the bus area overhead 2/spl middot/C pattern CODECs. Using these results, we characterize the area overhead versus crosstalk immunity achieved. A similar exercise is performed for 1/spl middot/C patterns. Our experimental results show that by using 2/spl middot/C crosstalk canceling techniques, buses can be sped up by up to a factor of 6 with an area overhead of about 200%, and that 1/spl middot/C techniques are not very robust. Chunjie Duan, Sunil P. Khatri |
DATE | 2 |
| 2004 | High-throughput VLSI implementations of iterative decoders and related code construction problemsabstractIn this paper, an efficient, fully-parallel network of programmable logic array (NPLA)-based realization of iterative decoders for structured LDPC codes is presented. The LDPC codes are developed in tandem with the underlying VLSI implementation technique, without compromising chip design constraints. The codes are based on a novel modification of array codes. This design methodology results in reduced routing congestion, a major problem in prior approaches. The operating power, delay and chip-size of the circuits are estimated, indicating that this implementation significantly outperforms presently used standard-cell based architectures. The described LDPC design method can accommodate widely different requirements, such as those arising from recording and wireless channel applications. Vijay Nagarajan, Nikhil Jayakumar, Sunil P. Khatri, Olgica Milenkovic |
GLOBECOM | 3 |
| 2004 | A metal and via maskset programmable VLSI design methodology using PLAsabstractIn recent times there has been a substantial increase in the cost and complexity of fabricating a VLSI chip. The lithography masks themselves can cost between /spl epsi/ and /spl ges/. It is conjectured that due to these increasing costs, the number of ASIC starts in the last few years has declined. We address this problem by using an array of dynamic PLAs which require only metal and via mask customization in order to implement a new design. This would allow several similar-sized designs to share the same base set of masks (right up to the metal layers) and only have different metal and via masks. We have implemented our methodology for both combinational and sequential designs, and demonstrate that our approach strikes a reasonable compromise between ASIC and field programmable design methodologies in terms of placed-and-routed area and delay. Our method has a 2.89/spl times/ (3.58/spl times/) delay overhead and a 4.96/spl times/ (3.44/spl times/) area overhead compared to standard cells for combinational (sequential) designs. Nikhil Jayakumar, Sunil P. Khatri |
ICCAD | 2 |
| 2004 | A novel clock distribution and dynamic de-skewing methodologyabstractIn present day VLSI ICs, intra-die processing variations are becoming harder to control, resulting in a large skew in the clock signals at the end of the clock distribution network. We describe a buffered H-tree technique to distribute the clock signal and to de-skew a clock network. The clock shielding wires (which are connected to GND in normal operation) are, in de-skewing mode, used to selectively return the clock signal for de-skewing, and for serial communication with the clock distribution sites for skew adjustment. Our forward and return clock networks are buffered, with identically sized and co-located wires and buffers. This results in both these networks exhibiting identical delay characteristics in the presence of intra-die process variations. Unlike existing approaches, our method utilizes a single phase detection circuit, and can achieve a very low maximum chip-level clock skew. This skew value is not dependent on the resolution of the phase detector. Further, our technique can be applied dynamically, either at boot time or periodically during the operation of the IC, as necessary. Additionally, our buffered H-tree enables us to implement efficient clock gating by allowing the user to turn off clocks in the distribution network itself, thus disabling entire sections of the clock network. We demonstrate the utility of our technique on a 6-level H-tree clock distribution network. In a clock distribution network which is initially skewed by up to 300ps, our technique can de-skew signals to within 4ps of each other. We show that the total wiring area of our clock distribution and de-skewing methodology is about 35% higher than a traditional H-tree (which does not have a deskewing functionality), while the active logic area overhead is about 25%. The power consumption of our network is 5% lower than that of a traditional H-tree network with no de-skewing functionality. Arjun Kapoor, Nikhil Jayakumar, Sunil P. Khatri |
ICCAD | 3 |
| 2004 | SPFD-based wire removal in standard-cell and network-of-PLA circuitsabstractWire removal is a technique by which the total number of wires between individual circuit nodes is reduced, either by removing wires or replacing them with other new wires. The wire removal techniques we describe in this paper are based on both binary and multivalued sets of pairs of functions to be distinguished (SPFDs). Recently, it was shown that a design style based on a multilevel network of approximately equal-sized programmable logic arrays (PLAs) results in a dense, fast, and crosstalk-resistant layout. This paper describes the application of SPFD-based wire removal techniques for circuit implementations utilizing networks of PLAs as well as standard-cells. In our first set of wire removal experiments (which utilize binary SPFD-based wire removal), we demonstrate that the benefit of SPFD-based wire removal is insignificant when the circuit is mapped using standard cells. We demonstrate that this technique is very effective in the context of a network of PLAs. In the next set of wire removal experiments, we focus only on circuits implemented using a network of PLAs. Three separate wire removal experiments are performed. Wire removal is invoked before clustering the original netlist into a network of PLAs, or after clustering, or both before and after clustering. For wire removal before clustering, binary SPFD-based wire removal is used. For wire removal after clustering, multivalued SPFD-based wire removal is used since the multioutput PLAs can be viewed as multivalued single output nodes. We demonstrate that these techniques are effective. The most effective approach is to perform wire removal both before and after clustering. Using these techniques, we obtain a reduction in placed and routed circuit area of about 11%. This reduction is significantly higher (about 20%) for the larger circuits we used in our experiments. Sunil P. Khatri, Subarnarekha Sinha, Robert K. Brayton, Alberto L. Sangiovanni-Vincentelli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2003 | A fast ternary CAM design for IP networking applicationsabstractIn this paper we describe a VLSI implementation and complete circuit design of a fast ternary CAM (TCAM). TCAMs are commonly used to perform routing lookups in the backbone of IP networks and small gateways. Our TCAM is designed to have a greater capacity and speed than any commercial offering at this time. In contrast with existing TCAM approaches, our TCAM allows complete flexibility in the location where any new entry is inserted. This is achieved by a novel longest prefix match (LPM) determination circuit, whose delay increases logarithmically with the number of bits to be looked up. We have implemented our TCAM with 512 bits of prefix entry with 512 bits of destination information, allowing it to implement large address lookups as well as quality of service mechanisms. This would make our TCAM design particularly suitable for IPv6 routing lookup applications. The speed improvement of our TCAM over currently available TCAMs results from various carefully selected VLSI architectural and implementation choices. The TCAM size is 21 Mb and is broken up into a regular grid of 13x13 smaller TCAM blocks for improved speed characteristics. Routing lookup operations use a heavily pipelined approach for maximum throughput, while ensuring a lookup latency of 3 clock cycles. Individual match lines in these blocks are split into 4 sections to reduce RC delay in the lookup process. Our LPM determination circuit is implemented using an efficient wired-NOR circuit for further reduced delay. Sense amplifiers are utilized in the LPM and SRAM sections of the TCAM and are located in the center of each TCAM subblock in order to improve lookup speed. We have implemented and validated our design using state-of-the-art circuit analysis and design tools. We have also generated mask layouts of the entire TCAM design using current layout tools. The complete TCAM circuit design is approximately 18mm on a side, with a total capacity of 21Mb. Our TCAM has an ability to perform routing lookups at a line rate of 76.8Gb/s which is twice as fast as the fastest commercially available TCAM today. Bruce Gamache, Zachary Pfeffer, Sunil P. Khatri |
ICCCN | 3 |
| 2003 | An ASIC design methodology with predictably low leakage, using leakage-immune standard cellsabstractIn this paper we introduce a low-leakage standard cell based ASIC design methodology which is based on the use of modified standard cells. These cells are designed to consume extremely low and predictable leakage currents in standby mode. For each cell in a standard cell library, we design two low-leakage variants of the cell. If the inputs of a cell during the standby mode of operation are such that the output has a high value, we minimize the leakage in the pull-down network, and vice versa. While technology mapping a circuit, we determine the particular variant to utilize in each instance, so as to minimize leakage of the final mapped design.We have designed and laid out our modified standard cells, and have performed experiments to compare placed-and-routed area, leakage and delays of our method against MTCMOS and a straightforward ASIC flow. Each design style we compare utilizes the same base standard cell library.Our results show that designs obtained using our methodology have better speed and area characteristics than designs implemented in MTCMOS. The exact leakage current obtained for MTCMOS is highly unpredictable, while our method exhibits leakage currents which are precisely estimable. The leakage current for HL designs can be dramatically lower than the worst-case leakage of MTCMOS based designs, and two orders of magnitude compared to traditional standard cells. Also, a design implemented in MTCMOS would require the use of separate power and ground supplies for latches and combinational logic, while our methodology does away with such a requirement. Nikhil Jayakumar, Sunil P. Khatri |
ISLPED | 2 |
| 2002 | An efficient and regular routing methodology for datapath designsusing net regularity extractionabstractWe present a new detailed routing methodology specifically designed for datapath layouts. In typical state-of-the-art microprocessor designs, datapaths comprise about 70% of the logic (excluding caches). However, most logic and layout synthesis research has targeted random-logic portions of the design. In general, techniques for random-logic placement and routing do not produce good results for datapath layouts. Although research on datapath placement and global routing has been reported, very little research has been reported in the area of detailed routing for datapaths. Our datapath routing methodology exploits the unique feature of datapaths, namely, their regularity. Datapaths typically comprise regular structures (or bit slices), which are replicated. The interconnections between these replicated bit slices are also typically very regular. Our datapath routing methodology utilizes new techniques to extract interconnect regularity among bit slices. We define a net cluster, which is collection of similarly structured nets present across different bit slices. We introduce two clustering schemes (footprint-driven clustering and instance-driven clustering) to extract such net clusters. Using these net clusters, we select one representative bit slice to perform a strap-based routing (which optimally finds the shortest path between two points if that path is available) on a member net of each net cluster. Then for each such net, we propagate its route to all other nets in its net cluster. Our algorithm is unique in that it performs the detailed routing on a single bit slice and infers the routing for all bit slices using the notion of net clusters. Since we only route a small fraction of nets present in the design, significant speedup is obtained. We demonstrate at least six times speedup for industrial 32- and 64-bit datapath designs. The regularity of the routes across the bit slices results in more predictable timing characteristics for the resulting layout. Sabyasachi Das, Sunil P. Khatri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2001 | Addressing the Timing Closure Problem by Integrating Logic Optimization and PlacementabstractTiming closure problems occur when timing estimates computed during logic synthesis do not match with timing estimates computed from the layout of the circuit. In such a situation, logic synthesis and layout synthesis are iterated until the estimates match. The number of such iterations is becoming larger as technology scales. Timing closure problems occur mainly due to the difficulty in accurately predicting interconnect delay during logic synthesis. In this paper, we present an algorithm that integrates logic synthesis and global placement to address the timing closure problem. We introduce technology independent algorithms as well as technology dependent algorithms. Our technology independent algorithms are based on the notion of "wire-planning". All these algorithms interleave their logic operations with local and incremental/full global placement, in order to maintain a consistent placement while the algorithm is run. We show that by integrating logic synthesis and placement, we avoid the need to predict interconnect delay during logic synthesis. We demonstrate that our scheme significantly enhances the predictability of wire delays, thereby solving the timing closure problem. This is the main result of our paper. Our results also show that our algorithms result in a significant reduction in total circuit delay. In addition, our technology independent algorithms result in a significant circuit area reduction. Wilsin Gosti, Sunil P. Khatri, Alberto L. Sangiovanni-Vincentelli |
ICCAD | 2 |
| 2001 | A regularity-driven fast gridless detailed router for high frequency datapath designsabstractWe present a new detailed routing methodology specifically designed for datapath layouts. In typical state-of-the-art microprocessor designs, datapaths comprise about 70\% of the logic (excluding caches). Although research on datapath placement and global routing has been reported, very little research has been reported in the area of detailed routing for datapaths. Sabyasachi Das, Sunil P. Khatri |
ISPD | 2 |
| 2000 | Cross-Talk Immune VLSI Design Using a Network of PLAs Embedded in a Regular Layout FabricabstractWe present a VLSI design methodology to address the cross-talk problem, which is becoming increasingly important in Deep Sub-Micron (DSM) IC design. In our approach, we implement the logic netlist in the form of a network of medium sized PLAs. We utilize two regular layout "fabrics" in our methodology, one for areas where PLA logic is implemented, and another for routing regions between such logic blocks. We show that a single PLA implemented in the first fabric style is not only cross-talk immune, but also about 2/spl times/ smaller and faster than a traditional standard cell based implementation of the same logic. The second fabric, utilized in the routing region between individual PLAs, is also highly cross-talk immune. Additionally, in this fabric, power and ground signals are essentially "pre-routed" all over the die. Our synthesis flow involves decomposing the design into a network of PLAs, each of which has a bounded width and height. The number of inputs and outputs of each PLA are flexible as long as the resulting PLA width is bounded. We perform folding of PLAs to achieve better logic density. Routing is performed using 2,3,4,5 and 6 routing layers. State-of-the-art commercial routing tools are utilized for the experiments involving the use of 3,4,5 and 6 routing layers. We have implemented the entire design flow using these ideas. Our scheme results in a reduction in the cross-talk between signal wires of between one and two orders of magnitude. As a result, for a 0.1 /spl mu/m process, the delay variation due to cross-talk dramatically drops from 2.47:1 to 1.02:1. Additionally, our methodology results in circuits that are extremely fast and dense, with a timing improvement of about 15% and an overall area penalty of about 3% compared to standard cells. The regular arrangement of metal conductors in our scheme results in low and highly predictable inductive and capacitive parasitics, resulting in highly predictable designs. The crosstalk immunity, high speed, low area overhead and high predictability of our methodology indicate that it is a strong candidate as the preferred design methodology in the DSM era. Sunil P. Khatri, Robert K. Brayton, Alberto L. Sangiovanni-Vincentelli |
ICCAD | 1 |
| 2000 | Binary and Multi-Valued SPFD-Based Wire Removal in PLA NetworksabstractThis paper describes the application of binary and multivalued SPFD-based wire removal techniques for circuit implementations utilizing networks of PLAs. It has been shown that a design style based on a multi-level network of approximately equal-sized PLAs results in a dense, fast, and crosstalk-resistant layout. Wire removal is a technique where the total number of wires between individual circuit nodes is reduced, either by removing wires, or replacing them with other existing wires. Three separate wire removal experiments are performed. Either wire removal is invoked before clustering the original netlist into a network of PLAs, or after clustering, or both before and after clustering. For wire removal before clustering, binary SPFD-based wire removal is used. For wire removal after clustering, multi-valued SPFD-based wire removal is used since the multi-output PLAs can be viewed as multi-valued single output nodes. We demonstrate that these techniques are effective. The most effective approach is to perform wire removal both before and after clustering. Using these techniques, we obtain a reduction in placed and routed circuit area of about 11%. This reduction is significantly higher (about 20%) for the larger circuits we used in our experiments. Subarnarekha Sinha, Sunil P. Khatri, Robert K. Brayton, Alberto L. Sangiovanni-Vincentelli |
ICCD | 2 |
| 1999 | A Novel VLSI Layout Fabric for Deep Sub-Micron ApplicationsabstractWe propose a new VLSI layout methodology which addresses the main problems faced in Deep Sub-Micron (DSM) integrated circuit design. Our layout “fabric ” scheme eliminates the conventional no-tion of power and ground routing on the integrated circuit die. In-stead, power and ground are essentially “pre-routed ” all over the die. By a clever arrangement of power/ground and signal pins, we almost completely eliminate the capacitive effects between signal wires. Ad-ditionally, we get a power and ground distribution network with a very low resistance at any point on the die. Another advantage of our scheme is that the arrangement of conductors ensures that on-chip inductances are uniformly negligible. Finally, characterization of the circuit delays, capacitances and resistances becomes extremely simple in our scheme, and needs to be done only once for a design. We show how the uniform parasitics of our fabric give rise to a reliable and predictable design. We have implemented our scheme using public domain layout software. Preliminary results show that it holds much promise as the layout methodology of choice in DSM integrated circuit design. 1 Sunil P. Khatri, Amit Mehrotra, Robert K. Brayton, Ralph H. J. M. Otten, Alberto L. Sangiovanni-Vincentelli |
DAC | 1 |
| 1996 | VIS: A System for Verification and Synthesis
Robert K. Brayton, Gary D. Hachtel, Alberto L. Sangiovanni-Vincentelli, Fabio Somenzi, Adnan Aziz, Szu-Tsung Cheng, Stephen A. Edwards, Sunil P. Khatri, Yuji Kukimoto, Abelardo Pardo, Shaz Qadeer, Rajeev Ranjan 0001, Shaker Sarwary, Thomas R. Shiple, Gitanjali Swamy, Tiziano Villa |
CAV | 8 |
| 1996 | Engineering Change in a Non-Deterministic FSM Settingabstractpersonal or class-room use is granted without fee provided that copies are not made or distributed for profit or commercial advantage, the copyright notice, the title of the publication and its date appear, and notice is given that copying is Sunil P. Khatri, Amit Narayan, Sriram C. Krishnan, Kenneth L. McMillan, Robert K. Brayton, Alberto L. Sangiovanni-Vincentelli |
DAC | 1 |
| 1996 | VIS
Robert K. Brayton, Gary D. Hachtel, Alberto L. Sangiovanni-Vincentelli, Fabio Somenzi, Adnan Aziz, Szu-Tsung Cheng, Stephen A. Edwards, Sunil P. Khatri, Yuji Kukimoto, Abelardo Pardo, Shaz Qadeer, Rajeev Ranjan 0001, Shaker Sarwary, Thomas R. Shiple, Gitanjali Swamy, Tiziano Villa |
FMCAD | 8 |
| 1996 | Decomposition Techniques for Efficient ROBDD Construction
Jawahar Jain, Amit Narayan, C. Coelho 0001, Sunil P. Khatri, Alberto L. Sangiovanni-Vincentelli, Robert K. Brayton |
FMCAD | 4 |