EDBT 2026 Demo / reviewers in the wild / expert
Srinivas Katkoori
dblp:01/4012
· DBLP profile ↗
38ranked-venue papers
5as first author
4since 2021 · last 2023
0000-0002-7589-5836ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 3Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | On Feasibility of Decision Trees for Edge Intelligence in Highly Constrained Internet-of-Things (IoT)abstractInternet-of-Things (IoT) edge devices have limited resources in terms of area and power. Machine Learning based intelligent filtering can be effective in reducing the data footprint. In this work, we report a feasibility study of using decision trees (DTs) on the edge. The main contribution of this work is to demonstrate that decision trees are equally effective compared to popular neural networks (multi-layer perceptrons). We trained four datasets (Iris, Heart Disease, Breast Cancer, and Credit Card) with supervised decision tree-based learning with accuracy comparable to that of MLPs. We synthesized the DTs to gate-level implementation with the Synopsys Design Compiler in 32 nm CMOS technology node. Compared to MLP implementations, DTs can save approximately 97-98% in both area and power. Raaga Sai Somesula, Rajeev Joshi, Srinivas Katkoori |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | Early Design Space Exploration Framework for Memristive Crossbar ArraysabstractFor memristive crossbar arrays, currently, no high-level design validation and early space exploration tools exist in the literature. Such tools are essential to quickly verify the design functionality as well as compare design alternatives in terms of power and performance. In this work, we propose a VHDL-based framework that enables us to quickly perform behavioral simulation as well as estimate dynamic energy consumption and speed of any large memristive crossbar array. We propose a high-level (VHDL) model of a memristor based on which crossbar architectures can be modeled. The individual memristor model is embedded with power and delay numbers obtained from a detailed memristor model. We demonstrate the framework for MAGIC-style memristive crossbars. We validate the framework against detailed Verilog-A based model on fifteen combinational benchmarks. For the single row model, we obtained 153x simulation speedup over HSPICE, average estimation errors of 6.64% and 0% for dynamic energy consumption and cycle-time, respectively. For the transpose model, we obtained average estimation errors of 5.51% and 10.90% for dynamic energy consumption and cycle-time, respectively. We also extend our framework to support another prominent logic style and validate through a case study. The proposed framework can be easily extended to other emerging technologies. Md. Adnan Zaman, Rajeev Joshi, Srinivas Katkoori |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2021 | Partial evaluation based triple modular redundancy for single event upset mitigation
Srinivas Katkoori, Sheikh Ariful Islam, Sujana Kakarla |
Integr. | 1 |
| 2021 | High-Level Synthesis of Key-Obfuscated RTL IP with Design Lockout and CamouflagingabstractWe propose three orthogonal techniques to secure Register-Transfer-Level (RTL) Intellectual Property (IP). In the first technique, the key-based RTL obfuscation scheme is proposed at an early design phase during High-Level Synthesis (HLS). Given a control-dataflow graph, we identify operations on non-critical paths and leverage synthesis information during and after HLS to insert obfuscation logic. In the second approach, we propose a robust design lockout mechanism for a key-obfuscated RTL IP when an incorrect key is applied more than the allowed number of attempts. We embed comparators on obfuscation logic output to check if the applied key is correct or not and a finite-state machine checker to enforce design lockout. Once locked out, only an authorized user (designer) can unlock the locked IP. In the third technique, we design four variants of the obfuscating module to camouflage the RTL design. We analyze the security properties of obfuscation, design lockout, and camouflaging. We demonstrate the feasibility on four datapath-intensive IPs and one crypto core for 32-, 64-, and 128-bit key lengths under three design corners (best, typical, and worst) with reasonable area, power, and delay overheads on both ASIC and FPGA platforms. Sheikh Ariful Islam, Love Kumar Sah, Srinivas Katkoori |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2020 | Basic Block Encoding Based Run-time CFI Check for Embedded SoftwareabstractModern control flow attacks circumvent existing defense mechanisms to transfer the program control to attacker chosen malicious code in the program, leaving application vulnerable to attack. Advanced attacks such as Return-Oriented Programming (ROP) attack and its variants, transfer program execution to gadgets (code-snippet that ends with return instruction). The code space to generate gadgets is large and attacks using these gadgets are Turing-complete. One big challenge to harden the program against ROP attack is to confine gadget selection to a limited locations, thus leaving the attacker to search entire code space according to payload criteria. In this paper, we present a novel approach to label the nodes of the Control-Flow Graph (CFG) of a program such that labels of the nodes on a valid control flow edge satisfy a Hamming distance property. The newly encoded CFG enables detection of illegal control flow transitions during the runtime in the processor pipeline. Experimentally, we have demonstrated that the proposed Control Flow Integrity (CFI) implementation is effective against control-flow hijacking and the technique can reduce the search space of the ROP gadgets upto 99.28%. We have also validated our technique on seven applications from MiBench and the proposed labeling mechanism incurs no instruction count overhead while, on average, it increases instruction width to a maximum of 12.13%. Love Kumar Sah, Srivarsha Polnati, Sheikh Ariful Islam, Srinivas Katkoori |
VLSI-SOC | 4 |
| 2020 | Design, Analysis and Application of Embedded Resistive RAM Based Strong Arbiter PUFabstractResistive Random Access Memory (RRAM) based Physical Unclonable Function (PUF) designs exploit either the probabilistic switching or the resistance variability during forming, SET and RESET processes of RRAM. Memory PUFs using RRAM are typically weak PUFs due to fewer number of challenge response pairs. We propose a strong arbiter PUF based on 1T-1R bit cell which is designed from conventional RRAM memory array with minimally invasive changes. Conventional voltage sense amplifier is repurposed to act like an arbiter and generate the response. Similarly, address and data lines are repurposed to act as challenge and response bits respectively. The PUF is simulated using 65 nm predictive technology models for CMOS and Verilog-A model for a hafnium oxide based RRAM. The proposed PUF architecture is evaluated for uniqueness, uniformity and reliability for various number of stages. It demonstrates mean intra-die Hamming Distance (HD) of 0.135 percent and inter-die HD of 51.4 percent, and passes the NIST tests. We study the vulnerability of proposed PUF to machine learning attacks. We also present an application of proposed PUF for data attestation in the internet of things. Proposed PUF-based data attestation consumes 9.88pJ of total energy per data block of 64-bits and offers a speed of 120.7 kbps. Rekha Govindaraj, Swaroop Ghosh, Srinivas Katkoori |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2020 | Interval Arithmetic and Self-Similarity Based RTL Input Vector Control for Datapath Leakage MinimizationabstractWith technology scaling, subthreshold leakage has dominated the overall power consumption in a design. Input vector control is an effective technique to minimize subthreshold leakage. Low leakage input vector determination is not often possible due to large design space and simulation time. Similarly, applying an appropriate minimum leakage vector (MLV) to each Register Transfer Level (RTL) module instance in a design often results in a low leakage state with significant area overhead. In this work, we propose a top-down and bottom-up approach for propagating the input vector interval to identify low leakage input vector at primary inputs of an RTL datapath. For each module, via Monte Carlo simulation, we identify a set of MLV intervals such that maximum leakage is within (say) 10% of the lowest leakage points. As the module bit width increases, exhaustive simulation to find the low leakage vector is not feasible. Further, we need to uniformly search the entire input space to obtain as many low leakage intervals as possible. Based on empirical observations, we observe self-similarity in the subthreshold leakage distribution of adder/multiplier modules with highly regular bit-slice architectures when input space is partitioned into smaller cells. This property enables the uniform search of low leakage vectors in the entire input space where the time taken for characterization increases linearly with the module size. We further process the reduced interval set with simulated annealing to arrive at the best low-leakage vector at the primary inputs. We also propose to reduce area overhead (in some cases to 0%) by choosing Primary Input (PI) MLVs such that resultant inputs to internal nodes are also MLVs. Compared to existing work, experimental results for DSP filters simulated in 16nm technology demonstrated leakage savings of 93.6% and 89.2% for top-down and bottom-up approaches with no area overhead. Shilpa Pendyala, Sheikh Ariful Islam, Srinivas Katkoori |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2019 | Locomotion in virtual reality for room scale tracked areas
Evren Bozgeyikli, Andrew Raij, Srinivas Katkoori, Rajiv V. Dubey |
Int. J. Hum. Comput. Stud. | 3 |
| 2018 | Minimizing Performance and Energy Overheads Due to Fanout In Memristor based Logic ImplementationsabstractTo enable memristor based in-memory computing, crossbar architecture is considered as one of the most preferred structures that provide high density. A given logic function is first synthesized as a netlist of NOR and NOT gates and then it is mapped to the crossbar architecture. In this approach, fanout incurs significant performance and energy overheads. As fanout helps in compact designs, it is essential that we investigate ways to minimize the overheads. In this work, a novel approach has been proposed to address this problem - instead of copying the logic value as inputs to the driven memristors, we propose that the controller reads the logic value and then applies it in parallel to the driven memristors. In comparison to recently published works, experimental evaluation on ISCAS85 benchmarks resulted in average performance improvements of 51.08%, 38.66%, and 63.18% considering three different mapping scenarios (average, best, and worst). In regards to energy dissipation, we have also obtained average improvements of 91.30%, 88.53%, and 74.04% considering the aforementioned scenarios. Md. Adnan Zaman, Srinivas Katkoori |
VLSI-SoC | 2 |
| 2018 | CSRO-Based Reconfigurable True Random Number Generator Using RRAMabstractIn this paper, we propose a high-speed (kilohertz-megahertz), reconfigurable current starved ring oscillator (CSRO)-based true random number generator (TRNG) design. The proposed TRNG exploits the intradevice stochastic variations in resistive RAM switching parameters and random telegraph noise (RTN). We demonstrate the effect of RTN on the jitter of CSRO oscillations. We also propose a methodology to reconfigure the TRNG to generate new random numbers. The proposed 10-bit TRNG is validated by NIST test suite for randomness in the data stream. Energy/bit is 22.8 fJ for generation, and the speed of random data generation is 6 MHz. Security vulnerabilities and countermeasures of the proposed TRNG are also investigated. Rekha Govindaraj, Swaroop Ghosh, Srinivas Katkoori |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Memory access pattern based insider threat detection in big data systemsabstractBig data platforms like Hadoop and Spark are being widely adopted both by academia and industry. In this paper, we propose a runtime intrusion detection technique that understands and works according to the memory properties of such distributed compute platforms. The proposed method is based on runtime analysis of memory access patterns of tasks running on the slave nodes of a distributed compute cluster. First, every slave node of the cluster creates a behavior profile for each task it executes. A behavior profile includes information representing the sizes of private & shared memory accesses made by a task during execution. Then, each process behavior profile is shared with other replica nodes that are scheduled to execute the same task on their copy of the same data. Next, these replica nodes verify their local tasks with the help of the information embedded in the received behavior profiles. This step is realized by running Principal Component Analysis (PCA) on the memory access patterns. Finally, nodes share their observations for consensus and report a possible intrusion to the master node if they find any discrepancy. This is a position paper and hence the proposed solution was tested and proved to work in real-time while executing the terasort mapreduce example on a small hadoop cluster. Santosh Aditham, N. Ranganathan, Srinivas Katkoori |
IEEE BigData | 3 |
| 2014 | Self similarity and interval arithmetic based leakage optimization in RTL datapathsabstractLow leakage input vector determination in data path intensive circuits is often not feasible through exhaustive simulation. Hence, top down interval propagation technique for low leakage vector determination is proposed in this paper. This technique is a variation to the heuristic used in [1]. For each RTL module, several low leakage intervals are identified. As the module size increases, exhaustive simulation to find the low leakage vector is not feasible. Further, we need to search the entire input space uniformly to obtain as many low leakage intervals as possible. Based on empirical observations, we observed self similarity in the leakage distribution of adder/multiplier modules when input space is partitioned into smaller cells. This property enables uniform search of low leakage vectors in the entire input space. Also, time taken for characterization increases linearly with the module size. Hence, this technique is scalable to higher bit width modules with acceptable characterization time. We propose a self similarity based Monte Carlo simulation for optimum low leakage interval characterization of RTL modules. The interval propagation is then implemented with the low leakage intervals obtained in the characterization. This yields a reduced low leakage interval set at the primary inputs. The reduced set of intervals is further processed with simulated annealing to arrive at the best low leakage vector at the primary inputs. By applying this low leakage vector, the entire circuit is put in low leakage state. Experimental results for DSP filters simulated in 16nm technology demonstrated leakage savings of 93.6% with no area overhead. Shilpa Pendyala, Srinivas Katkoori |
VLSI-SoC | 2 |
| 2013 | A multi-parameter functional side-channel analysis method for hardware trust verificationabstractA hardware Trojan is a modification to a hardware design which inserts undesired or malicious functionality. They pose a substantial security risk, and as such, rapid, reliable detection of these Trojans has become a critical necessity. In this paper, we propose a method for detecting compromised designs quickly and effectively. The method involves a multi-parameter analysis of the design and statistical analysis to determine which designs have been compromised. We also briefly discuss a supplemental method of confirming our results using a targeted method of FPGA design analysis. This method was proposed for consideration in the Cyber Security Awareness Week (CSAW) 2012 Embedded Systems Challenge (ESC) hosted by the Polytechnic Institute of New York University. This served to independently verify our results. The method was awarded first place in the competition. Christopher Bell, Matthew Lewandowski, Srinivas Katkoori |
VTS | 3 |
| 2012 | Interval arithmetic based input vector control for RTL subthreshold leakage minimization
Shilpa Pendyala, Srinivas Katkoori |
VLSI-SoC | 2 |
| 2011 | State-Retentive Power Gating of Register Files in Multicore Processors Featuring Multithreaded In-Order CoresabstractIn this work, we investigate state-retentive power gating of register files for leakage reduction in multicore processors supporting multithreading. In an in-order core, when a thread gets blocked due to a memory stall, the corresponding register file can be placed in a low leakage state through power gating for leakage reduction. When the memory stall gets resolved, the register file is activated for being accessed again. Since the contents of the register file are not lost and restored on wakeup, this is referred to as state-retentive power gating of register files. While state-retentive power gating in single cores has been studied in the literature, it is being investigated for multicore architectures for the first time in this work. We propose specific techniques to implement state-retentive power gating for three different multicore processor configurations based on the multithreading model: 1) coarse-grained multithreading, 2) fine-grained multithreading, and 3) simultaneous multithreading. The proposed techniques can be implemented as design extensions within the control units of the in-order cores. Each technique uses two different modes of leakage states: low-leakage savings and low wake-up and high-leakage savings and high wake-up latency. The overhead due to wake-up latency is completely avoided in two techniques while it is hidden for most part in the third approach, either by overlapping the wake-up process with the thread context switching latency or by executing instructions from other threads ready for execution. The proposed techniques were evaluated through simulations with multiprogrammed workloads comprised of SPEC 2000 integer benchmarks. Experimental results show that in an 8-core processor executing 64 threads, the average leakage savings were 42 percent in coarse-grained multithreading, while they were between seven percent and eight percent for finegrained and simultaneous multithreading. Soumyaroop Roy, N. Ranganathan, Srinivas Katkoori |
IEEE Trans. Computers | 3 |
| 2011 | Simultaneous Scheduling, Allocation, Binding, Re-Ordering, and Encoding for Crosstalk Pattern Minimization During High-Level SynthesisabstractOn-chip signal crosstalk is a function of switching activity pattern, coupling parasitics, and signal timing. We propose a simulated annealing (SA)-based high-level synthesis algorithm for crosstalk activity minimization for a given data environment. We target bus-based architectures as the bus-lines have well-defined neighborhood (aggressors). Our objective is to minimize worst case crosstalk patterns by exploring synthesis solutions with correlations that do not result in such worst case patterns. Besides synthesis moves, we also incorporate bus re-ordering and data transfer invert encoding. Experimental results for design under resource as well as latency constraints are promising. For a set of nine DSP benchmarks we reduce up to 75% of bus lines that require no shielding lines. The results also show that the designs synthesized through the proposed framework have an average performance improvement by 23.5% compared to un-optimized designs. Hariharan Sankaran, Srinivas Katkoori |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Customizable FPGA IP Core Implementation of a General-Purpose Genetic Algorithm EngineabstractHardware implementation of genetic algorithms (GAs) is gaining importance because of their proven effectiveness as optimization engines for real-time applications (e.g., evolvable hardware). Earlier hardware implementations suffer from major drawbacks such as absence of GA parameter programmability, rigid predefined system architecture, and lack of support for multiple fitness functions. In this paper, we report the design of an IP core that implements a general-purpose GA engine that addresses these problems. Specifically, the proposed GA IP core can be customized in terms of the population size, number of generations, crossover and mutation rates, random number generator seed, and the fitness function. It has been successfully synthesized and verified on a Xilinx Virtex II Pro Field programmable gate arrays device (xc2vp30-7ff896) with only 13% logic slice utilization, 1% block memory utilization for GA memory, and a clock speed of 50 MHz. The GA core has been used as a search engine for real-time adaptive healing but can be tailored to any given application by interfacing with the appropriate application-specific fitness evaluation module as well as the required storage memory and by programming the values of the desired GA parameters. The core is soft in nature i.e., a gate-level netlist is provided which can be readily integrated with the user's system. The performance of the GA core was tested using standard optimization test functions. In the hardware experiments, the proposed core either found the globally optimum solution or found a solution that was within 3.7% of the value of the globally optimal solution. The experimental test setup including the GA core achieved a speedup of around 5.16× over an analogous software implementation. Pradeep Fernando, Srinivas Katkoori, Didier Keymeulen, Ricardo Salem Zebulum, Adrian Stoica |
IEEE Trans. Evol. Comput. | 2 |
| 2010 | TABS: Temperature-Aware Layout-Driven Behavioral SynthesisabstractWith rising power densities in modern VLSI circuits, thermal effects are becoming important in the design of ICs. Elevated chip temperatures have an adverse impact on performance, reliability, power consumption, and cooling costs. To ensure adequate thermal management, all phases of the design flow must account for thermal effects on their design decisions. We present a two-stage simulated annealing-based high-level synthesis technique that combines power minimization with temperature-aware scheduling, binding, and floorplanning. In our technique, the first stage of the simulated annealing algorithm creates a low-power solution, which is then iteratively improved by the second stage to minimize estimated on-chip peak temperature using accurate module-level temperature estimation. We show that minimizing average power alone does not guarantee minimal peak temperatures. However, our approach consistently finds solutions that have lower on-chip peak temperatures and uniform on-chip temperature distributions, compared to a traditional low-power synthesis methodology that minimizes average power. Experiments show that our method reduces peak temperatures on average by 12% and up to 16%, compared to a traditional low-power synthesis algorithm that minimizes average power. These improvements in chip-level temperature distributions are achieved with a modest increase in chip area of under 15% on average. Vyas Krishnan, Srinivas Katkoori |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Compiler-directed leakage reduction in embedded microprocessorsabstractCompiler-directed power gating is an approach in which sleep instructions are inserted appropriately at compile time into the application code to selectively deactivate the functional units in microprocessors during their idle periods to reduce power dissipation due to leakage. Although the effect of code transformations on dynamic and system power has been investigated and reported in the literature, such a study is lacking in the context of power gating. In this paper, we investigate and report how the leakage savings in both integer and floating point units can be improved using machine-dependent and independent optimizations in a compiler-directed power gating framework. In our study, it is ensured that power gating is applied only when the leakage savings are considerably more than the various overheads incurred in its implementation. The target embedded processor is modeled on the ARMv4 architecture, which is modified to support the power gating of its arithmetic functional units. For experimentation, GCC is used as the compiler infrastructure and Simplescalar-ARM is used as the detailed architectural simulator for reporting power and performance metrics for embedded applications belonging to the MiBench and MediaBench benchmark suites. Experimental results suggest that the additional savings in leakage energy due to one or more of the optimizations may vary largely depending on the benchmark. Moreover, the overhead of sleep instructions can be reduced by up to 50 times by performing procedure inlining. Soumyaroop Roy, N. Ranganathan, Srinivas Katkoori |
ICCD | 3 |
| 2009 | Exploring Compiler Optimizations for Enhancing Power GatingabstractPower gating is a circuit level technique for reducing standby leakage in a circuit block by cutting off paths in it between the supply and the ground. A processor architecture that supports power gating of its resources may provide instructions that activate and deactivate those resources as part of the instruction set architecture level. Adequate compiler support is then required so that the power gating instructions can be inserted into the code to deactivate the resources that remain idle for long periods of time during program execution. However, the resource usage in a program depends on the code generated by the compiler. Thus, the code transformations performed by the compiler has an influence on the power gating opportunities of the processor resources. In this work, we explore target independent compiler optimizations that modify the functional unit usage in the loops of a procedure to enhance the opportunities to deactivate functional units in an embedded processor architecture. The optimizations performed on the code are sparse conditional constant propagation, lazy code motion, weak strength reduction, and operator strength reduction. Insertion of power gating instructions is performed by inspecting the idleness of the units in the regions enclosed within loops. We model the processor architecture with power gating support around an ARM core and use the SUIF framework for compiler support. Finally, we use the Simplescalar-ARM distribution to perform power and performance evaluation with a set of benchmarks from MiBench and MediaBench suites. Experimental results indicate that the integer multiplier in the processor core can be power gated for upto 99% of its idle cycles, for integer benchmarks, and upto 93%, for floating point benchmarks, when all the optimizations are performed. Moreover, the energy due to leakage in the functional units for the code with all the optimizations performed can be upto 51% lower, for integer benchmarks, and upto 21% lower, for floating point benchmarks, than that for the unoptimized code. Soumyaroop Roy, N. Ranganathan, Srinivas Katkoori |
ISCAS | 3 |
| 2009 | A Framework for Power-Gating Functional Units in Embedded MicroprocessorsabstractPower gating is a technique commonly used for leakage reduction in integrated circuits. In microprocessors, power gating is implemented by using sleep transistors to selectively deactivate circuit modules that remain idle for sustained periods of time during program execution. In this work, we develop a new framework for power gating the functional units in embedded system microprocessors without degradation in performance. The proposed framework includes an efficient algorithm for idle time estimation, appropriate insertion of sleep instructions within the code, and a method for reactivating the sleeping units only when needed without the use of wakeup instructions. We introduce the notion of loop hierarchy trees (LHTs) to represent the partial ordering of the nested loops within the program. From the control flow graph (CFG) representation of the source program, a forest of LHTs is constructed and is used to identify the maximal subgraphs representing the long idle periods for the functional units. For each subgraph thus identified, a sleep instruction is introduced in the program with a list of corresponding functional units to be deactivated. When an instruction is decoded, the functional units needed for that instruction are automatically activated by the control unit such that the units are ready before the instruction reaches the execute stage. This eliminates the need for wakeup instructions to be inserted into the object code reducing the overheads. In our implementation, the ARM processor architecture was modified and resynthesized to include power gating by developing a CMOS cell library of functional units with the above capabilities. Experimental results are reported for a set of 12 benchmarks chosen from the MiBench suite, which indicate that, on average, our technique reduces the leakage energy in functional units by 31.1% for integer benchmarks and 26.8% for floating-point benchmarks. Soumyaroop Roy, N. Ranganathan, Srinivas Katkoori |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | A customizable FPGA IP core implementation of a general purpose Genetic Algorithm engineabstractHardware implementation of Genetic Algorithms (GA) is gaining importance as genetic algorithms can be effectively used as an optimization engine for real-time applications (for e.g., evolvable hardware). In this work, we report the design of an IP core that implements a general purpose GA engine which has been successfully synthesized and verified on a Xilinx Virtex II Pro FPGA Device (XC2VP30). The placed and routed IP core has an area utilization of only 16% and clock period of 2.2ns (∼450MHz). The GA core can be customized in terms of the population size, number of generations, cross-over and mutation rates, and the random number generator seed. The GA engine can be tailored to a given application by interfacing with the application specific fitness evaluation module as well as the required storage memory (to store the current and new populations). The core is soft in nature i.e., a gate-level netlist is provided which can be readily integrated with the user’s system. Pradeep Fernando, Hariharan Sankaran, Srinivas Katkoori, Didier Keymeulen, Adrian Stoica, Ricardo Salem Zebulum, Rajeshuni Ramesham |
IPDPS | 3 |
| 2007 | Minimizing wire delays by net-topology aware binding during floorplan- driven high level synthesisabstractWith shrinking feature sizes in deep sub-micron technologies, interconnect delays play a dominant role in the cycle time of digital circuits. It is essential to consider the impact of physical design during high-level synthesis. No prior work exists in literature that accounts for the topology of nets resulting from binding decisions during high-level synthesis. This paper presents a novel floorplan-aware high-level synthesis technique that uses accurate net topologies and distributed wire-delay models to guide resource allocation and binding decisions during design-space exploration. The proposed approach tightly integrates a floorplanner with a high-level synthesis binding algorithm. The location of data path modules in the floorplan is used to determine the minimal length RSMT of every net, to which the delay model is applied to accurately estimate delays of multi-terminal nets. Our results show that, when compared to previous approaches, the synthesis technique proposed in this paper reduces wire delays by as much as 48.9% in 70nm technology with an average improvement of 38.6%, and an overhead of only 3.6% in chip area Vyas Krishnan, Srinivas Katkoori |
VLSI-SoC | 2 |
| 2006 | A genetic algorithm for the design space exploration of datapaths during high-level synthesisabstractHigh-level synthesis is comprised of interdependent tasks such as scheduling, allocation, and module selection. For today's very large-scale integration (VLSI) designs, the cost of solving the combined scheduling, allocation, and module selection problem by exhaustive search is prohibitive. However, to meet design objectives, an extensive design space exploration is often critical to obtaining superior designs. We present a framework for efficient design space exploration during high-level synthesis of datapaths for data-dominated applications. The framework uses a genetic algorithm (GA) to concurrently perform scheduling and allocation with the aim of finding schedules and module combinations that lead to superior designs while considering user-specified latency and area constraints. The GA uses a multichromosome representation to encode datapath schedules and module allocations and efficient heuristics to minimize functional and storage area costs, while minimizing circuit latencies. The framework provides the flexibility to perform resource-constrained scheduling, time-constrained scheduling, or a combination of the two, using a simple and fast list-scheduling technique. A graded penalty function is used as an objective function in evaluating the quality of designs to enable the GA to quickly reach areas of the search space where designs meeting user specified criteria are most likely to be found. Since GAs are population-based search heuristics, a unique feature of our framework is its ability to offer a large number of alternative datapath designs, all of which meet design specifications but differ in module, register, and interconnect configurations. Many experiments on well-known benchmarks show the effectiveness of our approach. Vyas Krishnan, Srinivas Katkoori |
IEEE Trans. Evol. Comput. | 2 |
| 2005 | System Level Energy Optimization for Location Aware ComputingabstractWe present a system-level energy optimization technique for a location-aware computing system that provides relevant information about the user’s current location. The system is initialized with a map (in the form of a graph) as well as audio files associated with several locations in the map. The system consists of: GPS receiver module, Serial port, Compact flash module, Stereo codec, Power manager module implementing three sub modules namely, GPS-to-real-world position conversion module (implements algorithm to convert GPS co-ordinates to graph nodes), Nearest-location-search module (implements modified Dijkstra’s algorithm), User speed estimation module. The power manager implements an algorithm that works as follows: at any given location, the algorithm predicts the user speed by exponential average approach. The attenuation factor of this approach can be varied to account for the user speed history. The estimated speed is used to predict the time (say T) required to reach the next nearest location determined by Nearest-location-search module implementing modified Dijkstra’s algorithm. The subsystems are shutdown or switched to low-power mode for time T. After time T, the system will wake up and re-execute the algorithm. Based on a system-level model (VHDL and C), compared to simple time-out and constant update policies,the proposed algorithm results in energy savings in the range 55-99%. Work is ongoing to implement the entire system as a single chip solution. Hariharan Sankaran, Srinivas Katkoori, Umadevi Kailasam |
PerCom | 2 |
| 2005 | Intrabus crosstalk estimation using word-level statisticsabstractWe propose two word-level statistical techniques to estimate the probability of crosstalk events on the signal lines of a system bus. Given the word-level statistical parameters, namely mean, standard deviation, and lag-one temporal correlation coefficient, we analytically estimate the bit-level crosstalk probability. To linearize the complexity and efficiently scale the estimation technique for large bus-widths, we modify the first technique by using a circular right shift procedure that maps disjoint values in a distribution to continuous values in a modified distribution. Experimental results for data streams from different data environments, compared against detailed HSPICE simulations, are presented. The statistical estimators yield average errors less than 7% and 12%, respectively, for bus-widths ranging from 8 to 32 bits. Compared to HSPICE, the execution times are reduced by factors of over 10 /spl times/ for the first technique and over two orders of magnitude for the second technique. The statistical approaches are shown to be compatible with existing bus reordering techniques. Suvodeep Gupta, Srinivas Katkoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2004 | A Fast Word-Level Statistical Estimator of Intra-Bus CrosstalkabstractGiven word-level statistics, namely mean, standard deviation, and lag-one temporal correlation of input data, we estimate the bit-level crosstalk probability on a system bus using a non-enumerative statistical approach. We introduce a sampling technique for fast evaluation of integrals during the estimation process. We had proposed two techniques previously - (a) a stream-based estimator that counts crosstalk events on a bus; and (b) a statistical enumeration technique that enumerates crosstalk-producing values on a bus and computes their occurrence probability. Both these techniques suffer from exponential time complexity with respect to the bus-width. In this work, we propose a statistical non-enumerative technique that has linear time complexity with respect to the bus-width. We achieve the linear complexity by resorting to: (1) manipulating the data stream to make the crosstalk-producing values contiguous and (2) sampling the distribution function and storing it as a lookup table. Experimental results for data streams from different data environments are presented, compared against the stream-based approach. Average errors of less than 12% are obtained for bus-widths ranging from 8b to 32b. Suvodeep Gupta, Srinivas Katkoori |
DATE | 2 |
| 2004 | Power minimization algorithms for LUT-based FPGA technology mappingabstractWe study the technology mapping problem for LUT-based FPGAs targeting at power minimization. The problem has been proved to be NP-hard previously. Therefore, we present an efficient heuristic algorithm to generate low-power mapping solutions. The key idea is to compute and select low-power K -feasible cuts by an efficient incremental network flow computation method. Experimental results show that our algorithm reduces power consumption as well as area over the best algorithms reported in the literature. In addition, we present an extension to compute depth-optimal low-power mappings. Compared with Cutmap, a depth-optimal mapper with simultaneous area minimization, we achieve a 14% power savings on average without any depth penalty. Srinivas Katkoori, Wai-Kei Mak |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2004 | Ant colony system application to macrocell overlap removalabstractWe present a novel macrocell overlap removal algorithm, based on the ant colony optimization metaheuristic. The procedure generates a feasible placement from a relative placement with overlaps produced by some placement algorithms such as quadratic programming and force-directed. It uses the concept of ant colonies, a set of agents that work together to improve an existing solution. Each ant in the colony will generate a placement based on the relative positions of the cells and feedback information about the best placements generated by previous colonies. The solution of each ant is improved by using a local optimization procedure which reduces the unused space. The worst runtime is O(n/sup 3/), but the average runtime can be reduced to O(n/sup 2/). Stelian Alupoaei, Srinivas Katkoori |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Efficient LUT-based FPGA technology mapping for power minimizationabstractWe study the technology mapping problem for LUT-based FPGAs targeting at power minimization. The problem has been proved to be NP-hard previously. Hence, we present an efficient heuristic to compute low-power mapping solutions. The major distinction of our work from previous ones is that while generating a LUT, we look ahead at the impact of the mapping selection of this LUT on the power consumption of the remaining network. We choose the mapping that results in the least estimated overall power consumption. The key idea is to compute low-power K-feasible cuts by an efficient incremental network flow computation method. Experimental results show that our algorithm reduces both power consumption and area over the previous algorithms reported in the literature. Wai-Kei Mak, Srinivas Katkoori |
ASP-DAC | 3 |
| 2003 | KnapBind: An Area-Efficient Binding Algorithm for Low-leakage DatapathsabstractLow leakage power datapaths can be synthesized using multithreshold CMOS (MTCMOS) modules. MTCMOS modules can be turned ON or OFF, using sleep signals. The controller in a digital system can be automatically synthesized to generate these sleep signals to turn OFF idle modules, thus minimizing leakage power. In order to sustain performance, the sleep transistor needs to be sized to large widths. This leads to a significant area overhead. We propose a binding algorithm, based on the 0-1 Knapsack algorithm, to selectively bind modules in a datapath to MTCMOS modules, achieving the optimizing leakage power within a given area constraint. We present results for five data dominated DSP circuits, at 100 nm technology node. Chandramouli Gopalakrishnan, Srinivas Katkoori |
ICCD | 2 |
| 2002 | An efficient register optimization algorithm for high-level synthesis from hierarchical behavioral specificationsabstractWe address the problem of register optimization that arises during high-level synthesis from modular hierarchical behavioral specifications. Register optimization is the process of grouping carriers such that each group can be safely allocated to a hardware register. Global register optimization by inline expansion involves flattening the module hierarchy and using a heuristic register optimization procedure on the flattened description. Although inline expansion yields a near-optimal number of registers, it is very time consuming due to the large number of carrier compatibility relationships that must be considered. We present an efficient register optimization algorithm that achieves nearly the same effect of inline expansion without actually inline expanding. The distinguishing feature of the proposed algorithm is that it employs a hierarchical optimization phase which effectively exploits the properties of the module call graph and information gathered during local carrier lifecycle analysis of each module. Experimental results on a number of benchmarks show that the proposed algorithm produces nearly the same number of registers as inline expansion based global optimization and is faster by a factor of 7.0. Ranga Vemuri, Srinivas Katkoori, Meenakshi Kaul, Jay Roy |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2002 | Net-based force-directed macrocell placement for wirelength optimizationabstractWe propose a net-based hierarchical macrocell placement such that "net placement" dictates the cell placement. The proposed approach has four phases. 1) Net clustering and net-level floorplanning phase: a weighted net dependency graph is built from the input register-transfer-level netlist. Clusters of nets are then formed by clique partitioning and a net-cluster level floorplan is obtained by simulated annealing. The floorplan defines the regions where the nets in each cluster must be routed. 2) Force-directed net placement phase: a force-directed net placement is performed which yields a coarse net-level placement without consideration for the cell placement. 3) Iterative net terminal and cell placement phase: a force-directed net and cell placement is performed iteratively. The terminals of a net are free to move under the influence of forces in the quest for optimal wire length. The cells with high net length cost may "jump" out of local minima by ignoring the rejection forces. The overlaps are reduced by employing electrostatic rejection forces. 4) Overlap removal and input/output (I/O) pin assignment phase: Overlap removal is performed by a grid-based heuristic. I/O pin assignment is performed by minimum-weight bipartite matching. Placements generated by the proposed approach are compared with those generated by Cadence Silicon Ensemble and the O-tree floorplanning algorithm. On average, the proposed approach improves both the total wire length and longest wire length by 18.9% and 28.3%, respectively, with an average penalty of 5.6% area overhead. Stelian Alupoaei, Srinivas Katkoori |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | Scheduling for low power under resource and latency constraintsabstractWe extend the Force-Directed List Scheduling (FDLS) algorithm proposed by Paulin and Knight (1989). In any time step, if the number of ready operations exceeds the available functional resources, then some operations must be deferred. The concept of "force" introduced by Paulin and Knight captures the effect of deferring an operation on the schedule length: larger the force, lower the likelihood of schedule length increase due to the operation's deferral. We develop a power cost function that captures the effect of an operation's deferral on the total power consumption of the design. The novelty of the work lies in heuristically determining the "best" time-step for an operation such that the overall power consumption is minimized without sacrificing the design throughput. The power-delay cost function proposed at the operation-level facilitates such an exploration. Experimental results show power savings of up to 60% with an average power savings of 23% at datapath level, 3% at the controller level, and 14% at the design-level. Srinivas Katkoori, Ranga Vemuri |
ISCAS | 1 |
| 1999 | Accurate Resource Estimation Algorithms for Behavioral SynthesisabstractGiven a scheduled data flow graph the functional, storage, and interconnect (multiplexors) resources are analytically estimated taking into account the effects of post-scheduling tasks. Complexity of the controller implementation is also estimated. The novelty of this work lies in predicting the effects of the post-scheduling task on the final amount of resources, the effects of data path resource optimization on the controller complexity. Experimental results show high correlation between estimated and actual numbers. Srinivas Katkoori, Ranga Vemuri |
Great Lakes Symposium on VLSI | 1 |
| 1997 | A constructive method for data path area estimation during high-level VLSI synthesisabstractIn this paper we present a fast and computationally efficient deterministic method for estimating the area of a register transfer level datapath obtained during high level VLSI synthesis. The estimation makes use of a RT level netlist along with a pre-synthesized library of RT level components. The layout area is estimated using a quadratic programming based framework to get a quick module allocation and generating a topological floorplan which is then followed by heuristic algorithms for mapping RTL modules and their interconnections on a standard cell based layout design style. Experiments on a suite of benchmark examples show promising results with reliable accuracy. Natesan Venkateswaran, Srinivas Katkoori, Dinesh Bhatia, Ranga Vemuri |
ASP-DAC | 3 |
| 1996 | Simulation based architectural power estimation for PLA-based controllersabstractWe present an architectural power simulation technique for PLA-based controllers. The contributions of this work are (1) a simple but efficient power characterization of PLAs; and (2) a strategy for developing a simulatable power model from the input description. Node Switching Capacitance (NSC) of a sub-component (such as AND plane) in a PLA is the average capacitance switched by a node in the sub-component, when the node undergoes a power consuming transition (0/spl rarr/1). Power characterization involves extracting NSC equations for different sub-components as a function of input size, output size and number of terms. Prototype PLAs whose are employed to derive NSC equations for a given technology. The input description is modified for power simulation by adding NSC equations with dependent variables instanced to the controller's parameters. For a given input sequence, the modified VHDL description is simulated to estimate the total power consumption. Experimental Results are obtained with average estimation error of 10.48% with a minimum error of 0.19% and maximum error of 21.90%. Srinivas Katkoori, Ranga Vemuri |
ISLPED | 1 |
| 1995 | High level profiling based low power synthesis techniqueabstractWe present a profiling based technique for power estimation. This technique is implemented in the PDSS (Profile Driven Synthesis System) for the synthesis of low power designs. Initially, each module in the module library is characterized for the average switching capacitance per input vector. The input description is simulated using user-specified set of input vectors to collect the profile data for various operators and carriers. The profile data, in conjunction with the pre-characterized module library is used to estimate the total capacitance switched by each of the valid schedules produced by the PDSS scheduler. A valid schedule is one which satisfies other constants such as area and delay. The schedule with the least switching capacitance estimate is further synthesized to the layout level. Results show an average deviation of 12% compared with the actual switching capacitance values at the layout level. Srinivas Katkoori, Nand Kumar, Ranga Vemuri |
ICCD | 1 |