Vivek Joy Kozhikkottu

dblp:36/10053 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Memory systems · 27% Emerging computing paradigms · 25% Electronic design automation · 24%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
logic synthesis
0.622020
Logic Synthesis of Approximate Circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
SALSA: systematic logic synthesis of approximate circuits · DAC 2012
Emerging computing paradigms
approximate computing
0.522020
Logic Synthesis of Approximate Circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
SALSA: systematic logic synthesis of approximate circuits · DAC 2012
Emerging computing paradigms › approximate computing
approximate circuit design
0.412020
Logic Synthesis of Approximate Circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Electronic design automation › logic synthesis › logic optimization
approximate logic synthesis
0.412020
Logic Synthesis of Approximate Circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Memory systems
cache design
0.212016
Cache Design with Domain Wall Memory · IEEE Trans. Computers 2016
Memory systems
non-volatile memory
0.212016
Cache Design with Domain Wall Memory · IEEE Trans. Computers 2016
Memory systems › emerging memory technologies
spintronic memory
0.212016
Cache Design with Domain Wall Memory · IEEE Trans. Computers 2016
Memory systems
cache management
0.212014
Variation Aware Cache Partitioning for Multithreaded Programs · DAC 2014
Memory systems › cache management
cache partitioning
0.212014
Variation Aware Cache Partitioning for Multithreaded Programs · DAC 2014
Processor architecture and microarchitecture
multicore design
0.212014
Variation Aware Cache Partitioning for Multithreaded Programs · DAC 2014
Processor architecture and microarchitecture › instruction scheduling
variation-aware scheduling
0.212014
Variation Aware Cache Partitioning for Multithreaded Programs · DAC 2014
Emerging computing paradigms › approximate computing
approximate circuit synthesis
0.112012
SALSA: systematic logic synthesis of approximate circuits · DAC 2012
Hardware reliability and fault tolerance
error recovery
0.112012
Recovery-based design for variation-tolerant SoCs · DAC 2012
Hardware reliability and fault tolerance › error recovery
recovery-based design
0.112012
Recovery-based design for variation-tolerant SoCs · DAC 2012
Integrated circuit design › variation-aware design
variation-tolerant design
0.112012
Recovery-based design for variation-tolerant SoCs · DAC 2012
Energy-efficient computing › memory energy efficiency
low-power cache design
0.112016
Cache Design with Domain Wall Memory · IEEE Trans. Computers 2016
Parallel and multicore computing › thread-level parallelism
multithreaded applications
0.112014
Variation Aware Cache Partitioning for Multithreaded Programs · DAC 2014
Integrated circuit design
system-on-chip
0.012012
Recovery-based design for variation-tolerant SoCs · DAC 2012

Methods — techniques the papers use, named apart from their topics

quality constraint encoding · 0.4don't-care-based optimization · 0.4write optimization · 0.2device-circuit-architecture co-design · 0.2boolean equivalence checking · 0.1
YearPublicationVenuePosition
2020 Logic Synthesis of Approximate Circuits
abstract
The ability of several important application domains to tolerate inexactness or approximations in a large fraction of their computations has lead to the advent of approximate computing, a new design paradigm that exploits the intrinsic error-resilient nature to optimize computing platforms for energy and performance. A promising approach to approximate computing is to design approximate circuits, or circuit implementations that are highly efficient but differ in functionality from their original specifications subject to a prespecified quality constraint. While a slew of manual design techniques for approximate circuits have demonstrated their significant potential, a key requirement for their mainstream adoption is to develop automatic methodologies and tools that are general and scalable to any given circuit and quality specification. In this article, we propose SALSA, a systematic methodology for automatic logic synthesis of approximate circuits. Given a golden RTL specification of a circuit and a quality constraint that defines the amount of error that may be introduced in the implementation, SALSA synthesizes an approximate version of the circuit that adheres to the prespecified quality bounds. We make two key contributions: 1) the rigorous formulation of the problem of approximate logic synthesis (ALS), enabling the generation of circuits that is corrected by construction and 2) mapping the problem of approximate synthesis into an equivalent traditional logic synthesis problem, thereby allowing the capabilities of existing synthesis tools to be fully utilized for ALS. In order to achieve these benefits, SALSA forms a virtual quality constraint circuit (QCC) that encodes the quality constraints using logic functions called Q-functions. It then captures the flexibility that engendered by them as approximation don't cares (ADCs), which are used for circuit simplification using traditional don't care-based optimization techniques. We utilized SALSA to automatically synthesize approximate circuits ranging from arithmetic building blocks (adders, multipliers, and MAC) to entire datapaths (DCT, FIR, IIR, SAD, FFT Butterfly, and Euclidean distance), demonstrating scalability and significant improvements in area (1.1× to 1.85× for tight error constraints, and 1.2× to 4.75× for relaxed error constraints) and power (1.15× to 1.75× for tight error constraints, and 1.3× to 5.25× for relaxed error constraints).
Swagath Venkataramani, Vivek Joy Kozhikkottu, Amit Sabne, Kaushik Roy 0001, Anand Raghunathan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 Cache Design with Domain Wall Memory
abstract
Domain wall memory (DWM) is a recently developed spin-based memory technology in which several bits of data are densely packed into the domains of a ferromagnetic wire. DWM has shown great promise in enabling non-volatile memory with very high density and energy efficiency, and has been explored for secondary storage and off-chip memory. In this work, we explore the use of DWM within the on-chip cache hierarchy of general purpose computing platforms. Our work is motivated by the fact that DWMs enable much higher density compared to SRAM, DRAM, and other spin-based memory technologies such as STT-MRAM. However, DWMs also pose the unique challenge of serial access to the bits stored in a cell, leading to large and variable access latencies. In addition, DWMs share the inherent write inefficiency of other spin-based memories. We propose TapeCache, a DWM-based cache design that employs device, circuit, and architectural techniques to address these challenges. At the device level, we perform write optimization by employing a new write mechanism based on domain wall shifts to achieve fast, energy-efficient writes in DWM. At the circuit level, we propose different DWM bit-cell designs that are tailored to the distinct architectural requirements of different levels in the cache hierarchy. At the architecture level, we propose a new cache organization and suitable management policies that mitigate the performance penalty arising from serial access to bits in a DWM cell. We show that the holistic device-circuit-architecture co-design enables all the levels in the cache hierarchy to be realized using DWM and benefit from its improved density. Over a wide range of SPEC CPU 2006 benchmarks, TapeCache achieves an average energy improvement of 7.5x, with virtually identical performance and 7.8x improvement in area, compared to an iso-capacity SRAM cache. Compared to an iso-capacity STT-MRAM cache, TapeCache obtains 3.1x improvement in area and 2x average energy savings along with 1.1 percent performance improvement.
Rangharajan Venkatesan, Vivek Joy Kozhikkottu, Mrigank Sharad, Charles Augustine, Arijit Raychowdhury, Kaushik Roy 0001, Anand Raghunathan
IEEE Trans. Computers2
2016 Emulation-Based Analysis of System-on-Chip Performance Under Variations
abstract
The scaling of integrated circuits into the nanometer regime has led to variations emerging as a primary design concern. Most efforts in the area of variation-tolerant design have focused on the physical, circuit, and logic levels of abstraction. However, inevitable increases in the magnitude of variations with scaling have elevated them to a design concern that must be addressed starting at the system level. We address the problem of analyzing the performance of system-onchip (SoC) architectures in the presence of variations. A modern SoC is a complex ensemble of components that are organized into multiple voltage and frequency domains or islands. The impact of variations on the clock frequencies of individual SoC components may be analyzed using existing tools, such as circuit-level statistical timing analysis. However, the key challenge that needs to be addressed is how to translate these component-level clock frequency distributions into a system-level performance distribution. This task is particularly complex and challenging due to the interdependences between components' execution, indirect effects of shared resources, and interactions between multiple system-level execution paths. We argue that an accurate variation-aware performance analysis requires Monte Carlo-based repeated system execution. We describe a framework variability emulation for SoC performance analysis (VESPA)-that leverages emulation to significantly speed up the performance analysis without sacrificing the generality and accuracy achieved by Monte Carlo-based simulation. We further improve the efficiency of VESPA by utilizing correlated sampling to reduce the number of samples needed for Monte Carlo simulations. We demonstrate the utility of VESPA by applying it to design variation-tolerant architectures for three example SoCs. Our experiments show the performance improvements of ~180× compared with the state-of-the-art hardware-software cosimulation tools and also underscore the potential of VESPA to enable variation-aware design and exploration at the system level.
Vivek Joy Kozhikkottu, Rangharajan Venkatesan, Anand Raghunathan, Sujit Dey
IEEE Trans. Very Large Scale Integr. Syst.1
2014 Variation Aware Cache Partitioning for Multithreaded Programs
abstract
Multithreaded programs are commonly written and optimized for homogeneous multi-core processors assuming equal performance from all the cores. This assumption greatly simplifies the partitioning and balancing of an application's workload across threads; however, it no longer holds when the frequencies of the cores differ due to within-die variations, leading to a degradation in performance. We observe that, in addition to the frequency of the core that it executes on, the performance of a thread is also dependent on the share of shared system resources, such as last-level cache, that it receives. We propose variation-aware cache partitioning as an approach to redress the variation-induced imbalance in the execution times of threads, thereby improving the performance of multi-threaded programs. We discuss the challenges involved in realizing our proposal, including synchronization (e.g., barriers) across threads, which results in faster threads being limited by slower threads, the complex and non-linear relationship between a thread's performance and the cache capacity allocated to it, and the fact that different program phases, can respond quite differently to varying cache capacity. We propose a runtime scheme to perform spatio-temporal cache partitioning while considering both chip characteristics (frequency variations) and program characteristics. We evaluate the proposed technique by applying it to an ensemble of variation-impacted multi-cores executing multi-threaded programs from the PARSEC and SPEC-OMP suites, and demonstrate that it results in an average performance improvement of 15% by mitigating the impact of frequency variations.
Vivek Joy Kozhikkottu, Abhisek Pan, Vijay S. Pai, Sujit Dey, Anand Raghunathan
DAC1
2014 Variation tolerant design of a vector processor for recognition, mining and synthesis
abstract
Variations have emerged as one of the most significant challenges facing the design of integrated circuits in nanoscale technologies. As a consequence, variation tolerant design has become essential at all levels of design abstraction.
Vivek Joy Kozhikkottu, Swagath Venkataramani, Sujit Dey, Anand Raghunathan
ISLPED1
2012 Recovery-based design for variation-tolerant SoCs
abstract
Parameter variations have emerged as a significant threat to continued CMOS scaling in the nanometer regime. Due to increasing performance penalties associated with worst-case design, recovery based design has emerged as a promising approach for dealing with the impact of variations. Previous work has applied recovery based design at the circuit and micro-architecture levels of abstraction. In this work, we address the problem of designing variation-tolerant SoCs using the recovery based design paradigm. We demonstrate that a monolithic implementation of recovery based design fails to scale for large SoCs. We propose the concept of recovery islands, wherein each island consists of one or more SoC components that can recover independent of the rest of the SoC, and demonstrate how our proposal can be easily realized via minor changes to a traditional SoC design flow. We study the tradeoffs involved in applying recovery based design at the system level. We demonstrate that it is critical to account for (i) the inherent diversity of the error-voltage profiles among various components in an SoC, and (ii) the impact of error recovery in a component on overall system performance. We then propose a systematic recovery-based SoC design methodology that partitions a given SoC into recovery islands and also computes the optimal operating points for each island, taking into account the various system level trade-offs involved. We evaluate our framework on three different SoC designs, an 802.11b MAC processor, an MPEG encoder and a Wireless Video Capture system and demonstrate an average of 32% energy savings over conventional designs.
Vivek Joy Kozhikkottu, Sujit Dey, Anand Raghunathan
DAC1
2012 SALSA: systematic logic synthesis of approximate circuits
abstract
Approximate computing has emerged as a new design paradigm that exploits the inherent error resilience of a wide range of application domains by allowing hardware implementations to forsake exact Boolean equivalence with algorithmic specifications. A slew of manual design techniques for approximate computing have been proposed in recent years, but very little effort has been devoted to design automation.
Swagath Venkataramani, Amit Sabne, Vivek Joy Kozhikkottu, Kaushik Roy 0001, Anand Raghunathan
DAC3
2012 TapeCache: a high density, energy efficient cache based on domain wall memory
abstract
Domain Wall Memory (DWM) is a recently developed spin-based memory technology in which several bits of data are densely packed into the domains of a ferromagnetic wire. DWM has shown great promise in enabling non-volatile memory with unprecedented density and high energy efficiency. In this work, we propose TapeCache, a first attempt to employ DWMs as last-level caches in general purpose computing platforms. DWMs enable much higher density compared to SRAM, DRAM, and other spin-based memory technologies such as STT-MRAM. However, they also pose unique challenges such as serial access to the bits stored in a DWM cell, leading to variable access latencies. We propose a novel circuit-architecture co-design for TapeCache, consisting of (i) a multi-port DWM macro-cell optimized for read operations considering the asymmetry in applications' read/write characteristics, and (ii) a new cache organization and suitable management policies that mitigate the performance penalty arising from serial access to bits in a macro-cell. Over a wide range of SPEC 2006 benchmarks, TapeCache achieves 7.8X improvement in area, an average energy improvement of 7.3X, and an average performance improvement of 1.2% compared to an iso-capacity SRAM cache. Compared to an iso-capacity STT-MRAM cache, TapeCache obtains 2.3X improvement in area and 1.4X average energy savings with virtually identical performance.
Rangharajan Venkatesan, Vivek Joy Kozhikkottu, Charles Augustine, Arijit Raychowdhury, Kaushik Roy 0001, Anand Raghunathan
ISLPED2
2011 VESPA: Variability emulation for System-on-Chip performance analysis
abstract
We address the problem of analyzing the performance of System-on-chip (SoC) architectures in the presence of variations. Existing techniques such as gate-level statistical timing analysis compute the distributions of clock frequencies of SoC components. However, we demonstrate that translating component-level characteristics into a system-level performance distribution is a complex and challenging problem due to the inter-dependencies between components' execution, indirect effects of shared resources, and interactions between multiple system-level “execution paths”. We argue that accurate variation-aware system-level performance analysis requires repeated system execution, which is prohibitively slow when based on simulation. Emulation is a widely-used approach to drastically speedup system-level simulation, but it has not been hitherto applied to variation analysis. We describe a framework - Variability Emulation for SoC Performance Analysis (VESPA) - that adapts and applies emulation to the problem of variation aware SoC performance analysis. The proposed framework consists of three phases: component variability characterization, variation-aware emulation setup, and Monte-carlo driven emulation. We demonstrate the utility of the proposed framework by applying it to design variation-aware architectures for two example SoCs - an 802.11 MAC processor and an MPEG encoder. Our results suggest that variability emulation has great potential to enable variation-aware design and exploration at the system level.
Vivek Joy Kozhikkottu, Rangharajan Venkatesan, Anand Raghunathan, Sujit Dey
DATE1