EDBT 2026 Demo / reviewers in the wild / expert
Srihari Cadambi
dblp:99/6719
· DBLP profile ↗
26ranked-venue papers
8as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 8 first-authorSoftware engineering, systems software and programming languages · 4Artificial intelligence and machine learning · 1Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
GPUs and heterogeneous computing · 26% Hardware accelerators and domain-specific architectures · 22% Parallel and multicore computing · 15% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 47% Program analysis · 26% Program verification · 26% | |
| Network and information security
1 paper |
Network security · 100% | |
| Computer networks
2 papers |
Routing and switching · 100% |
Topics — the 30 heaviest of 43, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.2 | 2 | 2010 | A dynamically configurable coprocessor for convolutional neural networks · ISCA 2010 A Massively Parallel Digital Learning Processor · NIPS 2008 |
Processor architecture and microarchitecture › many-core architecture
intel xeon phi |
0.2 | 1 | 2013 | COSMIC: middleware for high performance and reliable multiprocessing on xeon phi coprocessors · HPDC 2013 |
Parallel and multicore computing
parallel programming runtimes |
0.2 | 1 | 2013 | COSMIC: middleware for high performance and reliable multiprocessing on xeon phi coprocessors · HPDC 2013 |
Compilers and program optimization › loop transformation
loop fusion |
0.1 | 1 | 2012 | Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation · MICRO 2012 |
Compilers and program optimization › deep learning compiler
operator fusion |
0.1 | 1 | 2012 | Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation · MICRO 2012 |
GPUs and heterogeneous computing › GPU resource management
GPU cluster resource management |
0.1 | 1 | 2012 | Interference-driven resource management for GPU-based heterogeneous clusters · HPDC 2012 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2012 | Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation · MICRO 2012 |
GPUs and heterogeneous computing
GPU sharing |
0.1 | 1 | 2012 | Interference-driven resource management for GPU-based heterogeneous clusters · HPDC 2012 |
Parallel and multicore computing › parallel scheduling › resource-aware scheduling
interference-aware scheduling |
0.1 | 1 | 2012 | Interference-driven resource management for GPU-based heterogeneous clusters · HPDC 2012 |
GPUs and heterogeneous computing › GPU kernel optimization
kernel fusion |
0.1 | 1 | 2012 | Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation · MICRO 2012 |
Hardware accelerators and domain-specific architectures › accelerator architecture
programmable accelerator |
0.1 | 1 | 2012 | A Massively Parallel, Energy Efficient Programmable Accelerator for Learning and Classification · ACM Trans. Archit. Code Optim. 2012 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator |
0.1 | 1 | 2010 | A dynamically configurable coprocessor for convolutional neural networks · ISCA 2010 |
Reconfigurable computing and FPGAs › reconfigurable computing
reconfigurable accelerator |
0.1 | 1 | 2010 | A dynamically configurable coprocessor for convolutional neural networks · ISCA 2010 |
Program analysis › static analysis
abstract interpretation |
0.1 | 1 | 2008 | Bitwidth Reduction via Symbolic Interval Analysis for Software Model Checking · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Program analysis › static analysis › abstract interpretation
interval analysis |
0.1 | 1 | 2008 | Bitwidth Reduction via Symbolic Interval Analysis for Software Model Checking · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Program verification › model checking
software model checking |
0.1 | 1 | 2008 | Bitwidth Reduction via Symbolic Interval Analysis for Software Model Checking · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Program verification › model checking
state space reduction |
0.1 | 1 | 2008 | Bitwidth Reduction via Symbolic Interval Analysis for Software Model Checking · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Network security › intrusion detection and prevention
intrusion detection |
0.1 | 1 | 2007 | Memory-Efficient Regular Expression Search Using State Merging · INFOCOM 2007 |
Network security › intrusion detection and prevention › intrusion detection › pattern matching
regular expression matching |
0.1 | 1 | 2007 | Memory-Efficient Regular Expression Search Using State Merging · INFOCOM 2007 |
Routing and switching › IP lookup
longest prefix matching |
0.1 | 1 | 2006 | Chisel: A Storage-efficient, Collision-free Hash-based Network Processing Architecture · ISCA 2006 |
Routing and switching
routing |
0.1 | 1 | 2006 | Chisel: A Storage-efficient, Collision-free Hash-based Network Processing Architecture · ISCA 2006 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.1 | 1 | 2006 | Signature-based workload estimation for mobile 3D graphics · DAC 2006 |
Energy-efficient computing
power management |
0.1 | 1 | 2006 | Signature-based workload estimation for mobile 3D graphics · DAC 2006 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2006 | Signature-based workload estimation for mobile 3D graphics · DAC 2006 |
Query processing and optimization › query execution › hardware-accelerated query processing
GPU-accelerated query processing |
0.0 | 1 | 2012 | Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation · MICRO 2012 |
High-performance computing › cluster computing
heterogeneous clusters |
0.0 | 1 | 2012 | Interference-driven resource management for GPU-based heterogeneous clusters · HPDC 2012 |
Memory systems
processing-in-memory |
0.0 | 1 | 2012 | A Massively Parallel, Energy Efficient Programmable Accelerator for Learning and Classification · ACM Trans. Archit. Code Optim. 2012 |
High-performance computing › collective communication
reduction operations |
0.0 | 1 | 2012 | A Massively Parallel, Energy Efficient Programmable Accelerator for Learning and Classification · ACM Trans. Archit. Code Optim. 2012 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.0 | 1 | 2002 | A fast, inexpensive and scalable hardware acceleration technique for functional simulation · DAC 2002 |
Electronic design automation › hardware simulation
functional simulation |
0.0 | 1 | 2002 | A fast, inexpensive and scalable hardware acceleration technique for functional simulation · DAC 2002 |
Methods — techniques the papers use, named apart from their topics
producer-consumer dependence classification · 0.4VLIW compilation · 0.2middleware · 0.2transition labeling · 0.1state merging · 0.1interference modeling · 0.1prefix collapsing · 0.1bloomier filter · 0.1symbolic interval analysis · 0.1multiply-accumulate · 0.1SIMD · 0.1signature-based prediction · 0.1FPGA prototyping · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | A Coprocessor Sharing-Aware Scheduler for Xeon Phi-Based Compute ClustersabstractWe propose a cluster scheduling technique for compute clusters with Xeon Phi coprocessors. Even though the Xeon Phi runs Linux which allows multiprocessing, cluster schedulers generally do not allow jobs to share coprocessors because sharing can cause oversubscription of coprocessor memory and thread resources. It has been shown that memory or thread oversubscription on a many core like the Phi results in job crashes or drastic performance loss. We first show that such an exclusive device allocation policy causes severe coprocessor underutilization: for typical workloads, on average only 38% of the Xeon Phi cores are busy across the cluster. Then, to improve coprocessor utilization, we propose a scheduling technique that enables safe coprocessor sharing without resource oversubscription. Jobs specify their maximum memory and thread requirements, and our scheduler packs as many jobs as possible on each coprocessor in the cluster, subject to resource limits. We solve this problem using a greedy approach at the cluster level combined with a knapsack-based algorithm for each node. Every coprocessor is modeled as a knapsack and jobs are packed into each knapsack with the goal of maximizing job concurrency, i.e., as many jobs as possible executing on each coprocessor. Given a set of jobs, we show that this strategy of packing for high concurrency is a good proxy for (i) reducing make span, without the need for users to specify job execution times and (ii) reducing coprocessor footprint, or the number of coprocessors required to finish the jobs without increasing make span. We implement the entire system as a seamless add on to Condor, a popular distributed job scheduler, and show make span and footprint reductions of more than 50% across a wide range of workloads. Giuseppe Coviello, Srihari Cadambi, Srimat T. Chakradhar |
IPDPS | 2 |
| 2013 | COSMIC: middleware for high performance and reliable multiprocessing on xeon phi coprocessors
Srihari Cadambi, Giuseppe Coviello, Cheng-Hong Li, Rajat Phull, Kunal Rao, Murugan Sankaradass, Srimat T. Chakradhar |
HPDC | 1 |
| 2012 | Interference-driven resource management for GPU-based heterogeneous clustersabstractGPU-based clusters are increasingly being deployed in HPC environments to accelerate a variety of scientific applications. Despite their growing popularity, the GPU devices themselves are under-utilized even for many computationally-intensive jobs. This stems from the fact that the typical GPU usage model is one in which a host processor periodically offloads computationally intensive portions of an application to the coprocessor. Since some portions of code cannot be offloaded to the GPU (for example, code performing network communication in MPI applications), this usage model results in periods of time when the GPU is idle. GPUs could be time-shared across jobs to "fill" these idle periods, but unlike CPU resources such as the cache, the effects of sharing the GPU are not well understood. Specifically, two jobs that time-share a single GPU will experience resource contention and interfere with each other. The resulting slow-down could lead to missed job deadlines. Current cluster managers do not support GPU-sharing, but instead dedicate GPUs to a job for the job's lifetime. Rajat Phull, Cheng-Hong Li, Kunal Rao, Srihari Cadambi, Srimat T. Chakradhar |
HPDC | 4 |
| 2012 | Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU ComputationabstractData warehousing applications represent an emerging application arena that requires the processing of relational queries and computations over massive amounts of data. Modern general purpose GPUs are high bandwidth architectures that potentially offer substantial improvements in throughput for these applications. However, there are significant challenges that arise due to the overheads of data movement through the memory hierarchy and between the GPU and host CPU. This paper proposes data movement optimizations to address these challenges. Inspired in part by loop fusion optimizations in the scientific computing community, we propose kernel fusion as a basis for data movement optimizations. Kernel fusion fuses the code bodies of two GPU kernels to i) reduce data footprint to cut down data movement throughout GPU and CPU memory hierarchy, and ii) enlarge compiler optimization scope. We classify producer consumer dependences between compute kernels into three types, i) fine-grained thread-to-thread dependences, ii) medium-grained thread block dependences, and iii) coarse-grained kernel dependences. Based on this classification, we propose a compiler framework, Kernel Weaver, that can automatically fuse relational algebra operators thereby eliminating redundant data movement. The experiments on NVIDIA Fermi platforms demonstrate that kernel fusion achieves 2.89x speedup in GPU computation and a 2.35x speedup in PCIe transfer time on average across the micro-benchmarks tested. We present key insights, lessons learned, measurements from our compiler implementation, and opportunities for further improvements. Haicheng Wu, Gregory Frederick Diamos, Srihari Cadambi, Sudhakar Yalamanchili |
MICRO | 3 |
| 2012 | A Massively Parallel, Energy Efficient Programmable Accelerator for Learning and ClassificationabstractApplications that use learning and classification algorithms operate on large amounts of unstructured data, and have stringent performance constraints. For such applications, the performance of general purpose processors scales poorly with data size because of their limited support for fine-grained parallelism and absence of software-managed caches. The large intermediate data in these applications also limits achievable performance on many-core processors such as GPUs. To accelerate such learning applications, we present a programmable accelerator that can execute multiple learning and classification algorithms. To architect such an accelerator, we profile five representative workloads, and find that their computationally intensive portions can be formulated as matrix or vector operations generating large amounts of intermediate data, which are then reduced by a secondary operation such as array ranking, finding max/min and aggregation. Our proposed accelerator, called MAPLE, has hundreds of simple processing elements (PEs) laid out in a two-dimensional grid, with two key features. First, it uses dynamic in-memory processing where on-chip memory blocks perform the secondary reduction operations. Second, MAPLE uses banked off-chip memory, and organizes its PEs into independent groups each with its own off-chip memory bank. These two features allow MAPLE to scale its performance with data size. We also present an Atom based energy-efficient heterogeneous system with MAPLE as the accelerator that satisfies the application’s performance requirements at a lower system power. This article describes the MAPLE architecture, explores its design space with a simulator, illustrates how to automatically map application kernels to the hardware, and presents its performance improvement and energy benefits over classic server-based implementations. We implement a 512-PE FPGA prototype of MAPLE and find that it is 1.5-10x faster than a 2.5 GHz quad-core Xeon processor despite running at a modest 125 MHz clock rate. With MAPLE connected to a 1.6GHz dual-core Atom, we show an energy improvement of 38-84% over the Xeon server coupled to a 1.3 GHz 240 core Tesla GPU. Abhinandan Majumdar, Srihari Cadambi, Michela Becchi, Srimat T. Chakradhar, Hans Peter Graf |
ACM Trans. Archit. Code Optim. | 2 |
| 2011 | Symphony: A Scheduler for Client-Server Applications on Coprocessor-Based Heterogeneous ClustersabstractCoprocessors such as GPUs are increasingly being deployed in clusters to process scientific and compute-intensive jobs. In this work, we study if GPU-based heterogeneous clusters can benefit client-server applications. Specifically, we consider the practical situation where multiple client-server applications share a heterogeneous cluster (multi-tenancy), and experience unpredictable variations in incoming client request rates, including steep load spikes. Even for "compute-intensive" client-server applications, it is unclear if a GPU-based cluster can seamlessly deliver acceptable response times in the presence of multi-tenancy and load spikes. We argue that a cluster-level scheduler that is aware of application load, request deadlines and the heterogeneity is necessary in this situation. We propose a novel scheduler called Symphony that enables efficient, dynamic sharing of a GPU-based heterogeneous cluster across multiple concurrently-executing client-server applications, each with arbitrary load spikes. Symphony performs three key tasks: it (i) monitors the load on each application, (ii) collects past performance data and dynamically builds simple performance models of available processing resources and (iii) computes a priority for pending requests based on the above parameters and the requests' slack. Based on this, it reorders client requests across different applications to achieve acceptable response times. We also define how client-server applications should interact with a scheduler such as Symphony, and develop an API to this end. We deploy Symphony as user-space middleware on a high-end heterogeneous cluster with dual quad-core Xeon CPUs and dual NVIDIA Fermi GPUs. An evaluation using representative applications shows that in the presence of load spikes (i) Symphony incurs 2-20× fewer requests that do not meet response time constraints compared with other schedulers, and (ii) in order to achieve the same performance as Symphony, other schedulers need 2× more cluster nodes. M. Mustafa Rafique, Srihari Cadambi, Kunal Rao, Ali Raza Butt, Srimat T. Chakradhar |
CLUSTER | 2 |
| 2010 | A programmable parallel accelerator for learning and classificationabstractFor learning and classification workloads that operate on large amounts of unstructured data with stringent performance constraints, general purpose processor performance scales poorly with data size. In this paper, we present a programmable accelerator for this workload domain. To architect the accelerator, we profile five representative workloads, and find that their computationally intensive portions can be formulated as matrix or vector operations generating large amounts of intermediate data, which are then reduced by a secondary operation such as array ranking, finding max/min and aggregation. The proposed accelerator, called MAPLE, has hundreds of simple processing elements (PEs) laid out in a two-dimensional grid, with two key features. First, it uses in-memory processing where on-chip memory blocks perform the secondary reduction operations. By doing so, the intermediate data are dynamically processed and never stored or sent off-chip. Second, MAPLE uses banked off-chip memory, and organizes its PEs into independent groups each with its own off-chip memory bank. These two features together allow MAPLE to scale its performance with data size. This paper describes the MAPLE architecture, explores its design space with a simulator, and illustrates how to automatically map application kernels to the hardware. We also implement a 512-PE FPGA prototype of MAPLE and find that it is 1.5-10x faster than a 2.5 GHz quad-core Xeon processor despite running at a modest 125 MHz. Srihari Cadambi, Abhinandan Majumdar, Michela Becchi, Srimat T. Chakradhar, Hans Peter Graf |
PACT | 1 |
| 2010 | A dynamically configurable coprocessor for convolutional neural networksabstractConvolutional neural networks (CNN) applications range from recognition and reasoning (such as handwriting recognition, facial expression recognition and video surveillance) to intelligent text applications such as semantic text analysis and natural language processing applications. Two key observations drive the design of a new architecture for CNN. First, CNN workloads exhibit a widely varying mix of three types of parallelism: parallelism within a convolution operation, intra-output parallelism where multiple input sources (features) are combined to create a single output, and inter-output parallelism where multiple, independent outputs (features) are computed simultaneously. Workloads differ significantly across different CNN applications, and across different layers of a CNN. Second, the number of processing elements in an architecture continues to scale (as per Moore's law) much faster than off-chip memory bandwidth (or pin-count) of chips. Based on these two observations, we show that for a given number of processing elements and off-chip memory bandwidth, a new CNN hardware architecture that dynamically configures the hardware on-the-fly to match the specific mix of parallelism in a given workload gives the best throughput performance. Our CNN compiler automatically translates high abstraction network specification into a parallel microprogram (a sequence of low-level VLIW instructions) that is mapped, scheduled and executed by the coprocessor. Compared to a 2.3 GHz quad-core, dual socket Intel Xeon, 1.35 GHz C870 GPU, and a 200 MHz FPGA implementation, our 120 MHz dynamically configurable architecture is 4x to 8x faster. This is the first CNN architecture to achieve real-time video stream processing (25 to 30 frames per second) on a wide range of object detection and recognition tasks. Srimat T. Chakradhar, Murugan Sankaradass, Venkata Jakkula, Srihari Cadambi |
ISCA | 4 |
| 2010 | Data-aware scheduling of legacy kernels on heterogeneous platforms with distributed memoryabstractIn this paper, we describe a runtime to automatically enhance the performance of applications running on heterogeneous platforms consisting of a multi-core (CPU) and a throughput-oriented many-core (GPU). The CPU and GPU are connected by a non-coherent interconnect such as PCI-E, and as such do not have shared memory. Heterogeneous platforms available today such as [9] are of this type. Our goal is to enable the programmer to seamlessly use such a system without rewriting the application and with minimal knowledge of the underlying architectural details. Assuming that applications perform function calls to computational kernels with available CPU and GPU implementations, our runtime achieves this goal by automatically scheduling the kernels and managing data placement. In particular, it intercepts function calls to well-known computational kernels and schedules them on CPU or GPU based on their argument size and location. To improve performance, it defers all data transfers between the CPU and the GPU until necessary. By managing data placement transparently to the programmer, it provides a unified memory view despite the underlying separate memory sub-systems. Michela Becchi, Surendra Byna, Srihari Cadambi, Srimat T. Chakradhar |
SPAA | 3 |
| 2009 | A Massively Parallel Coprocessor for Convolutional Neural NetworksabstractWe present a massively parallel coprocessor for accelerating Convolutional Neural Networks (CNNs), a class of important machine learning algorithms. The coprocessor functional units, consisting of parallel 2D convolution primitives and programmable units performing sub-sampling and non-linear functions specific to CNNs, implement a ldquometa-operatorrdquo to which a CNN may be compiled to. The coprocessor is serviced by distributed off-chip memory banks with large data bandwidth. As a key feature, we use low precision data and further increase the effective memory bandwidth by packing multiple words in every memory operation, and leverage the algorithmpsilas simple data access patterns to use off-chip memory as a scratchpad for intermediate data, critical for CNNs. A CNN is mapped to the coprocessor hardware primitives with instructions to transfer data between the memory and coprocessor. We have implemented a prototype of the CNN coprocessor on an off-the-shelf PCI FPGA card with a single Xilinx Virtex5 LX330T FPGA and 4 DDR2 memory banks totaling 1 GB. The coprocessor prototype can process at the rate of 3.4 billion multiply accumulates per second (GMACs) for CNN forward propagation, a speed that is 31x faster than a software implementation on a 2.2 GHz AMD Opteron processor. For a complete face recognition application with the CNN on the coprocessor and the rest of the image processing tasks on the host, the prototype is 6-10times faster, depending on the host-coprocessor bandwidth. Murugan Sankaradass, Venkata Jakkula, Srihari Cadambi, Srimat T. Chakradhar, Igor Durdanovic, Eric Cosatto, Hans Peter Graf |
ASAP | 3 |
| 2009 | A Massively Parallel FPGA-Based Coprocessor for Support Vector MachinesabstractWe present a massively parallel FPGA-based coprocessor for Support Vector Machines (SVMs), a machine learning algorithm whose applications include recognition tasks such as learning scenes, situations and concepts, and reasoning tasks such as analyzing the recognized scenes and semantics. The coprocessor architecture, targeted at both SVM training and classification, is based on clusters of vector processing elements (VPEs) operating in single-instruction multiple data (SIMD) mode to take advantage of large amounts of data parallelism in the application. We use the FPGA's DSP elements as parallel multiply-accumulators (MACs), a core computation in SVMs. A key feature of the architecture is that it is customized to low precision arithmetic which permits one DSP unit to perform two or more MACs in parallel. Low precision also reduces the required number of parallel off-chip memory accesses by packing multiple data words on the FPGA-memory bus. We have built a prototype using an off-the-shelf PCI-based FPGA card with a Xilinx Virtex 5 FPGA and 1 GB DDR2 memory. For SVM training, we observe application-level end-to-end computation speeds of over 9 billion multiply-accumulates per second (GMACs). For SVM classification, using data packing, the application speed increases to 14 GMACs. The FPGA-based system is about 20times faster than a dual Opteron 2.2 GHz processor CPU, and dissipates around 10 W of power. Srihari Cadambi, Igor Durdanovic, Venkata Jakkula, Murugan Sankaradass, Eric Cosatto, Srimat T. Chakradhar, Hans Peter Graf |
FCCM | 1 |
| 2009 | Using hardware transactional memory for data race detectionabstractWidespread emergence of multicore processors will spur development of parallel applications, exposing programmers to degrees of hardware concurrency hitherto unavailable. Dependable multithreaded software will have to rely on the ability to dynamically detect non-deterministic and notoriously hard to reproduce synchronization bugs manifested through data races. Previous solutions to dynamic data race detection have required specialized hardware, at additional power, design and area costs. We propose RaceTM, a novel approach to data race detection that exploits hardware that will likely be present in future multiprocessors, albeit for a different purpose. In particular, we show how emerging hardware support for transactional memory can be leveraged to aid data race detection. We propose the concept of lightweight debug transactions that exploit the conflict detection mechanisms of transactional memory systems to perform data race detection. We present a proof-of-concept simulation prototype, and evaluate it on data races injected into applications from the SPLASH-2 suite. Our experiments show that this technique is effective at discovering data races and has low performance overhead. Florin Sultan, Srihari Cadambi, Franjo Ivancic, Martin Rötteler |
IPDPS | 3 |
| 2009 | A hybrid nano-CMOS architecture for defect and fault toleranceabstractAs the end of the semiconductor roadmap for CMOS approaches, architectures based on nanoscale molecular devices are attracting attention. Among several alternatives, silicon nanowires and carbon nanotubes are the two most promising nanotechnologies according to the ITRS. These technologies may enable scaling deep into the nanometer regime. However, they suffer from very defect-prone manufacturing processes. Although the reconfigurability property of the nanoscale devices can be used to tolerate high defect rates, it may not be possible to locate all defects. With very high device densities, testing each component may not be possible because of time or technology restrictions. This points to a scenario in which even though the devices are tested, the tests are not very comprehensive at locating defects, and hence the shipped chips are still defective. Moreover, the devices in the nanometer range will be susceptible to transient faults which can produce arbitrary soft errors. Despite these drawbacks, it is possible to make nanoscale architectures practical and realistic by introducing defect and fault tolerance. In this article, we propose and evaluate a hybrid nanowire-CMOS architecture that addresses all three problems—namely high defect rates, unlocated defects, and transient faults—at the same time. This goal is achieved by using multiple levels of redundancy and majority voters. A key aspect of the architecture is that it contains a judicious balance of both nanoscale and traditional CMOS components. A companion to the architecture is a compiler with heuristics to quickly determine if logic can be mapped onto partially defective nanoscale elements. The heuristics make it possible to introduce defect-awareness in placement and routing. The architecture and compiler are evaluated by applying the complete design flow to several benchmarks. Muzaffer O. Simsir, Srihari Cadambi, Franjo Ivancic, Martin Rötteler, Niraj K. Jha |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2008 | A Massively Parallel Digital Learning ProcessorabstractWe present a new, massively parallel architecture for accelerating machine learning algorithms, based on arrays of variable-resolution arithmetic vector processing elements (VPE). Groups of VPEs operate in SIMD (single instruction multiple data) mode, and each group is connected to an independent memory bank. In this way memory bandwidth scales with the number of VPE, and the main data flows are local, keeping power dissipation low. With 256 VPEs, implemented on two FPGA (field programmable gate array) chips, we obtain a sustained speed of 19 GMACS (billion multiply-accumulate per sec.) for SVM training, and 86 GMACS for SVM classification. This performance is more than an order of magnitude higher than that of any FPGA implementation reported so far. The speed on one FPGA is similar to the fastest speeds published on a Graphics Processor for the MNIST problem, despite a clock rate of the FPGA that is six times lower. High performance at low clock rates makes this massively parallel architecture particularly attractive for embedded applications, where low power dissipation is critical. Tests with Convolutional Neural Networks and other learning algorithms are under way now. Hans Peter Graf, Srihari Cadambi, Igor Durdanovic, Venkata Jakkula, Murugan Sankaradass, Eric Cosatto, Srimat T. Chakradhar |
NIPS | 2 |
| 2008 | RaceTM: detecting data races using transactional memoryabstractWidespread emergence of multicore processors will spur development of parallel applications, exposing programmers to more hardware concurrency. Dependable multithreaded software will have to rely on the ability to dynamically detect data races, which are non-deterministic and notoriously hard to reproduce symptoms of synchronization bugs. In this paper, we propose RaceTM, a novel approach that exploits transactional memory support to detect data races. We introduce the concept of lightweight debug transactions that exploit the conflict detection mechanisms of transactional memory systems to perform data race detection. Debug transactions differ from regular transactions in that they do not need to be rolled back, and therefore require no versioning or checkpointing support. Debug transactions do not overlap with a regular transaction, thus providing a transparent mechanism to leverage existing transactional memory support for data race detection. Florin Sultan, Srihari Cadambi, Franjo Ivancic, Martin Rötteler |
SPAA | 3 |
| 2008 | Bitwidth Reduction via Symbolic Interval Analysis for Software Model CheckingabstractThis paper presents a lightweight interval analysis technique for determining the lower and upper bounds for program variables and its application in improving software model checking techniques. The experiments demonstrate that it is an effective approach to alleviate the state explosion problem in software model checking. Aleksandr Zaks, Zijiang Yang 0006, Ilya Shlyakhter, Franjo Ivancic, Srihari Cadambi, Malay K. Ganai, Aarti Gupta, Pranav Ashar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2007 | Memory-Efficient Regular Expression Search Using State MergingabstractPattern matching is a crucial task in several critical network services such as intrusion detection and policy management. As the complexity of rule-sets increases, traditional string matching engines are being replaced by more sophisticated regular expression engines. To keep up with line rates, deal with denial of service attacks and provide predictable resource provisioning, the design of such engines must allow examining payload traffic at several gigabits per second and provide worst case speed guarantees. While regular expression matching using deterministic finite automata (DFA) is a well studied problem in theory, its implementation either in software or specialized hardware is complicated by prohibitive memory requirements. This is especially true for DFAs representing complex regular expressions present in practical rule-sets. In this paper, we introduce a novel method to drastically reduce the DFA memory requirement and still provide worst-case speed guarantees. Specifically, we merge several "non-equivalent" states in a DFA by introducing labels on their input and output transitions. We then propose a data structure to represent the merged states and the transition labels. We show that, with very few assumptions about the original DFA, such a transformation results in significant compression in the DFA representation. We have implemented a state merging and transition labeling algorithm for DFAs, and show that for Snort and Bro security rule-sets, state merging results in memory reductions of an order of magnitude. Michela Becchi, Srihari Cadambi |
INFOCOM | 2 |
| 2006 | Signature-based workload estimation for mobile 3D graphicsabstractUntil recently, most 3D graphics applications had been regarded as too computationally intensive for devices other than desktop computers and gaming consoles. This notion is rapidly changing due to improving screen resolutions and computing capabilities of mass-market handheld devices such as cellular phones and PDAs. As the mobile 3D gaming industry is poised to expand, significant innovations are required to provide users with high-quality 3D experience under limited processing, memory and energy budgets that are characteristic of the mobile domain.Energy saving schemes such as Dynamic Voltage and Frequency Scaling (DVFS), as well as system-level power and performance optimization methods for mobile devices require accurate and fast workload prediction. In this paper, we address the problem of workload prediction for mobile 3D graphics. We propose and describe a signature-based estimation technique for predicting 3D graphics workloads. By analyzing a gaming benchmark, we show that monitoring specific parameters of the 3D pipeline provides better prediction accuracy over conventional approaches. We describe how signatures capture such parameters concisely to make accurate workload predictions. Signature-based prediction is computationally efficient because first, signatures are compact, and second, they do not require elaborate model evaluations. Thus, they are amenable to efficient, real-time prediction. A fundamental difference between signatures and standard history-based predictors is that signatures capture previous outcomes as well as the cause that led to the outcome, and use both to predict future outcomes. We illustrate the utility of signature-based workload estimation technique by using it as a basis for DVFS in 3D graphics pipelines. Bren Mochocki, Kanishka Lahiri, Srihari Cadambi, Xiaobo Sharon Hu |
DAC | 3 |
| 2006 | Power analysis of mobile 3D graphicsabstractThe world of 3D graphics, until recently restricted to high-end workstations and game consoles, is rapidly expanding into the domain of mobile platforms such as cellular phones and PDAs. Even as the mobile chip market is poised to exceed production of 500 million chips per year, incorporation of 3D graphics in handhelds poses several serious challenges to the hardware designer. Compared with other platforms, graphics on handhelds have to contend with limited energy supplies and lower computing horsepower. Nevertheless, images must still be rendered at high quality since handheld screens are typically held closer to the observer's eye, making imperfections and approximations very noticeable. In this paper, we provide an in-depth quantitative analysis of the power consumption of mobile 3D graphics pipelines. We analyze the effects of various 3D graphics factors such as resolution, frame rate, level of detail, lighting and texture maps on power consumption. We demonstrate that significant imbalance exists across the workloads of different graphics pipeline stages. In addition, we illustrate how this imbalance may vary dynamically, depending on the characteristics of the graphics application. Based on this observation, we identify and compare the benefits of candidate dynamic voltage and frequency scaling (DVFS) schemes for mobile 3D graphics pipelines. In our experiments we observe that DVFS for mobile 3D graphics reduces energy by as much as 50% Bren Mochocki, Kanishka Lahiri, Srihari Cadambi |
DATE | 3 |
| 2006 | Chisel: A Storage-efficient, Collision-free Hash-based Network Processing ArchitectureabstractLongest prefix matching (LPM) is a fundamental part of various network processing tasks. Previously proposed approaches for LPM result in prohibitive cost and power dissipation (TCAMs) or in large memory requirements and long lookup latencies (tries), when considering future line-rates, table sizes and key lengths (e.g., IPv6). Hash-based approaches appear to be an excellent candidate for LPM with the possibility of low power, compact storage, and O(1) latencies. However, there are two key problems that hinder their practical deployment as LPM solutions. First, naive hash tables incur collisions and resolve them using chaining, adversely affecting worst-case lookup-rate guarantees that routers must provide. Second, hash functions cannot directly operate on wildcard bits, a requirement for LPM, and current solutions require either considerably complex hardware or large storage space. In this paper we propose a novel architecture which successfully addresses for the first time, both key problems in hash based LPM - making the following contributions: (1) We architect an LPM solution based upon a recently-proposed, collision-free hashing scheme called Bloomier filter, by eliminating its false positives in a storage efficient way. (2) We propose a novel scheme called prefix collapsing, which provides support for wildcard bits with small additional storage and reduced hardware complexity. (3) We exploit prefix collapsing and key characteristics found in real update traces to support fast and incremental updates, a feature generally not available in collision-free hashing schemes Jahangir Hasan, Srihari Cadambi, Venkata Jakkula, Srimat T. Chakradhar |
ISCA | 2 |
| 2002 | A fast, inexpensive and scalable hardware acceleration technique for functional simulationabstractWe introduce a novel approach to accelerating functional simulation. The key attributes of our approach are high-performance, low-cost, scalability and low turn-around-time (TAT). We achieve speedups between 25 and 2000x over zero delay event-driven simulation and between 75 and 1000x over cycle-based simulation on benchmark and industrial circuits while maintaining the cost, scalability and TAT advantages of simulation. Owing to these attributes, we believe that such an approach has potential for very wide deployment as replacement or enhancement for existing simulators. Our technology relies on a VLIW-like virtual simulation processor (SimPLE) mapped to a single FPGA on an off-the-shelf PCI board. Primarily responsible for the speed are (i) parallelism in the processor architecture (ii) high pin count on the FPGA enabling large instruction bandwidth and (iii) high speed (124 MHz on Xilinx Virtex-II) single-FPGA implementation of the processor with regularity driven efficient place and route. Companion to the processor is the very fast SimPLE compiler which achieves compilation rates of 4 million gates/hour. In order to simulate the netlist, the compiled instructions are streamed through the FPGA, along with the simulation vectors. This architecture plugs in naturally into any existing HDL simulation environment. We have a working prototype based on a commercially available PCI-based FPGA board. Srihari Cadambi, Chandra Mulpuri, Pranav Ashar |
DAC | 1 |
| 2001 | Static Profile-Driven Compilation for FPGAs
Srihari Cadambi, Seth Copen Goldstein |
FPL | 1 |
| 2000 | Efficient Place and Route for Pipeline Reconfigurable ArchitecturesabstractIn this paper, we present a fast and efficient compilation methodology for pipeline reconfigurable architectures. Our compiler back-end is much faster than conventional CAD tools, and fairly efficient. We represent pipeline reconfigurable architectures by a generalized VLIW-like model. The complex architectural constraints are effectively expressed in terms of a single graph parameter: the routing path length (RPL). Compiling to our model using RPL, we demonstrate fast compilation times and show speedups of between 10x and 200x on a pipeline reconfigurable architecture when compared to an UltraSparc-II. Srihari Cadambi, Seth Copen Goldstein |
ICCD | 1 |
| 1999 | CPR: A Configuration Profiling ToolabstractIn this paper we describe a Configuration PRofiling tool (CPR) and show how it can be used to aid compiler designers, FPGA architects and in the construction of macro-generator libraries. CPR uses subgraph matching to identify the parts of an application which are most important to achieve high performance. Using CPR as a guide we implemented a few macros for a macro-generator library, which yielded significant improvement in both the quality of configurations and speed of compilation. Srihari Cadambi, Seth Copen Goldstein |
FCCM | 1 |
| 1999 | PipeRench: A Coprocessor for Streaming multimedia AccelerationabstractFuture computing workloads will emphasize an architecture's ability to perform relatively simple calculations on massive quantities of mixed-width data. This paper describes a novel reconfigurable fabric architecture, PipeRench, optimized to accelerate these types of computations. PipeRench enables fast, robust compilers, supports forward compatibility, and virtualizes configurations, thus removing the fixed size constraint present in other fabrics. For the first time we explore how the bit-width of processing elements affects performance and show how the PipeRench architecture has been optimized to balance the needs of the compiler against the realities of silicon. Finally, we demonstrate extreme performance speedup on certain computing kernels (up to 190x versus a modern RISC processor), and analyze how this acceleration translates to application speedup. Seth Copen Goldstein, Herman Schmit, Matthew Moe, Mihai Budiu, Srihari Cadambi, R. Reed Taylor, Ronald Laufer |
ISCA | 5 |
| 1998 | Managing Pipeline-Reconfigurable FPGAsabstractWhile reconfigurable computing promises to deliver incomparable performance, it is still a marginal technology due to the high cost of developing and upgrading applications. Hardware virtualization can be used to significantly reduce both these costs. In this paper we describe the benefits of hardware virtualization, and show how it can be achieved using a combination of pipeline reconfiguration and run-time scheduling of both configuration streams and data streams. The result is PipeRench, an architecture that supports robust compilation and provides forward compatibility. Our preliminary performance analysis predicts that PipeRench will outperform commercial FPGAs and DSPs in both overall performance and in performance per mm2. Srihari Cadambi, Jeffrey Weener, Seth Copen Goldstein, Herman Schmit, Donald E. Thomas |
FPGA | 1 |