Nitin Chandrachoodan

dblp:55/3765 · DBLP profile ↗
← Back
27ranked-venue papers
1as first author
8since 2021 · last 2023
0000-0002-9258-7317ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 7 since 2021Security and privacy · 2 · 1 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Work-in-Progress: QRCNN: Scalable CNNs
abstract
Dropping the features/kernels in the convolutional layer of convolutional neural networks is a popular variant of structured pruning to reduce the computational load, but this comes at the cost of retraining and performance loss. In this work, we propose the QRCNN framework that shows graceful degradation of performance while dropping features in the convolutional layers without needing retraining. The framework allows trimming the network at inference time, thereby scaling with available computational power. The proposed method can achieve a reduction in the number of MAC computations of 1.22 -- 2.28X with a median of 1.575X. The speedup is measured on three compute platforms namely 8GB RAM Raspberry Pi 4B (embedded platform), octa-core Intel i7 processor with 64GB RAM, and NVidia Quadro K2200 (GPU) with 4GB memory.
Dara Nagaraju, Nitin Chandrachoodan
CASES2
2023 ViTA: A Vision Transformer Inference Accelerator for Edge Applications
abstract
Vision Transformer models, such as ViT, Swin Transformer, and Transformer-in-Transformer, have recently gained significant traction in computer vision tasks due to their ability to capture the global relation between features which leads to superior performance. However, they are compute-heavy and difficult to deploy in resource-constrained edge devices. Existing hardware accelerators, including those for the closely-related BERT transformer models, do not target highly resource-constrained environments. In this paper, we address this gap and propose ViTA - a configurable hardware accelerator for inference of vision transformer models, targeting resource-constrained edge computing devices and avoiding repeated off-chip memory accesses. We employ a head-level pipeline and inter-layer MLP optimizations, and can support several commonly used vision transformer models with changes solely in our control logic. We achieve nearly 90% hardware utilization efficiency on most vision transformer models, report a power of 0.88W when synthesised with a clock of 150 MHz, and get reasonable frame rates - all of which makes ViTA suitable for edge applications.
Shashank Nag, Gourav Datta, Souvik Kundu 0002, Nitin Chandrachoodan, Peter A. Beerel
ISCAS4
2023 Snoopy: A Webpage Fingerprinting Framework With Finite Query Model for Mass-Surveillance
abstract
Internet users are vulnerable to privacy attacks despite the use of encryption. Webpage fingerprinting, an attack that analyzes encrypted traffic, can identify the webpages visited by a user. The key challenges in performing mass-scale webpage fingerprinting arise from (i) the sheer number of combinations of user behavior and preferences to account for, and; (ii) the bound on the number of website queries imposed by the defense mechanisms (e.g., DDoS defense) deployed at the website. These constraints preclude the use of conventional data-intensive ML-based techniques. In this work, we propose Snoopy, a first-of-its-kind framework, that performs webpage fingerprinting for a large number of users visiting a website. Snoopy caters to the generalization requirements of mass-surveillance while complying with a bound on the number of website accesses (finite query model) for traffic sample collection. We show that Snoopy achieves$\approx 90\%$accuracy when evaluated on most websites, across various browsing contexts. A simple ensemble of Snoopy and an ML-based technique achieves$\approx 97\%$accuracy while adhering to the finite query model, in cases when Snoopy alone does not perform well.
Gargi Mitra, Prasanna Karthik Vairam, Sandip Saha, Nitin Chandrachoodan, V. Kamakoti 0001
IEEE Trans. Dependable Secur. Comput.4
2022 Split-Knit Convolution: Enabling Dense Evaluation of Transpose and Dilated Convolutions on GPUs
abstract
Transpose convolutions occur in several image-based neural network applications, especially those involving segmentation or image generation. Unlike regular (forward) convolutions, they result in data access and computation patterns that are less regular, and generally have poorer performance when implemented in software. We present split-knit convolution (SKConv) – a technique to replace transpose convolutions with multiple regular convolutions followed by interleaving. We show how existing software frameworks for GPU implementation of deep neural networks can be adapted to realize this computation, and compare against the standard techniques used by such frameworks.
Arjun Menon Vadakkeveedu, Debabrata Mandal, Pradeep Ramachandran, Nitin Chandrachoodan
HIPC4
2022 Layerwise Disaggregated Evaluation of Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) have attracted considerable attention due to their suitability to processing temporal input streams, as well as the emergence of highly power-efficient neuromorphic hardware platforms. The computational cost of evaluating a Spiking Neural Network (SNN) is strongly correlated with the number of timesteps for which it is evaluated. To improve the computational efficiency of SNN evaluation, we propose layerwise disaggregated SNNs (LD-SNNs), wherein the number of timesteps is independently optimized for each layer of the network. In effect, LD-SNNs allow for a better allocation of computational effort across layers in a network, resulting in an improved tradeoff between accuracy and efficiency. We propose a methodology to design optimized LD-SNNs from any given SNN. Across four benchmark networks, LD-SNNs achieve 1.67-3.84x reduction in synaptic updates and 1.2-2.56x reduction in neurons evaluated. These improvements translate to 1.25-3.45x faster inference on four different hardware platforms including two server-class platforms, a desktop platform and an edge SoC.
Abinand Nallathambi, Sanchari Sen, Anand Raghunathan, Nitin Chandrachoodan
ISLPED4
2022 Reduced Memory Viterbi Decoding for Hardware-accelerated Speech Recognition
abstract
Large Vocabulary Continuous Speech Recognition systems require Viterbi searching through a large state space to find the most probable sequence of phonemes that led to a given sound sample. This needs storing and updating of a large Active State List (ASL) in the on-chip memory (OCM) at regular intervals (called frames), which poses a major performance bottleneck for speech decoding. Most works use hash tables for OCM storage while beam-width pruning to restrict the ASL size. To achieve a decent accuracy and performance, a large OCM, numerous acoustic probability computations, and DRAM accesses are incurred. We propose to use a binary search tree for ASL storage and a max heap data structure to track the worst cost state and efficiently replace it when a better state is found. With this approach, the ASL size can be reduced from over 32K to 512 with minimal impact on recognition accuracy for a 7,000-word vocabulary model. This, combined with a caching technique for acoustic scores, reduced the DRAM data accessed by 31 \( \times \) and the acoustic probability computations by 26 \( \times \) . The approach has also been implemented in hardware on a Xilinx Zynq FPGA at 200 MHz using the Vivado SDS compiler. We study the tradeoffs among the amount of OCM used, word error rate, and decoding speed to show the effectiveness of the approach. The resulting implementation is capable of running faster than real time with 91% lesser block-RAMs.
Pani Prithvi Raj, Akhil Reddy Pakala, Nitin Chandrachoodan
ACM Trans. Embed. Comput. Syst.3
2021 A Smoothed LASSO-Based DNN Sparsification Technique
abstract
Deep Neural Networks (DNNs) are increasingly being used in a variety of applications. However, DNNs have huge computational and memory requirements. One way to reduce these requirements is to sparsify DNNs by using smoothed LASSO (Least Absolute Shrinkage and Selection Operator) functions. In this paper, we show that irrespective of error profile, the sparsity values obtained using various smoothed LASSO functions are similar, provided the maximum error of these functions with respect to the LASSO function is the same. We also propose a layer-wise DNN pruning algorithm, where the layers are pruned based on their individual allocated accuracy loss budget, determined by estimates of the reduction in number of multiply-accumulate operations (in convolutional layers) and weights (in fully connected layers). Further, the structured LASSO variants in both convolutional and fully connected layers are explored within the smoothed LASSO framework and the tradeoffs involved are discussed. The efficacy of proposed algorithm in enhancing the sparsity within the allowed degradation in DNN accuracy and results obtained on structured LASSO variants are shown on MNIST, SVHN, CIFAR-10, and Imagenette datasets and on larger networks such as ResNet-50 and Mobilenet.
Basava Naga Girish Koneru, Nitin Chandrachoodan, Vinita Vasudevan
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 Optimization of Signal Processing Applications Using Parameterized Error Models for Approximate Adders
abstract
Approximate circuit design has gained significance in recent years targeting error-tolerant applications. In the literature, there have been several attempts at optimizing the number of approximate bits of each approximate adder in a system for a given accuracy constraint. For computational efficiency, the error models used in these routines are simple expressions obtained using regression or by assuming inputs or the error is uniformly distributed. In this article, we first demonstrate that for many approximate adders, these assumptions lead to an inaccurate prediction of error statistics for multi-level circuits. We show that mean error and mean square error can be computed accurately if static probabilities of adders at all stages are taken into account. Therefore, in a system with a certain type of approximate adder, any optimization framework needs to take into account not just the functionality of the adder but also its position in the circuit, functionality of its parents, and the number of approximate bits in the parent blocks. We propose a method to derive parameterized error models for various types of approximate adders. We incorporate these models within an optimization framework and demonstrate that the noise power is computed accurately.
Celia Dharmaraj, Vinita Vasudevan, Nitin Chandrachoodan
ACM Trans. Embed. Comput. Syst.3
2020 Depending on HTTP/2 for Privacy? Good Luck!
abstract
HTTP/2 introduced multi-threaded server operation for performance improvement over HTTP/1.1. Recent works have discovered that multi-threaded operation results in multiplexed object transmission, that can also have an unanticipated positive effect on TLS/SSL privacy. In fact, these works go on to design privacy schemes that rely heavily on multiplexing to obfuscate the sizes of the objects based on which the attackers inferred sensitive information. Orthogonal to these works, we examine if the privacy offered by such schemes work in practice. In this work, we show that it is possible for a network adversary with modest capabilities to completely break the privacy offered by the schemes that leverage HTTP/2 multiplexing. Our adversary works based on the following intuition: restricting only one HTTP/2 object to be in the server queue at any point of time will eliminate multiplexing of that object and any privacy benefit thereof. In our scheme, we begin by studying if (1) packet delays, (2) network jitter, (3) bandwidth limitation, and (4) targeted packet drops have an impact on the number of HTTP/2 objects processed by the server at an instant of time. Based on these insights, we design our adversary that forces the server to serialize object transmissions, thereby completing the attack. Our adversary was able to break the privacy of a real-world HTTP/2 website 90% of the time, the code for which will be released. To the best of our knowledge, this is the first privacy attack on HTTP/2.
Gargi Mitra, Prasanna Karthik Vairam, Patanjali SLPSK, Nitin Chandrachoodan, V. Kamakoti 0001
DSN4
2020 EASpiNN: Effective Automated Spiking Neural Network Evaluation on FPGA
abstract
Neural networks (NNs) have been widely used in many machine learning algorithms and have been deployed for various industrial applications like image classification, speech recognition, and automated control. Spiking neural network (SNN), known as the third-generation neural network, incorporates timing information in the network and is more biologically plausible [1]. Compared to today's artificial and convolutional neural networks (ANN and CNN) where all neurons in each layer will always be activated and computed, SNN only activates those neurons whose membrane potential exceed the threshold potential [2]. As a result, SNN requires fewer computation resources and less data communication between network layers due to its event-driven nature. Although SNN has been blamed for the relatively lower accuracy, recent studies on converted SNNs have improved its accuracy to a similar level of ANN and CNN for smaller network models like MNIST and CIFAR-10, and have demonstrated the great potential of SNN in future deep learning systems [2].
Sathish Panchapakesan, Zhenman Fang, Nitin Chandrachoodan
FCCM3
2020 Energy Reduction in Turbo Decoding through Dynamically Varying Bit-Widths
abstract
We investigate the impact of dynamically changing the numerical bit-width across iterations in a BCJR component decoder of a turbo decoder. We show that by performing initial iterations with a larger number of bits but thereafter reducing the number of bits, it is possible to have minimal impact on decoding accuracy. At the same time, this reduction in bit-width can be exploited through appropriate hardware changes to consume less power in the later iterations. We propose a new state encoding mechanism as well as an organization of memory blocks that enables this power reduction, and quantify the effects. When combined with other schemes for early termination, the overall energy consumed per decoding operation can be reduced by between 10-20%.
Sundarrajan Rangachari, Nitin Chandrachoodan
ISCAS2
2019 Data Subsetting: A Data-Centric Approach to Approximate Computing
abstract
Approximate Computing (AxC), which leverages the intrinsic resilience of applications to approximations in their underlying computations, has emerged as a promising approach to improving computing system efficiency. Most prior efforts in AxC take a compute-centric approach and approximate arithmetic or other compute operations through design techniques at different levels of abstraction. However, emerging workloads such as machine learning, search and data analytics process large amounts of data and are significantly limited by the memory sub-systems of modern computing platforms.In this work, we shift the focus of approximations from computations to data, and propose a data-centric approach to AxC, which can boost the performance of memory-subsystem-limited applications. The key idea is to modulate the application's data-accesses in a manner that reduces off-chip memory traffic. Specifically, we propose a data-access approximation technique called data subsetting, in which all accesses to a data structure are redirected to a subset of its elements so that the overall footprint of memory accesses is decreased. We realize data subsetting in a manner that is transparent to hardware and requires only minimal changes to application software. Recognizing that most applications of interest represent and process data as multi-dimensional arrays or tensors, we develop a templated data structure called SubsettableTensor that embodies mechanisms to define the accessible subset and to suitably redirect accesses to elements outside the subset. As a further optimization, we observe that data subsetting may cause some computations to become redundant and propose a mechanism for application software to identify and eliminate such computations. We implement SubsettableTensor as a C++ class and evaluate it using parallel software implementations of 7 machine learning applications on a 48-core AMD Opteron server. Our experiments indicate that data subsetting enables 1.33×-4.44× performance improvement with <;0.5% loss in application-level quality, underscoring its promise as a new approach to approximate computing.
Swagath Venkataramani, Nitin Chandrachoodan, Anand Raghunathan
DATE3
2019 A Simulation-Based Metric to Guide Glitch Power Reduction in Digital Circuits
abstract
In this paper, we propose an algorithm to classify spurious transitions in the activity of a digital circuit as generated and propagated glitches during logic simulation. Using the activities obtained, we compute a criticality metric to identify the nets where glitch minimization techniques are likely to provide the maximum benefit. The proposed metric provides insight into which techniques are best suited for use in glitch reduction for a given circuit. This enables targeted application of glitch reduction techniques. Experiments with several glitch intensive benchmarks show a faster convergence within fewer iterations to solutions with reduced glitch activity. We validate this observation by using the proposed metric to guide the application of some glitch reduction techniques and quantify the resultant savings. The proposed algorithm can be seamlessly incorporated in modern event-driven logic simulators.
Shivani Bathla, Rahul M. Rao, Nitin Chandrachoodan
IEEE Trans. Very Large Scale Integr. Syst.3
2018 Optimizing power-accuracy trade-off in approximate adders
abstract
Approximate circuit design has gained significance in recent years targeting applications like media processing where full accuracy is not required. In this paper, we propose an approximate adder in which the approximate part of the sum is obtained by finding a single optimal level that minimises the mean error distance. Therefore hardware needed for the approximate part computation can be removed, which effectively results in very low power consumption. We compare the proposed adder with various approximate adders in the literature in terms of power and accuracy metrics. The power savings of our adder is shown to be 17% to 55% more than power savings of the existing approximate adders over a significant range of accuracy values. Further, in an image addition application, this adder is shown to provide the best trade-off between PSNR and power.
D. Celia, Vinita Vasudevan, Nitin Chandrachoodan
DATE3
2018 Lossless Parallel Implementation of a Turbo Decoder on GPU
abstract
Turbo decoders use the recursive BCJR algorithm which is computationally intensive and hard to parallelise. The branch metric and extrinsic log-likelihood ratio computations are easily parallelisable, but the forward and backward metric computation is not parallelisable without compromising bit error rate. This paper proposes a lossless parallelisation technique for Turbo decoders on Graphics Processing Units (GPU). The recursive forward and backward metric computation is formulated as prefix (scan) matrix multiplication problem which is computed on the GPU using parallel prefix sum computation technique. Overall, this method achieves a throughput of 73 Mbps for a 3GPP LTE compliant turbo decoder without any BER loss and latency as low as 61 μs.
Karthikeyan Natarajan, Nitin Chandrachoodan
HiPC2
2018 Probabilistic Error Modeling for Two-part Segmented Approximate Adders
abstract
Approximate adders are used in applications that are error tolerant to save on power and area. We consider the class of two-part segmented approximate adders, where the upper part of the sum is computed accurately and the lower part of the sum is approximated. In this paper, we model the error of various two-part segmented approximate adders using probabilistic analysis and derive expressions for some basic error metrics used in literature. We compare the results obtained using our expressions for various error metrics with those using Monte Carlo simulations for different input distributions. Further, in an image addition application, we use our expression derived for mean square error and show that it predicts the PSNR correctly.
D. Celia, Vinita Vasudevan, Nitin Chandrachoodan
ISCAS3
2017 FPGA Implementation of Non-Uniform DFT for Accelerating Wireless Channel Simulations (Abstract Only)
Srinivas Siripurapu, Aman Gayasen, Padmini Gopalakrishnan, Nitin Chandrachoodan
FPGA4
2017 Scenario-Aware Dynamic Power Reduction Using Bias Addition
abstract
Typical modern communication systems operate over a wide dynamic range of signal strengths. We consider the approach of adding a bias as offset to reduce switching activity, and study the average bit toggle for signals with different distributions, as a function of signal span and the actual bias value added. From the analysis, we provide guidelines to choose an optimal bias value, based on the system scenario, to obtain the lowest power consumption. Simulation results confirm the accuracy of the theoretical analysis. Various finite-impulse response filter architectures are evaluated, and we propose suitable enhancements to them to enable improved power savings. We apply the bias addition technique and the proposed architectural enhancements to a wireless local area network digital receiver chain, and demonstrate that over 25% power savings can be achieved under different signal conditions.
Sundarrajan Rangachari, Jaiganesh Balakrishnan, Nitin Chandrachoodan
IEEE Trans. Very Large Scale Integr. Syst.3
2015 iitRACE: A Memory Efficient Engine for Fast Incremental Timing Analysis and Clock Pessimism Removal
abstract
We describe a timing analysis engine for efficient processing of incremental changes to a circuit. The engine uses a block-based approach for incremental slack propagation. Logic cones affected by incremental changes to the design are identified and used to restrict the scope of the computation. Incremental block-based clock-pessimism removal and reporting of worst paths in the circuit is implemented using a novel dynamic path reduction technique. The engine is very efficient in memory usage compared to other known academic timers while maintaining a very high accuracy of reported path slacks when compared to a standard industrial timing engine. Certain paths are intentionally omitted from reporting in order to save on runtime, while ensuring that all paths with highest criticality are covered. Our timer (iitRACE) placed overall third in TAU 2015 contest on incremental timing analysis. Experimental results on industrial benchmarks from TAU 2015 contest have justified that iitRACE has average memory requirement 2X and 30X lower than that of first and second place timers respectively.
Chaitanya Peddawad, Aman Goel, B. Dheeraj, Nitin Chandrachoodan
ICCAD4
2015 DFT Assisted Techniques for Peak Launch-to-Capture Power Reduction during Launch-On-Shift At-Speed Testing
abstract
Scan-based testing is crucial to ensuring correct functioning of chips. In this scheme, the scan and capture phases are interleaved. It is well known that for large designs, excessive switching activity during the launch-to-capture window leads to high voltage droop on the power grid, ultimately resulting in false delay failures during at-speed test. This article proposes a new design-for-testability (DFT) scheme for launch-on-shift (LOS) testing, which ensures that the combinational logic remains undisturbed between the interleaved capture phases, providing computer-aided-design (CAD) tools with extra search space for minimizing launch-to-capture switching activity through test pattern ordering (TPO). We further propose a new TPO algorithm that keeps track of the don't cares during the ordering process, so that the don't care filling step after the ordering process yields a better reduction in launch-to-capture switching activity compared to any other technique in the literature. The proposed DFT-assisted technique, when applied to circuits in ITC99 benchmark suite, produces an average reduction of 17.68% in peak launch-to-capture switching activity (CSA) compared to the best known lowpower TPO technique. Even for circuits whose test cubes are not rich in don't care bits, the proposed technique produces an average reduction of 15% in peak CSA, while for the circuits with test cubes rich in don't care bits (≥75%), the average reduction is 24%. The proposed technique also reduces the average power dissipation (considering both scan cells and combinational logic) during the scan phase by about 43.5% on an average, compared to the adjacent filling technique.
Seetal Potluri, Satya Trinadh, Sobhan Babu Chintapalli, V. Kamakoti 0001, Nitin Chandrachoodan
ACM Trans. Design Autom. Electr. Syst.5
2013 Speeding up computation of the max/min of a set of gaussians for statistical timing analysis and optimization
abstract
Statistical static timing analysis (SSTA) involves computation of maximum (max) and minimum (min) of Gaussian random variables. Typically, the max or min of a set of Gaussians is performed iteratively in a pair-wise fashion, wherein the result of each pair-wise max or min operation is approximated to a Gaussian by matching moments of the true result obtained using Clark's approach [1]. The approximation error in the final result is thus a function of the order in which the pair-wise operations are performed.
Vimitha A. Kuruvilla, Debjit Sinha, Jeff Piaget, Chandu Visweswariah, Nitin Chandrachoodan
DAC5
2013 PinPoint: An algorithm for enhancing diagnostic resolution using capture cycle power information
abstract
Conventional ATPG tools help in detecting only the equivalence class to which a fault belongs and not the fault itself. This paper presents PinPoint, a technique that further divides the equivalence class into smaller sets based on the capture power consumed by the circuit under test in the presence of different faults in it, thus aiding in narrowing down on the fault. Applying the technique on ITC benchmark circuits yielded significant improvement in diagnostic resolution.
Seetal Potluri, Satya Trinadh, Roopashree Baskaran, Nitin Chandrachoodan, V. Kamakoti 0001
ETS4
2013 Scalable low power digital filter architectures for varying input dynamic range
abstract
Architecture level optimizations for reducing power consumption in commonly used digital filter implementations are studied. We look at two optimizations (data shifting and offset addition) when the operating condition is very different from the worst case scenarios for which the filters are designed, and show how to use them to exploit this behavior. We study the reduction in switching activity for different input conditions. The power consumption of the proposed architectures are compared against reference designs using prelayout synthesis netlist with a 45 nm CMOS library and the results are tabulated.
Sundarrajan Rangachari, Nitin Chandrachoodan
ISCAS2
2012 FPGA-Based High-Performance and Scalable Block LU Decomposition Architecture
abstract
Decomposition of a matrix into lower and upper triangular matrices (LU decomposition) is a vital part of many scientific and engineering applications, and the block LU decomposition algorithm is an approach well suited to parallel hardware implementation. This paper presents an approach to speed up implementation of the block LU decomposition algorithm using FPGA hardware. Unlike most previous approaches reported in the literature, the approach does not assume the matrix can be stored entirely on chip. The memory accesses are studied for various FPGA configurations, and a schedule of operations for scaling well is shown. The design has been synthesized for FPGA targets and can be easily retargeted. The design outperforms previous hardware implementations, as well as tuned software implementations including the ATLAS and MKL libraries on workstations.
Manish Kumar Jaiswal, Nitin Chandrachoodan
IEEE Trans. Computers2
2007 A Novel Event Based Simulation Algorithm for Sequential Digital Circuit Simulation
abstract
An algorithm and architecture for a hardware based simulation accelerator is presented. The accelerator can perform full timing simulation of synchronous digital circuits described at the gate level. The simulator makes use of a cycle based processor core in conjunction with event queues to execute the simulation. By ensuring that the gates are evaluated in rank order, the problem of sorting event queues is avoided. Static scheduling of the gates at compile time allows a simple control structure for the run time. The simple architecture allows large processing arrays to be implemented on relatively simple hardware. Results on ISCAS89 benchmark circuits are presented to demonstrate the scalability with hardware resources.
Karthick Parashar, Nitin Chandrachoodan
FPL2
2003 Algorithm and VLSI architecture for high performance adaptive video scaling
abstract
We propose an efficient high-performance scaling algorithm based on the oriented polynomial image model. We develop a simple classification scheme that classifies the region around a pixel as an oriented or nonoriented block. Based on this classification, a nonlinear oriented interpolation is performed to obtain high quality video scaling. In addition, we also propose a generalization that can perform scaling for arbitrary scaling factors. Based on this algorithm, we develop an efficient architecture for image scaling. Specifically, we consider an architecture for scaling a Quarter Common Intermediate Format (QCIF) image to 4CIF format. We show the feasibility of the architecture by describing the various computation units in a hardware description language (Verilog) and synthesizing the design into a netlist of gates. The synthesis results show that an application specific integrated circuit (ASIC) design which meets the throughput requirements can be built with a reasonable silicon area.
Arun Raghupathy, Nitin Chandrachoodan, K. J. Ray Liu
IEEE Trans. Multim.2
2001 An efficient timing model for hardware implementation of multirate dataflow graphs
abstract
We consider the problem of representing timing information associated with functions in a dataflow graph used to represent a signal processing system in the context of high-level hardware (architectural) synthesis. This information is used for synthesis of appropriate architectures for implementing the graph. Conventional models for timing suffer from shortcomings that make it difficult to represent timing information in a hierarchical manner, especially for multirate signal processing systems. We identify some of these shortcomings, and provide an alternate model that does not have these problems. We show that with some reasonable assumptions on the way hardware implementations of multirate systems operate, we can derive general hierarchical descriptions of multirate systems similarly to single rate systems. Several analytical results such as the computation of the iteration period bound, that previously applied only to single rate systems can also easily be extended to multirate systems under the new assumptions. We have applied our model to several multirate signal processing applications, and obtained favorable results. We present results of the timing information computed for several multirate DSP applications that show how the new treatment can streamline the problem of performance analysis and synthesis of such systems.
Nitin Chandrachoodan, Shuvra S. Bhattacharyya, K. J. Ray Liu
ICASSP1