Niraj K. Jha

dblp:j/NirajKJha · DBLP profile ↗
← Back
335ranked-venue papers
25as first author
27since 2021 · last 2026
0000-0002-1539-0369ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 308 · 24 first-author · 17 since 2021Software engineering, systems software and programming languages · 18Artificial intelligence and machine learning · 9 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 since 2021Computer networks · 7 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1Theory of computation · 1
YearPublicationVenuePosition
2026 COMFORT: A Continual Fine-Tuning Framework for Foundation Models Targeted at Consumer Healthcare
abstract
Wearable medical sensors (WMSs) are revolutionizing smart healthcare by enabling continuous, real-time monitoring of user physiological signals, especially in the field of consumer healthcare. The integration of WMSs and modern machine learning enables unprecedented solutions to efficient early-stage disease detection. Despite the success of Transformers in various fields, their application to sensitive domains, such as smart healthcare, remains underexplored due to limited data accessibility and privacy concerns. A key challenge in disease detection using Transformers with WMS data is the difficulty of obtaining sufficient labeled training data from individuals suffering from a specific disease, as this requires labeling by experts and extensive time and effort devoted to data collection and data privacy regulation compliance. The phrase “foundation models” conjures up an image of a large language model trained on text corpora. However, foundation models need not be large, nor necessarily trained on text. In this work, we propose a foundation model trained on WMS data under a continual fine-tuning framework called COMFORT. COMFORT introduces a novel approach for pre-training a Transformer-based foundation model on a large dataset of physiological signals exclusively collected from healthy individuals with commercially available WMSs. We adopt a masked data modeling (MDM) objective to pre-train this health foundation model. MDM is inspired by the mask language modeling approach employed in self-supervised training of language models. Our WMS-based foundation model can then be rapidly adapted to multiple disease detection tasks through various parameter-efficient fine-tuning (PEFT) methods, such as Low-Rank Adaptation and its variants. COMFORT continually stores the low-rank decomposition matrices and classifiers obtained using the PEFT algorithms to construct a library for multi-disease detection. This library enables scalable and memory-efficient disease detection on edge devices. Our experimental results demonstrate that COMFORT achieves highly competitive performance while reducing memory overhead by up to 52% relative to conventional methods. Thus, COMFORT paves the way for personalized and proactive solutions to efficient and effective early-stage disease detection.
Chia-Hao Li, Niraj K. Jha
ACM Trans. Embed. Comput. Syst.2
2026 Learning Interpretable Differentiable Logic Networks for Tabular Regression
abstract
Neural networks (NNs) achieve outstanding performance in many domains; however, their decision processes are often opaque and their inference can be computationally expensive in resource-constrained environments. We recently proposed Differentiable Logic Networks (DLNs) to address these issues for tabular classification based on relaxing discrete logic into a differentiable form, thereby enabling gradient-based learning of networks built from binary logic operations. DLNs offer interpretable reasoning and substantially lower inference cost. We extend the DLN framework to supervised tabular regression. We first redesign the final output layer (the SumLayer) to support continuous targets. More critically, we find the original two-phase training procedure used for classification is suboptimal for regression, and thus develop a unified, single-stage optimization procedure. We also demonstrate that temperature annealing of the network’s differentiable relaxations is decisive for achieving stable convergence and high accuracy. We evaluate the resulting model on 15 public regression benchmarks, comparing it with modern neural networks and classical regression baselines. Regression DLNs match or exceed baseline accuracy while preserving interpretability and fast inference. Our results show that DLNs are a viable, cost-effective alternative for regression tasks, especially where model transparency and computational efficiency are important.
Chang Yue, Niraj K. Jha
ACM Trans. Design Autom. Electr. Syst.2
2025 LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
abstract
Text-to-video generation enhances content creation but is highly computationally intensive: The computational cost of Diffusion Transformers (DiTs) scales quadratically in the number of pixels. This makes minute-length video generation extremely expensive, limiting most existing models to generating videos of only 10-20 seconds length. We propose a Linear-complexity text-to-video Generation (Lin-Gen) framework whose cost scales linearly in the number of pixels. For the first time, LinGen enables high-resolution minute-length video generation on a single GPU without compromising quality. It replaces the computationally-dominant and quadratic-complexity block, self-attention, with a linear-complexity block called MATE, which consists of an MA-branch and a TE-branch. The MA-branch targets short-to-long-range correlations, combining a bidirectional Mamba2 block with our token rearrangement method, Rotary Major Scan, and our review tokens developed for long video generation. The TE-branch is a novel TEmporal Swin Attention block that focuses on temporal correlations between adjacent tokens and medium-range tokens. The MATE block addresses the adjacency preservation issue of Mamba and improves the consistency of generated videos significantly. Experimental results show that LinGen outperforms DiT (with a 75.6% win rate) in video quality with up to 15× (11.5×) FLOPs (latency) reduction. Furthermore, both automatic metrics and human evaluation demonstrate that our LinGen-4B yields comparable video quality to state-of-the-art models (with a 50.5%, 52.1%, 49.1% win rate with respect to Gen-3, LumaLabs, and Kling, respectively). This paves the way for hour-length movie generation and real-time interactive video generation. Project website: https://lineargen.github.io/.
Hongjie Wang 0002, Chih-Yao Ma, Yen-Cheng Liu, Ji Hou, Jialiang Wang 0001, Felix Juefei-Xu, Yaqiao Luo, Peizhao Zhang, Tingbo Hou, Peter Vajda, Niraj K. Jha, Xiaoliang Dai
CVPR12
2025 Generating, Fast and Slow: Scalable Parallel Video Generation with Video Interface Networks
Bhishma Dedhia, David Bourgin, Krishna Kumar Singh, Niraj K. Jha, Yuchen Liu 0002
ICCV7
2024 Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers
abstract
Deployment of Transformer models on edge devices is becoming increasingly challenging due to the exponentially growing inference cost that scales quadratically with the number of tokens in the input sequence. Token pruning is an emerging solution to address this challenge due to its ease of deployment on various Transformer backbones. However, most token pruning methods require computationally expensive fine-tuning, which is undesirable in many edge deployment cases. In this work, we propose Zero-TPrune, the first zero-shot method that considers both the importance and similarity of tokens in performing token pruning. It leverages the attention graph of pre-trained Transformer models to produce an importance distribution for tokens via our proposed Weighted Page Rank (WPR) algorithm. This distribution further guides token partitioning for efficient similarity-based pruning. Due to the elimination of the fine-tuning overhead, Zero-TPrune can prune large models at negligible computational cost, switch between different pruning configurations at no computational cost, and perform hyperparameter tuning efficiently. We evaluate the performance of Zero-TPrune on vision tasks by applying it to various vision Transformer backbones and testing them on ImageNet. Without any fine-tuning, Zero-TPrune reduces the FLOPs cost of DeiT-S by 34.7% and improves its throughput by 45.3% with only 0.4% accuracy loss. Compared with state-of-the-art pruning methods that require fine-tuning, Zero-TPrune not only eliminates the need for fine-tuning after pruning but also does so with only 0.1% accuracy loss. Compared with state-of-the-art fine-tuning-free pruning methods, Zero-TPrune reduces accuracy loss by up to 49% with similar FLOPs budgets. Project webpage: https://jha-lab.github.io/zerotprune.
Hongjie Wang 0002, Bhishma Dedhia, Niraj K. Jha
CVPR3
2024 Attention-Driven Training-Free Efficiency Enhancement of Diffusion Models
abstract
Diffusion models (DMs) have exhibited superior performance in generating high-quality and diverse images. How-ever, this exceptional performance comes at the cost of expensive generation process, particularly due to the heavily used attention module in leading models. Existing works mainly adopt a retraining process to enhance DM efficiency. This is computationally expensive and not very scalable. To this end, we introduce the Attention-driven Training-free Efficient Diffusion Model (AT-EDM) framework that leverages attention maps to perform run-time pruning of redundant tokens, without the need for any retraining. Specifically, for single-denoising-step pruning, we develop a novel ranking algorithm, Generalized Weighted Page Rank (G-WPR), to identify redundant tokens, and a similarity-based recovery method to restore tokens for the convolution operation. In addition, we propose a Denoising-Steps-Aware Pruning (DSAP) approach to adjust the pruning budget across different denoising timesteps for better generation quality. Extensive evaluations show that AT-EDM performs favorably against prior art in terms of efficiency (e.g., 38.8% FLOPs saving and up to 1.53× speed-up over Stable Diffusion XL) while maintaining nearly the same FID and CLIP scores as the full model. Project webpage: https://atedm.github.io.
Hongjie Wang 0002, Difan Liu, Yijun Li 0001, Zhe Lin 0001, Niraj K. Jha, Yuchen Liu 0002
CVPR6
2024 DynaMo: Accelerating Language Model Inference with Dynamic Multi-Token Sampling
abstract
Shikhar Tuli, Chi-Heng Lin, Yen-Chang Hsu, Niraj Jha, Yilin Shen, Hongxia Jin. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Shikhar Tuli, Chi-Heng Lin, Yen-Chang Hsu, Niraj K. Jha, Yilin Shen, Hongxia Jin
NAACL-HLT4
2024 DOCTOR: A Multi-Disease Detection Continual Learning Framework Based on Wearable Medical Sensors
abstract
Modern advances in machine learning (ML) and wearable medical sensors (WMSs) in edge devices have enabled ML-driven disease detection for smart healthcare. Conventional ML-driven methods for disease detection rely on customizing individual models for each disease and its corresponding WMS data. However, such methods lack adaptability to distribution shifts and new task classification classes. In addition, they need to be rearchitected and retrained from scratch for each new disease. Moreover, installing multiple ML models in an edge device consumes excessive memory, drains the battery faster, and complicates the detection process. To address these challenges, we propose DOCTOR, a multi-disease detection continual learning (CL) framework based on WMSs. It employs a multi-headed deep neural network (DNN) and a replay-style CL algorithm. The CL algorithm enables the framework to continually learn new missions in which different data distributions, classification classes, and disease detection tasks are introduced sequentially. It counteracts catastrophic forgetting with either a data preservation (DP) method or a synthetic data generation (SDG) module. The DP method preserves the most informative subset of real training data from previous missions for exemplar replay. The SDG module models the probability distribution of the real training data and generates synthetic data for generative replay while retaining data privacy. The multi-headed DNN enables DOCTOR to detect multiple diseases simultaneously based on user WMS data. We demonstrate DOCTOR’s efficacy in maintaining high disease classification accuracy with a single DNN model in various CL experiments. In complex scenarios, DOCTOR achieves 1.43× better average test accuracy, 1.25× better F1-score, and 0.41 higher backward transfer than the naïve fine-tuning framework, with a small model size of less than 350 KB.
Chia-Hao Li, Niraj K. Jha
ACM Trans. Embed. Comput. Syst.2
2024 EdgeTran: Device-Aware Co-Search of Transformers for Efficient Inference on Mobile Edge Platforms
abstract
Automated design of efficient transformer models has recently attracted significant attention from industry and academia. However, most works only focus on certain metrics while searching for the best-performing transformer architecture. Furthermore, running traditional, complex, and large transformer models on low-compute edge platforms is a challenging problem. In this work, we propose a framework, called ProTran, to profile the hardware performance measures for a design space of transformer architectures and a diverse set of edge devices. We use this profiler in conjunction with the proposed co-search technique to obtain the best-performing models that have high accuracy on the given task and minimize latency, energy consumption, and peak power draw to enable edge deployment. We refer to our framework for co-optimizing accuracy and hardware performance measures as EdgeTran. It searches for the best transformer model and edge device pair. Finally, we propose GPTran, a multi-stage block-level grow-and-prune post-processing step that further improves accuracy in a hardware-aware manner. The obtained transformer model is 2.8× smaller and has a 0.8% higher GLUE score than the baseline (BERT-Base). Inference with it on the selected edge device enables 15.0% lower latency, 10.0× lower energy, and 10.8× lower peak power draw compared to an off-the-shelf GPU.
Shikhar Tuli, Niraj K. Jha
IEEE Trans. Mob. Comput.2
2023 Smart Healthcare
abstract
The Internet-of-Things (IoT) era promises hundreds of billions of devices or physical objects connected to the Internet. These objects include sensors, actuators, and processing elements that help us gather data, make intelligent decisions, and optimize processes. IoT is expected to have a potential economic impact of 3-6 trillion per year by 2025, with 1-2.5 trillion of this economic impact (its largest fraction) coming from smart healthcare applications. These healthcare applications will be enabled by various technologies, e.g., (i) neural network based detection and differential diagnosis using wearable medical sensors present in smartwatches and smartphones that will communicate with a health server to enable a physician to keep track of an individual's health, (ii) image analytics, powered by accurate convolutional neural networks, with applications to radiology, pathology, ophthalmology, dermatology, cardiology, gastroenterology, etc., and (iii) personalized medical decision-making. However, many challenges remain in making this vision a reality. In this talk, we will explore how machine learning models employed at different layers of the healthcare hierarchy can begin to realize the above vision.
Niraj K. Jha
ACM Great Lakes Symposium on VLSI1
2023 Im-Promptu: In-Context Composition from Image Prompts
abstract
Large language models are few-shot learners that can solve diverse tasks from a handful of demonstrations. This implicit understanding of tasks suggests that the attention mechanisms over word tokens may play a role in analogical reasoning. In this work, we investigate whether analogical reasoning can enable in-context composition over composable elements of visual stimuli. First, we introduce a suite of three benchmarks to test the generalization properties of a visual in-context learner. We formalize the notion of an analogy-based in-context learner and use it to design a meta-learning framework called Im-Promptu. Whereas the requisite token granularity for language is well established, the appropriate compositional granularity for enabling in-context generalization in visual stimuli is usually unspecified. To this end, we use Im-Promptu to train multiple agents with different levels of compositionality, including vector representations, patch representations, and object slots. Our experiments reveal tradeoffs between extrapolation abilities and the degree of compositionality, with non-compositional representations extending learned composition rules to unseen domains but performing poorly on combinatorial tasks. Patch-based representations require patches to contain entire objects for robust extrapolation. At the same time, object-centric tokenizers coupled with a cross-attention module generate consistent and high-fidelity solutions, with these inductive biases being particularly crucial for compositional generalization. Lastly, we demonstrate a use case of Im-Promptu as an intuitive programming interface for image generation.
Bhishma Dedhia, Michael Chang 0003, Jake Snell, Thomas L. Griffiths 0001, Niraj K. Jha
NeurIPS5
2023 SCouT: Synthetic Counterfactuals via Spatiotemporal Transformers for Actionable Healthcare
abstract
The synthetic control method has pioneered a class of powerful data-driven techniques to estimate the counterfactual reality of a unit from donor units. At its core, the technique involves a linear model fitted on the pre-intervention period that combines donor outcomes to yield the counterfactual. However, linearly combining spatial information at each time instance using time-agnostic weights fails to capture important inter-unit and intra-unit temporal contexts and complex nonlinear dynamics of real data. We instead propose an approach to use local spatiotemporal information before the onset of the intervention as a promising way to estimate the counterfactual sequence. To this end, we suggest a Transformer model that leverages particular positional embeddings, a modified decoder attention mask, and a novel pre-training task to perform spatiotemporal sequence-to-sequence modeling. Our experiments on synthetic data demonstrate the efficacy of our method in the typical small donor pool setting and its robustness against noise. We also generate actionable healthcare insights at the population and patient levels by simulating a state-wide public health policy to evaluate its effectiveness, an in silico trial for asthma medications to support randomized controlled trials, and a medical intervention for patients with Friedreich’s ataxia to improve clinical decision making and promote personalized therapy (code is available at https://github.com/JHA-Lab/scout ).
Bhishma Dedhia, Roshini Balasubramanian, Niraj K. Jha
ACM Trans. Comput. Heal.3
2023 FlexiBERT: Are Current Transformer Architectures too Homogeneous and Rigid?
abstract
The existence of a plethora of language models makes the problem of selecting the best one for a custom task challenging. Most state-of-the-art methods leverage transformer-based models (e.g., BERT) or their variants. However, training such models and exploring their hyperparameter space is computationally expensive. Prior work proposes several neural architecture search (NAS) methods that employ performance predictors (e.g., surrogate models) to address this issue; however, such works limit analysis to homogeneous models that use fixed dimensionality throughout the network. This leads to sub-optimal architectures. To address this limitation, we propose a suite of heterogeneous and flexible models, namely FlexiBERT, that have varied encoder layers with a diverse set of possible operations and different hidden dimensions. For better-posed surrogate modeling in this expanded design space, we propose a new graph-similarity-based embedding scheme. We also propose a novel NAS policy, called BOSHNAS, that leverages this new scheme, Bayesian modeling, and second-order optimization, to quickly train and use a neural surrogate model to converge to the optimal architecture. A comprehensive set of experiments shows that the proposed policy, when applied to the FlexiBERT design space, pushes the performance frontier upwards compared to traditional models. FlexiBERT-Mini, one of our proposed models, has 3% fewer parameters than BERT-Mini and achieves 8.9% higher GLUE score. A FlexiBERT model with equivalent performance as the best homogeneous model has 2.6× smaller size. FlexiBERT-Large, another proposed model, attains state-of-the-art results, outperforming the baseline models by at least 5.7% on the GLUE benchmark.
Shikhar Tuli, Bhishma Dedhia, Shreshth Tuli, Niraj K. Jha
J. Artif. Intell. Res.4
2023 TUTOR: Training Neural Networks Using Decision Rules as Model Priors
abstract
The human brain has the ability to carry out new tasks with limited experience. It utilizes prior learning experiences to adapt the solution strategy to new domains. On the other hand, deep neural networks (DNNs) generally need large amounts of data and computational resources for training. However, this requirement is not met in many settings. To address these challenges, we propose the TUTOR (training neural networks using decision rules as model priors) DNN synthesis framework. TUTOR targets tabular datasets. It synthesizes accurate DNN models with limited available data and reduced memory/computational requirements. It consists of three sequential steps. The first step involves generation, verification, and labeling of synthetic data. The synthetic data generation module targets both categorical and continuous features. TUTOR generates synthetic data from the same probability distribution as the real data. It then verifies the integrity of the generated synthetic data using a semantic integrity classifier module. It labels synthetic data based on a set of rules extracted from the real dataset. Next, TUTOR uses two training schemes that combine synthetic and training data to learn the DNN model parameters. These two schemes focus on two different ways in which synthetic data can be used to derive a prior on the model parameters and, hence, provide a better DNN initialization for training with real data. In the third step, TUTOR employs a grow-and-prune synthesis paradigm to learn both the weights and the architecture of the DNN to reduce model size while ensuring its accuracy. We evaluate the performance of TUTOR on nine datasets of various sizes. We show that in comparison to fully connected DNNs, TUTOR, on an average, reduces the need for data by$5.9\times $(geometric mean), improves accuracy by 3.4%, and reduces the number of parameters (floating-point operations) by$4.7\times $($4.3\times $) (geometric mean). Thus, TUTOR enables less data-hungry, more accurate, and more compact DNN synthesis.
Shayan Hassantabar, Prerit Terway, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 INFORM: Inverse Design Methodology for Constrained Multiobjective Optimization
abstract
Many system design methods use population-based optimization or a surrogate model for solving constrained multiobjective optimization (CMOO). When designing a system with multiple objectives and constraints, the designer may first be interested in understanding the tradeoffs among different objectives from a small number of simulations. In the next step, the designer may focus on specific regions of interest in the design space near a set of nondominated solutions to further improve performance on the targeted objectives. This may help make the search process sample-efficient. We propose inverse design methodology for constrained multiobjective optimization (INFORM): a two-step approach for sample-efficient CMOO of real-world nonlinear systems. In the first step, we modify a genetic algorithm (GA) to make the design process sample-efficient. We inject candidate solutions into the GA population using inverse design methods instead of determining the candidate solutions for the next generation using only crossover and mutation, as is done in standard GA. We present three types of inverse design techniques based on: 1) a neural network (NN) verifier; 2) NN; and 3) Gaussian mixture model. The candidate solutions for the next generation are thus a mix of those generated using crossover/mutation and solutions generated using inverse design. At the end of the first step, we obtain a set of nondominated solutions. In the second step, we choose the regions of interest around the nondominated solutions to further improve the objective function values using inverse design methods. We demonstrate the efficacy of INFORM through synthesis of nonlinear systems and analog circuits. The experimental results show that INFORM reduces synthesis time by up to$29\times $and improves the value of the objective function by up to 33% compared to a state-of-the-art baseline design methodology.
Prerit Terway, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 AccelTran: A Sparsity-Aware Accelerator for Dynamic Inference With Transformers
abstract
Self-attention-based transformer models have achieved tremendous success in the domain of natural language processing. Despite their efficacy, accelerating the transformer is challenging due to its quadratic computational complexity and large activation sizes. Existing transformer accelerators attempt to prune its tokens to reduce memory access, albeit with high compute overheads. Moreover, previous works directly operate on large matrices involved in the attention operation, which limits hardware utilization. In order to address these challenges, this work proposes a novel dynamic inference scheme, DynaTran, which prunes activations at runtime with low overhead, substantially reducing the number of ineffectual operations. This improves the throughput of transformer inference. We further propose tiling the matrices in transformer operations along with diverse dataflows to improve data reuse, thus, enabling higher energy efficiency. To effectively implement these methods, we propose AccelTran, a novel accelerator architecture for transformers. Extensive experiments with different models and benchmarks demonstrate that DynaTran achieves higher accuracy than the state-of-the-art top-$k$hardware-aware pruning strategy while attaining up to$1.2\times $higher sparsity. One of our proposed accelerators, AccelTran-Edge, achieves 330K$\times $higher throughput with 93K$\times $lower energy requirement when compared to a Raspberry Pi device. On the other hand, AccelTran-Server achieves$5.73\times $higher throughput and$3.69\times $lower energy consumption compared to the state-of-the-art transformer co-processor, Energon. The simulation source code is available athttps://github.com/jha-lab/acceltran.
Shikhar Tuli, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 TransCODE: Co-Design of Transformers and Accelerators for Efficient Training and Inference
abstract
Automated co-design of machine learning models and evaluation hardware is critical for efficiently deploying such models at scale. Despite the state-of-the-art performance of transformer models, they are not yet ready for execution on resource-constrained hardware platforms. High memory requirements and low parallelizability of the transformer architecture exacerbate this problem. Recently proposed accelerators attempt to optimize the throughput and energy consumption of transformer models. However, such works are either limited to a one-sided search of the model architecture or a restricted set of off-the-shelf devices. Furthermore, previous works only accelerate model inference and not training, which incurs substantially higher memory and compute resources, making the problem even more challenging. To address these limitations, this work proposes a dynamic training framework, called DynaProp, that speeds up the training process and reduces memory consumption. DynaProp is a low-overhead pruning method that prunes activations and gradients at runtime. To effectively execute this method on hardware for a diverse set of transformer architectures, we propose a flexible BERT accelerator, a framework that simulates transformer inference and training on a design space of accelerators. We use this simulator in conjunction with the proposed co-design technique, called TransCODE, to obtain the best-performing models with high accuracy on the given task and minimize latency, energy consumption, and chip area. The obtained transformer–accelerator pair achieves 0.3% higher accuracy than the state-of-the-art pair while incurring$5.2\times $lower latency and$3.0\times $lower energy consumption.
Shikhar Tuli, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 REPAIRS: Gaussian Mixture Model-based Completion and Optimization of Partially Specified Systems
abstract
Most system optimization techniques focus on finding the values of the system components to achieve the best performance. Searching over all component values gives the search methodology the freedom to explore the entire design space to determine the best system configuration. However, real-world systems often require searching in a restricted space over only a subset of component values while freezing some of the components to fixed values. Rather than optimizing from scratch to search over the subset of components, incorporating the past simulation logs (search performed when all components were allowed to vary) enables the optimization mechanism to utilize knowledge from past system behavior. In addition, when the system gives the same response over different combinations of input values, the designer may prefer one combination over another. Furthermore, real-world data often contain errors. To avoid catastrophic consequences of making decisions based on incorrect data points, we need a mechanism to identify and correct the resulting error. We propose REPAIRS, a methodology to complete/optimize partially specified systems. It also performs data integrity checks and identifies/corrects errors after detecting an anomaly in the data. We use a Gaussian mixture model to learn the joint distribution of the system inputs and the corresponding output response (objectives/constraints). We use the learned model to complete a partially specified system where only a subset of the component values and/or the system response is specified. When the system response exhibits multiple modes (e.g., same response for different combinations of input values), REPAIRS determines the combinations of input values that correspond to the several modes. Using past simulation logs, it searches over various subsets of system inputs to improve the performance of the reference solution. We also present a framework for verifying the integrity of a given data instance. When the integrity check fails, we provide a mechanism to identify the error location and correct the error. REPAIRS provides an explanation for the decision it makes for the different use cases described in this article. We provide results of REPAIRS in the context of completion, partial optimization, and data integrity check of real-world systems. REPAIRS achieves a hypervolume that is better than that obtained using a baseline method by up to 50%. It successfully identifies the error location and predicts the correct value of the erroneous feature with an error less than 0.2%. It detects error locations with a mean accuracy of up to 95% even when three feature values have an error.
Prerit Terway, Niraj K. Jha
ACM Trans. Embed. Comput. Syst.2
2023 CODEBench: A Neural Architecture and Hardware Accelerator Co-Design Framework
abstract
Recently, automated co-design of machine learning (ML) models and accelerator architectures has attracted significant attention from both the industry and academia. However, most co-design frameworks either explore a limited search space or employ suboptimal exploration techniques for simultaneous design decision investigations of the ML model and the accelerator. Furthermore, training the ML model and simulating the accelerator performance is computationally expensive. To address these limitations, this work proposes a novel neural architecture and hardware accelerator co-design framework, called CODEBench. It comprises two new benchmarking sub-frameworks, CNNBench and AccelBench, which explore expanded design spaces of convolutional neural networks (CNNs) and CNN accelerators. CNNBench leverages an advanced search technique, Bayesian Optimization using Second-order Gradients and Heteroscedastic Surrogate Model for Neural Architecture Search, to efficiently train a neural heteroscedastic surrogate model to converge to an optimal CNN architecture by employing second-order gradients. AccelBench performs cycle-accurate simulations for diverse accelerator architectures in a vast design space. With the proposed co-design method, called Bayesian Optimization using Second-order Gradients and Heteroscedastic Surrogate Model for Co-Design of CNNs and Accelerators, our best CNN–accelerator pair achieves 1.4% higher accuracy on the CIFAR-10 dataset compared to the state-of-the-art pair while enabling 59.1% lower latency and 60.8% lower energy consumption. On the ImageNet dataset, it achieves 3.7% higher Top1 accuracy at 43.8% lower latency and 11.2% lower energy consumption. CODEBench outperforms the state-of-the-art framework, i.e., Auto-NBA, by achieving 1.5% higher accuracy and 34.7× higher throughput while enabling 11.0× lower energy-delay product and 4.0× lower chip area on CIFAR-10.
Shikhar Tuli, Chia-Hao Li, Ritvik Sharma, Niraj K. Jha
ACM Trans. Embed. Comput. Syst.4
2022 CURIOUS: Efficient Neural Architecture Search Based on a Performance Predictor and Evolutionary Search
abstract
Neural networks (NNs) have been successfully deployed in various applications of artificial intelligence. However, architectural design of NNs is still a challenging problem. This is due to the need to navigate a search space based on a large number of hyperparameters. This forces the search space of possible architectures to grow exponentially. Using a trial-and-error design approach is very time consuming and leads to suboptimal architectures. In addition, approaches, such as neural architecture search based on reinforcement learning and differentiable gradient-based architecture search, often incur huge computational costs or significant memory requirements. To address these challenges, we propose the CURIOUS NN synthesis methodology. It uses a performance predictor to efficiently navigate the architectural search space with an evolutionary search process. The predictor is built using quasi Monte-Carlo sampling, boosted decision tree regression, and an intelligent iterative sampling method. It is designed to be sample efficient. CURIOUS starts from a base architecture and explores the architectural search space to obtain a variant of the base architecture with the highest performance. This search framework is general and covers all important NN architecture types, e.g., feedforward NNs (FFNNs), convolutional NNs (CNNs), recurrent NNs (RNNs), and transformers. We evaluate the performance of CURIOUS on various datasets and base architectures. Through these experiments, we demonstrate significant performance improvements over the baseline architectures. For the MNIST dataset, our CNN architecture achieves an error rate of 0.66%, with$8.6\times $fewer parameters compared to the LeNet-5 baseline. For the CIFAR-10 dataset, we use the ResNet architectures and residual networks with Shake-Shake regularization as the baselines. Our synthesized ResNet-18 has a 2.52% accuracy improvement over the original ResNet-18, 1.74% over ResNet-101, and 0.16% over ResNet-1001, while requiring comparable number of parameters and floating-point operations to the original ResNet-18. This result shows that instead of just increasing the number of layers to increase accuracy, an alternative is to use a better NN architecture with a small number of layers. In addition, CURIOUS achieves an error rate of just 2.69% with a variant of the residual architecture with Shake-Shake regularization. We also use the set of optimized hyperparameters found for ResNet-18 on the CIFAR-10 dataset to train and evaluate the model on the ImageNet dataset, and show 3.43% (1.83%) improvement in the top-1 (top-5) error rate compared to the original ResNet-18 model. CURIOUS also obtains the highest accuracy for various other FFNNs that are geared toward edge devices and IoT sensors. In addition, we use CURIOUS to search for deep RNN architectures for the SICK dataset for sentence similarity evaluation. It achieves a mean-squared error of only 0.2060, improving upon the base network performance, without the need to stack multiple long short-term memories. We also use CURIOUS to search for a better NN classifier for the sentiment analysis task on the Stanford sentiment treebank dataset using a pretrained BERT model and again demonstrate improvements in performance.
Shayan Hassantabar, Xiaoliang Dai, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 SCANN: Synthesis of Compact and Accurate Neural Networks
abstract
Deep neural networks (DNNs) have become the driving force behind recent artificial intelligence (AI) research. With the help of a vast amount of training data, neural networks can perform better than traditional machine learning algorithms in many applications. An important problem with implementing a neural network is the design of its architecture. Typically, such an architecture is obtained manually by exploring its hyperparameter space and kept fixed during training. This approach is both time consuming and inefficient. Another issue is that modern neural networks often contain millions of parameters, whereas many applications require small inference models due to imposed resource constraints, such as energy constraints on battery-operated devices. However, efforts to migrate DNNs to such devices typically entail a significant loss of classification accuracy. To address these challenges, we propose a two-step neural network synthesis methodology, called DR+SCANN, that combines two complementary approaches to design compact and accurate DNNs. At the core of our framework is the SCANN methodology that uses three basic architecture-changing operations, namely, connection growth, neuron growth, and connection pruning, to synthesize feedforward architectures with arbitrary structure. These neural networks are not limited to the multilayer perceptron structure. SCANN encapsulates three synthesis methodologies that apply a repeated grow-and-prune paradigm to three architectural starting points. DR+SCANN combines the SCANN methodology with dataset dimensionality reduction to alleviate the curse of dimensionality. We demonstrate the efficacy of SCANN and DR+SCANN on various image and nonimage datasets. We evaluate SCANN on MNIST, CIFAR-10, and ImageNet benchmarks. Without any loss in accuracy, SCANN generates a$46.3\times $smaller network than the LeNet-5 Caffe model. We also compare SCANN-synthesized networks with a state-of-the-art fully connected (FC) feedforward model for MNIST, and show$20\times $($19.9\times $) reduction in the number of parameters (floating-point operations) with little drop in accuracy. For the CIFAR-10 dataset, we target AlexNet and VGG-16 baseline architectures. SCANN reduces the number of parameters in AlexNet by$10.1\times $without any drop in accuracy. It reduces the number of parameters in the FC layers of VGG-16 to only 2.5k while increasing accuracy by 1.05%. On the ImageNet dataset, for the VGG-16 and MobileNetV2 architectures, we reduce network parameters by$8.0\times $and$1.3\times $, respectively, with a similar or improved performance over their respective baselines. We also evaluate the efficacy of using dimensionality reduction alongside SCANN (DR+SCANN) on nine small-to-medium-size datasets. Using this methodology enables us to reduce the number of connections in the network by up to$5078.7\times $(geometric mean:$82.1\times $), with little to no drop in accuracy. On seven out of nine datasets, we show 0.41%–10.09% accuracy improvements over the FC baseline models. We also show that our synthesis methodology yields neural networks that are much better at navigating the accuracy versus energy efficiency space. This can enable neural network-based inference even on Internet-of-Things sensors.
Shayan Hassantabar, Zeyu Wang 0004, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Fast Design Space Exploration of Nonlinear Systems: Part I
abstract
System design tools are often only available as input–output blackboxes: for a given design as input, they compute an output representing system behavior. Blackboxes are intended to be run in the forward direction. This article presents a new method of solving the “inverse design problem,” namely, given requirements or constraints on output, find an input that also optimizes an objective function. This problem is challenging for several reasons. First, blackboxes are not designed to be run in reverse. Second, inputs and outputs can be discrete and continuous. Third, finding designs concurrently satisfying a set of requirements is hard because designs satisfying individual requirements may conflict with each other. Fourth, blackbox evaluations can be expensive. Finally, evaluations can sometimes fail to produce an output due to nonconvergence of underlying numerical algorithms. This article presents CNMA, a new method of solving the inverse problem that overcomes these challenges. CNMA tries to sample only the part of the design space relevant to solving the inverse problem, leveraging the power of neural networks, mixed-integer linear programs, and a new learning-from-failure feedback loop. This article also presents a parallel version of CNMA that improves the efficiency and quality of solutions over the sequential version and tries to steer it away from local optima. CNMA’s performance is evaluated against conventional optimization methods for seven nonlinear design problems of 8 (two problems), 10, 15, 36 and 60 real-valued dimensions and one with 186 binary dimensions. Conventional methods evaluated are stable, off-the-shelf implementations of the Bayesian optimization with the Gaussian Processes, Nelder–Mead, and Random Search. The first two do not even produce a solution for problems that are high dimensional, having both discrete and continuous variables or whose blackboxes fail to return values for some inputs. CNMA produces solutions for all problems. When conventional methods do produce solutions, CNMA improves upon their performance by 1%–87%.
Sanjai Narain, Emily Mak, Dana Chee, Brendan J. Englot, Kishore Pochiraju, Niraj K. Jha, Karthik S. Narayan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2022 Fast Design Space Exploration of Nonlinear Systems: Part II
abstract
Nonlinear system design is often a multiobjective optimization problem involving the search for a design that satisfies a number of predefined constraints. The design space is typically very large since it includes all possible system architectures with different combinations of components composing each architecture. In this article, we address nonlinear system design space exploration through a two-step approach encapsulated in a framework called fast design space exploration of nonlinear systems (ASSENT). In the first step, we use a genetic algorithm to search for system architectures that allow discrete choices for component values or else only component values for a fixed architecture. This step yields a coarse design since the system may or may not meet the target specifications. In contrast to prior works on design space exploration that rely on forward design, we use an inverse design in step 2 to search over a continuous space and fine-tune the component values with the goal of improving the value of the objective function. We use a neural network (NN) to model the system response. The NN is converted into a mixed-integer linear program for active learning to sample component values efficiently. We illustrate the efficacy of ASSENT on problems ranging from nonlinear system design to the design of electrical circuits. Experimental results show that ASSENT achieves the same or better value of the objective function compared to various other optimization techniques for nonlinear system design by up to 53%. We improve sample efficiency by 6–$12\times $compared to reinforcement learning-based synthesis of electrical circuits.
Prerit Terway, Kenza Hamidouche, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 MHDeep: Mental Health Disorder Detection System Based on Wearable Sensors and Artificial Neural Networks
abstract
Mental health problems impact the quality of life of millions of people around the world. However, diagnosis of mental health disorders is a challenging problem that often relies on self-reporting by patients about their behavioral patterns and social interactions. Therefore, there is a need for new strategies for diagnosis and daily monitoring of mental health conditions. The recent introduction of body-area networks consisting of a plethora of accurate sensors embedded in smartwatches and smartphones and edge-compatible deep neural networks (DNNs) points toward a possible solution. Such wearable medical sensors (WMSs) enable continuous monitoring of physiological signals in a passive and non-invasive manner. However, disease diagnosis based on WMSs and DNNs, and their deployment on edge devices, such as smartphones, remains a challenging problem. These challenges stem from the difficulty of feature engineering and knowledge distillation from the raw sensor data, as well as the computational and memory constraints of battery-operated edge devices. To this end, we propose a framework called MHDeep that utilizes commercially available WMSs and efficient DNN models to diagnose three important mental health disorders: schizoaffective, major depressive, and bipolar. MHDeep uses eight different categories of data obtained from sensors integrated in a smartwatch and smartphone. These categories include various physiological signals and additional information on motion patterns and environmental variables related to the wearer. MHDeep eliminates the need for manual feature engineering by directly operating on the data streams obtained from participants. Because the amount of data is limited, MHDeep uses a synthetic data generation module to augment real data with synthetic data drawn from the same probability distribution. We use the synthetic dataset to pre-train the weights of the DNN models, thus imposing a prior on the weights. We use a grow-and-prune DNN synthesis approach to learn both architecture and weights during the training process. We use three different data partitions to evaluate the MHDeep models trained with data collected from 74 individuals. We conduct two types of evaluations: at the data instance level and at the patient level. MHDeep achieves an average test accuracy, across the three data partitions, of 90.4%, 87.3%, and 82.4%, respectively, for classifications between healthy and schizoaffective disorder instances, healthy and major depressive disorder instances, and healthy and bipolar disorder instances. At the patient level, MHDeep DNN models achieve an accuracy of 100%, 100%, and 90.0% for the three mental health disorders, respectively, based on inference that uses 40, 16, and 22 minutes of sensor data collection from each patient.
Shayan Hassantabar, Joe Zhang, Hongxu Yin, Niraj K. Jha
ACM Trans. Embed. Comput. Syst.4
2021 SECRET: Semantically Enhanced Classification of Real-World Tasks
abstract
Supervised machine learning (ML) algorithms are aimed at maximizing classification performance under available energy and storage constraints. They try to map the training data to the corresponding labels while ensuring generalizability to unseen data. However, they do not integrate meaning-based relationships among labels in the decision process. On the other hand, natural language processing (NLP) algorithms emphasize the importance of semantic information. In this article, we synthesize the complementary advantages of supervised ML and NLP algorithms into one method that we refer to as SECRET (Semantically Enhanced Classification of REal-world Tasks). SECRET performs classifications by fusing the semantic information of the labels with the available data: it combines the feature space of the supervised algorithms with the semantic space of the NLP algorithms and predicts labels based on this joint space. Experimental results indicate that, compared to traditional supervised learning, SECRET achieves up to 14.0 percent accuracy and 13.1 percent F1 score improvements. Moreover, compared to ensemble methods, SECRET achieves up to 12.7 percent accuracy and 13.3 percent F1 score improvements. This points to a new research direction for supervised classification based on incorporation of semantic information.
Ayten Ozge Akmandor, Jorge Ortiz 0001, Irene Manotas, Bong Jun Ko, Niraj K. Jha
IEEE Trans. Computers5
2021 Software-Defined Design Space Exploration for an Efficient DNN Accelerator Architecture
abstract
Deep neural networks (DNNs) have been shown to outperform conventional machine learning algorithms across a wide range of applications, e.g., image recognition, object detection, robotics, and natural language processing. However, the high computational complexity of DNNs often necessitates extremely fast and efficient hardware. The problem gets worse as the size of neural networks grows exponentially. As a result, customized hardware accelerators have been developed to accelerate DNN processing without sacrificing model accuracy. However, previous accelerator design studies have not fully considered the characteristics of the target applications, which may lead to sub-optimal architecture designs. On the other hand, new DNN models have been developed for better accuracy, but their compatibility with the underlying hardware accelerator is often overlooked. In this article, we propose an application-driven framework for architectural design space exploration of DNN accelerators. This framework is based on a hardware analytical model of individual DNN operations. It models the accelerator design task as a multi-dimensional optimization problem. We demonstrate that it can be efficaciously used in application-driven accelerator architecture design: we use the framework to optimize the accelerator configurations for eight representative DNNs and select the configuration with the highest geometric mean performance. The geometric mean performance improvement of the selected DNN configuration relative to the architectural configuration optimized only for each individual DNN ranges from 12.0 to 117.9 percent. Given a target DNN, the framework can generate efficient accelerator design solutions with optimized performance and area. Furthermore, we explore the opportunity to use the framework for accelerator configuration optimization under simultaneous diverse DNN applications. The framework is also capable of improving neural network models to best fit the underlying hardware resources. We demonstrate that it can be used to analyze the relationship between the operations of the target DNNs and the corresponding accelerator configurations, based on which the DNNs can be tuned for better processing efficiency on the given accelerator without sacrificing accuracy.
Ye Yu 0003, Yingmin Li, Shuai Che, Niraj K. Jha, Weifeng Zhang 0003
IEEE Trans. Computers4
2021 YSUY: Your Smartphone Understands You - Using Machine Learning to Address Fundamental Human Needs
abstract
Most machine learning (ML) models are geared toward improving some desired metric like classification accuracy or inference latency. Given the significant successes of such models, especially supervised ones, in addressing such metrics, the time has come to ask if they can also be put to direct use in understanding the users and addressing basic human needs, e.g., subsistence, protection, affection, understanding, participation, leisure, creation, identity, and freedom. A prerequisite to addressing such human needs is that the ML models exhibit a basic grasp of human psychology. In this article, we present ML models that can be embedded in our smartphone and enable it to understand us. We call this system your smartphone understands you (YSUY). YSUY uses wearable medical sensors to understand our physical, mental, and four-class (two-class) emotional states with 90.0%, 90.3%, and 98.4% (99.5%) accuracy, respectively. After verifying YSUY’s ability to understand the human condition from various perspectives, we evaluate the relationship between the different states and discuss how YSUY can be taken one step further to start playing a helpful role in addressing human needs of the types mentioned above from four different perspectives: “being,” “doing,” “having,” and “interacting.” We show that YSUY is a a promising candidate for adapting ML models to human-centric needs. We view this only as an initial step, hopefully, spurring other researchers to investigate the relationship between ML models and fulfilling human needs in much greater depth.
Ayten Ozge Akmandor, Xiaoliang Dai, Niraj K. Jha
IEEE Trans. Syst. Man Cybern. Syst.3
2020 Dreaming to Distill: Data-Free Knowledge Transfer via DeepInversion
abstract
We introduce DeepInversion, a new method for synthesizing images from the image distribution used to train a deep neural network. We ``invert'' a trained network (teacher) to synthesize class-conditional input images starting from random noise, without using any additional information about the training dataset. Keeping the teacher fixed, our method optimizes the input while regularizing the distribution of intermediate feature maps using information stored in the batch normalization layers of the teacher. Further, we improve the diversity of synthesized images using Adaptive DeepInversion, which maximizes the Jensen-Shannon divergence between the teacher and student network logits. The resulting synthesized images from networks trained on the CIFAR-10 and ImageNet datasets demonstrate high fidelity and degree of realism, and help enable a new breed of data-free applications - ones that do not require any real images or labeled data. We demonstrate the applicability of our proposed method to three tasks of immense practical importance - (i) data-free network pruning, (ii) data-free knowledge transfer, and (iii) data-free continual learning.
Hongxu Yin, Pavlo Molchanov 0001, José M. Álvarez 0004, Zhizhong Li 0001, Arun Mallya, Derek Hoiem, Niraj K. Jha, Jan Kautz
CVPR7
2020 INVITED: Efficient Synthesis of Compact Deep Neural Networks
abstract
Deep neural networks (DNNs) have been deployed in myriad machine learning applications. However, advances in their accuracy are often achieved with increasingly complex and deep network architectures. These large, deep models are often unsuitable for real-world applications, due to their massive computational cost, high memory bandwidth, and long latency. For example, autonomous driving requires fast inference based on Internet-of-Things (IoT) edge devices operating under run-time energy and memory storage constraints. In such cases, compact DNNs can facilitate deployment due to their reduced energy consumption, memory requirement, and inference latency. Long short-term memories (LSTMs) are a type of recurrent neural network that have also found widespread use in the context of sequential data modeling. They also face a model size vs. accuracy trade-off. In this paper, we review major approaches for automatically synthesizing compact, yet accurate, DNN/LSTM models suitable for real-world applications. We also outline some challenges and future areas of exploration.
Wenhan Xia, Hongxu Yin, Niraj K. Jha
DAC3
2020 Grow and Prune Compact, Fast, and Accurate LSTMs
abstract
Long short-term memory (LSTM) has been widely used for sequential data modeling. Researchers have increased LSTM depth by stacking LSTM cells to improve performance. This incurs model redundancy, increases run-time delay, and makes the LSTMs more prone to overfitting. To address these problems, we propose a hidden-layer LSTM (H-LSTM) that adds hidden layers to LSTM's original one-level nonlinear control gates. H-LSTM increases accuracy while employing fewer external stacked layers, thus reducing the number of parameters and run-time latency significantly. We employ grow-and-prune (GP) training to iteratively adjust the hidden layers through gradient-based growth and magnitude-based pruning of connections. This learns both the weights and the compact architecture of H-LSTM control gates. We have GP-trained H-LSTMs for image captioning, speech recognition, and neural machine translation applications. For the NeuralTalk architecture on the MSCOCO dataset, our three models reduce the number of parameters by 38.7× [floating-point operations (FLOPs) by 45.5×], run-time latency by 4.5×, and improve the CIDEr-D score by 2.8 percent, respectively. For the DeepSpeech2 architecture on the AN4 dataset, the first model we generated reduces the number of parameters by 19.4× and run-time latency by 37.4 percent. The second model reduces the word error rate (WER) from 12.9 to 8.7 percent. For the encoder-decoder sequence-to-sequence network on the IWSLT 2014 German-English dataset, the first model we generated reduces the number of parameters by 10.8× and run-time latency by 14.2 percent. The second model increases the BLEU score from 30.02 to 30.98. Thus, GP-trained H-LSTMs can be seen to be compact, fast, and accurate.
Xiaoliang Dai, Hongxu Yin, Niraj K. Jha
IEEE Trans. Computers3
2020 McPAT-Monolithic: An Area/Power/Timing Architecture Modeling Framework for 3-D Hybrid Monolithic Multicore Systems
abstract
Three-dimensional integrated circuits (3-D ICs) have the potential to push Moore's law further by accommodating more transistors per unit footprint area along with a reduction in power consumption, interconnect length, and the number of repeaters. Monolithic 3-D integration is particularly promising in this regard as it offers a very high connectivity between vertical transistor layers owing to its nanoscale monolithic intertier vias. Monolithic integration can be realized at block-, gate-, and transistor-level granularity. A hybrid monolithic (HM) design aims to further optimize area, power, and performance of the chip by combining different monolithic styles. In this article, we introduce McPAT-monolithic, a framework for modeling HM multicore architectures. We use the OpenSPARC T2 processor as a case study to compare different monolithic implementation styles and explore the benefits of HM design. Our simulations show that, under the same timing constraint, an HM design offers 47.2% reduction in footprint area and 5.3% in power consumption compared to a 2-D design at the cost of slightly higher on-chip temperature.
Abdullah Guler, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2019 ChamNet: Towards Efficient Network Design Through Platform-Aware Model Adaptation
abstract
This paper proposes an efficient neural network (NN) architecture design methodology called Chameleon that honors given resource constraints. Instead of developing new building blocks or using computationally-intensive reinforcement learning algorithms, our approach leverages existing efficient network building blocks and focuses on exploiting hardware traits and adapting computation resources to fit target latency and/or energy constraints. We formulate platform-aware NN architecture search in an optimization framework and propose a novel algorithm to search for optimal architectures aided by efficient accuracy and resource (latency and/or energy) predictors. At the core of our algorithm lies an accuracy predictor built atop Gaussian Process with Bayesian optimization for iterative sampling. With a one-time building cost for the predictors, our algorithm produces state-of-the-art model architectures on different platforms under given constraints in just minutes. Our results show that adapting computation resources to building blocks is critical to model performance. Without the addition of any special features, our models achieve significant accuracy improvements relative to state-of-the-art handcrafted and automatically designed architectures. We achieve 73.8% and 75.3% top-1 accuracy on ImageNet at 20ms latency on a mobile CPU and DSP. At reduced latency, our models achieve up to 8.2% (4.8%) and 6.7% (9.3%) absolute top-1 accuracy improvements compared to MobileNetV2 and MnasNet, respectively, on a mobile CPU (DSP), and 2.7% (4.6%) and 5.6% (2.6%) accuracy gains over ResNet-101 and ResNet-152, respectively, on an Nvidia GPU (Intel CPU).
Xiaoliang Dai, Peizhao Zhang, Bichen Wu, Hongxu Yin, Fei Sun 0002, Yanghan Wang, Marat Dukhan, Yunqing Hu, Yiming Wu 0013, Yangqing Jia, Peter Vajda, Matthew Uyttendaele, Niraj K. Jha
CVPR13
2019 NeST: A Neural Network Synthesis Tool Based on a Grow-and-Prune Paradigm
abstract
Deep neural networks (DNNs) have begun to have a pervasive impact on various applications of machine learning. However, the problem of finding an optimal DNN architecture for large applications is challenging. Common approaches go for deeper and larger DNN architectures but may incur substantial redundancy. To address these problems, we introduce a network growth algorithm that complements network pruning to learn both weights and compact DNN architectures during training. We propose a DNN synthesis tool (NeST) that combines both methods to automate the generation of compact and accurate DNNs. NeST starts with a randomly initialized sparse network called the seed architecture. It iteratively tunes the architecture with gradient-based growth and magnitude-based pruning of neurons and connections. Our experimental results show that NeST yields accurate, yet very compact DNNs, with a wide range of seed architecture selection. For the LeNet-300-100 (LeNet-5) architecture, we reduce network parameters by 70.2× (74.3×) and floating-point operations (FLOPs) by 79.4× (43.7×). For the AlexNet, VGG-16, and ResNet-50 architectures, we reduce network parameters (FLOPs) by 15.7× (4.6×), 33.2× (8.9×), and 4.1× (2.1×) respectively. NeST's grow-and-prune paradigm delivers significant additional parameter and FLOPs reduction relative to pruning-only methods.
Xiaoliang Dai, Hongxu Yin, Niraj K. Jha
IEEE Trans. Computers3
2019 Three-Dimensional Monolithic FinFET-Based 8T SRAM Cell Design for Enhanced Read Time and Low Leakage
abstract
FinFETs have replaced planar MOSFETs due to their superior performance, power efficiency, and scalability. However, even FinFETs are expected to reach their scaling limits due to physical limits, process variations, and short-channel effects. As an alternative to device scaling, 3-D integrated circuits (ICs) can increase the number of transistors in a chip in the same footprint area. Among 3-D technologies, monolithic 3-D integration promises the highest density, performance, and power efficiency owing to its high-density monolithic intertier vias. A transistor-level monolithic implementation enables an independent optimization of transistor layers. However, it requires a new 3-D cell library. In this paper, we propose two new 3-D monolithic FinFET-based 8T SRAM cells and compare them with previously reported 6T and 8T SRAM cells implemented in 2-D/3-D. Both the proposed cells use pFinFET access transistors for better area efficiency in 3-D and low leakage current. One of the proposed cells utilizes independent-gate pFinFETs as pull-up transistors whose back gates are tied to the supply voltage for better writeability. This cell has 28.1% and 43.8% smaller footprint area, 31.6% and 43.2% smaller leakage current, and 53.2% and 29.0% lower read time compared with conventional 2-D 6T SRAM and 2-D 8T SRAM cells, respectively.
Abdullah Guler, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2018 Statistical Optimization of FinFET Processor Architectures under PVT Variations Using Dual Device-Type Assignment
abstract
With semiconductor technology scaling to the 22nm node and beyond, fin field-effect transistor (FinFET) has started replacing complementary metal-oxide semiconductor (CMOS), thanks to its superior control of short-channel effects and much lower leakage current. However, process, supply voltage, and temperature (PVT) variations across the integrated circuit (IC) become worse with technology scaling. Thus, to analyze timing, leakage power, and dynamic power under PVT variations, statistical analysis/optimization techniques are more suitable than traditional static timing/power analysis and optimization counterparts. In this article, we propose a statistical optimization framework using dual device-type assignment at the architecture level under PVT variations that takes spatial correlations into account and leverages circuit-level statistical analysis techniques. To the best of our knowledge, this is the first work to study statistical optimization at the system level under PVT variations. Simulation results show that leakage power yield and dynamic power yield at the mean value of the baseline can be improved by up to 44.2% and 21.2%, respectively, with no loss in timing yield for a single-core processor and up to 43.0% and 50.0%, respectively, without any loss in timing yield for an 8-core chip multiprocessor (CMP), at little area overhead. Under the same (99.0%) power yield constraints, leakage power and dynamic power are reduced by up to 91.2% and 4.3%, respectively, for a single-core processor, and up to 44.6% and 12.5%, respectively, for an 8-core CMP, with no loss in timing yield. We also show that optimizations performed without taking module-to-module and core-to-core spatial correlations into account overestimate yield, establishing the importance of taking such correlations into account.
Ye Yu 0003, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2018 Genetic Programming for Energy-Efficient and Energy-Scalable Approximate Feature Computation in Embedded Inference Systems
abstract
With the increasing interest in deploying embedded sensors in a range of applications, there is also interest in deploying embedded inference capabilities. Doing so under the strict and often variable energy constraints of the embedded platforms requires algorithmic, in addition to circuit and architectural, approaches to reducing energy. A broad approach that has recently received considerable attention in the context of inference systems is approximate computing. This stems from the observation that many inference systems exhibit various forms of tolerance to data noise. While some systems have demonstrated significant approximation-versus-energy knobs to exploit this, they have been applicable to specific kernels and architectures; the more generally available knobs have been relatively weak, resulting in large data noise for relatively modest energy savings (e.g., voltage overscaling, bit-precision scaling). In this work, we explore the use of genetic programming (GP) to compute approximate features. Further, we leverage a method that enhances tolerance to feature-data noise through directed retraining of the inference stage. Previous work in GP has shown that it generalizes well to enable approximation of a broad range of computations, raising the potential for broad applicability of the proposed approach. The focus on feature extraction is deliberate because they involve diverse, often highly nonlinear, operations, challenging general applicability of energy-reducing approaches. We evaluate the proposed methodologies through two case studies, based on energy modeling of a custom low-power microprocessor with a classification accelerator. The first case study is on electroencephalogram-based seizure detection. We find that the choice of two primitive functions (square root, subtraction) out of seven possible primitive functions (addition, subtraction, multiplication, logarithm, exponential, square root, and square) enables us to approximate features in 0.41$mJ$per feature vector (FV), as compared to 4.79$mJ$per FV required for baseline feature extraction. This represents a feature extraction energy reduction of 11.68$\times$. The important system-level performance metrics for seizure detection are sensitivity, latency, and number of false alarms per hour. Our set of GP models achieves 100 percent sensitivity, 4.37 second latency, and 0.15 false alarms per hour. The baseline performance is 100 percent sensitivity, 3.84 second latency, and 0.06 false alarms per hour. The second case study is on electrocardiogram-based arrhythmia detection. In this case, just one primitive function (multiplication) suffices to approximate features in 1.13$\mu J$per FV, as compared to 11.69$\mu J$per FV required for baseline feature extraction. This represents a feature extraction energy reduction of 10.35$\times$. The important system-level metrics in this case are sensitivity, specificity, and accuracy. Our set of GP models achieves 81.17 percent sensitivity, 80.63 percent specificity, and 81.86 percent accuracy, whereas the baseline achieves 82.05 percent sensitivity, 88.12 percent specificity, and 87.92 percent accuracy. These case studies demonstrate the possibility of a significant reduction in feature extraction energy at the expense of a slight degradation in system performance.
Hongyang Jia, Naveen Verma, Niraj K. Jha
IEEE Trans. Computers4
2018 Hybrid Monolithic 3-D IC Floorplanner
Abdullah Guler, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Internet-of-Medical-Things
abstract
We have arrived at the dawn of the Internet-of-Things (IoT) era. 25 billion devices (things or physical objects) are already connected to the Internet, and this number is expected to grow to 50 billion by 2020. IoT is a network of physical objects. These objects contain sensors, actuators, and processing elements that enable us to gather data, monitor the health of the object, make intelligent decisions, and optimize processes. IoT is expected to have a potential economic impact of $3-6 trillion per year by 2025, with $1-2.5 trillion of this economic impact (its largest fraction) coming from smart healthcare applications. These applications will be enabled by a personal healthcare system consisting of implantable and wearable medical sensors and devices connected to a personal health hub (e.g., a smartphone or smartwatch) that is connected to the Internet.
Niraj K. Jha
ACM Great Lakes Symposium on VLSI1
2017 Automated Quantum Circuit Synthesis and Cost Estimation for the Binary Welded Tree Oracle
abstract
Quantum computing is a new computational paradigm that promises an exponential speed-up over classical algorithms. To develop efficient quantum algorithms for problems of a non-deterministic nature, random walk is one of the most successful concepts employed. In this article, we target both continuous-time and discrete-time random walk in both the classical and quantum regimes. Binary Welded Tree (BWT), or glued tree, is one of the most well-known quantum walk algorithms in the continuous-time domain. Prior work implements quantum walk on the BWT with static welding. In this context, static welding is randomized but case-specific. We propose a solution to automatically generate the circuit for the Oracle for welding. We implement the circuit using the Quantum Assembly Language, which is a language for describing quantum circuits. We then optimize the generated circuit using the Fault-Tolerant Quantum Logic Synthesis tool for any BWT instance. Automatic welding enables us to provide a generalized solution for quantum walk on the BWT.
Mrityunjay Ghosh, Amlan Chakrabarti, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.3
2017 CABA: Continuous Authentication Based on BioAura
abstract
Most computer systems authenticate users only once at the time of initial login, which can lead to security concerns. Continuous authentication has been explored as an approach for alleviating such concerns. Previous methods for continuous authentication primarily use biometrics, e.g., fingerprint and face recognition, or behaviometrics, e.g., key stroke patterns. We describe CABA, a novel continuous authentication system that is inspired by and leverages the emergence of sensors for pervasive and continuous health monitoring. CABA authenticates users based on their BioAura, an ensemble of biomedical signal streams that can be collected continuously and non-invasively using wearable medical devices. While each such signal may not be highly discriminative by itself, we demonstrate that a collection of such signals, along with robust machine learning, can provide high accuracy levels. We demonstrate the feasibility of CABA through analysis of traces from the MIMIC-II dataset. We propose various applications of CABA, and describe how it can be extended to user identification and adaptive access control authorization. Finally, we discuss possible attacks on the proposed scheme and suggest corresponding countermeasures.
Arsalan Mosenia, Susmita Sur-Kolay, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Computers4
2017 Improving Convergence and Simulation Time of Quantum Hydrodynamic Simulation: Application to Extraction of Best 10-nm FinFET Parameter Values
abstract
As electronic devices enter the deep nanometer regime, accurate and efficient device simulations become necessary to account for the emerging quantum effects. The traditional drift-diffusion and hydrodynamic (HD) device simulation models are not accurate in this regime. It is important to use the quantum HD (QHD) simulation model. However, this model suffers from poor convergence and high CPU times. To overcome these obstacles, in this paper, we propose a novel method to replace part of the QHD simulation that exhibits poor convergence behavior and high CPU time with HD simulation. In order to implement this, we capture the device states from the classical HD model and then apply the results as the initial guess to the QHD simulation, which is then solved by the Newton-Raphson method. This leads to significant improvements. The nonconvergence rate and the simulation time are reduced by 86.0% and 30.2%, respectively. As an application of the proposed methodology, we extract the best parameter values of both bulk and silicon-on-insulator FinFETs at the 10-nm technology node from their vast device design space.
Xiaoliang Dai, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Using a Device State Library to Boost the Performance of TCAD Mixed-Mode Simulation
abstract
Technology computer-aided design (TCAD) device and small circuit simulations use numerical and physics models to investigate the properties and performance of circuits before they undergo fabrication. Thus, these simulations play a significant role in VLSI design, optimization, and verification. However, they suffer from poor convergence and high CPU times, especially when performing TCAD mixed-mode simulations. In this paper, we propose a new simulation flow to address this challenge. We use device states captured from single devices to build a device state library. Then, we leverage device-level solutions to form a global initial guess for circuit-level simulations that are based on the full Newton algorithm. This leads to a significant efficiency enhancement. The average speedup for quasi-stationary (or dc) operating point establishment for a standard cell library and a mirror adder is 6.9× and 21.2×, respectively, whereas the CPU time for static random access memory static noise margin extraction is reduced by 47.0%.
Xiaoliang Dai, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Ultra-Low-Leakage and High-Performance Logic Circuit Design Using Multiparameter Asymmetric FinFETs
abstract
Recently, multigate field-effect transistors have started replacing traditional planar MOSFETs to keep pace with Moore’s Law in deep submicron technology. Among different multigate transistors, FinFETs have become the preferred choice of the semiconductor industry owing to low fabrication cost, superior performance, lower leakage, and design flexibility. The back and front gates of a FinFET can either be shorted or remain independent, leading to two modes of operation: Shorted-Gate (SG) and Independent-Gate (IG). For a given mode of operation, the physical parameters of the FinFET can either be symmetric or asymmetric in nature. In this article, for the first time, we analyze multiparameter asymmetric SG FinFETs and illustrate their potential for implementing logic gates and circuits that are both ultra-low-leakage and high-performance simultaneously. We restrict this work to SG devices because IG FinFETs (symmetric/asymmetric) suffer from severely degraded on-current, which makes them unattractive for high-performance designs. We first compare head-to-head all viable single- and multiparameter symmetric/asymmetric SG FinFETs. Among all such FinFETs, the traditional SG (which are symmetric in nature), Asymmetric Workfunction Shorted-Gate (AWSG), and Asymmetric Workfunction-Underlap Shorted-Gate (AWUSG) FinFETs show the most promise. We characterize these devices under process variations in gate length ( L G ), fin thickness ( T SI ), gate-oxide thickness ( T OX ), gate underlap ( L UN ), and gate-workfunction (Φ G ) as well as supply voltage ( V DD ) variations, followed by a gate-level leakage/delay analysis at different temperatures. Although AWSG FinFETs consume very low leakage power, they do suffer from performance degradation relative to SG FinFETs. Similarly, our study reveals that no other single-parameter asymmetric FinFET provides a good combination of low-power and high-performance design. We show that gates/circuits based on AWUSG FinFETs are faster, yet consume much less leakage power and less area than gates/circuits based on traditional SG FinFETs. We observe 53.4% (30.2%) maximum (average) reduction in total power at temperature T = 348 K while meeting the same delay constraint, with 14.2% (13.5%) reduction in area for AWUSG circuits relative to SG circuits. At T = 373 K , we see 68.6% (46.9%) maximum (average) reduction in total power.
Sourindra Chaudhuri, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2016 Ultra-low-leakage, Robust FinFET SRAM Design Using Multiparameter Asymmetric FinFETs
abstract
Memory arrays consisting of Static Random Access Memory (SRAM) cells occupy the largest area on chip and are responsible for significant leakage power consumption in modern microprocessors. With the transition from planar Complementary Metal-Oxide-Semiconductor (CMOS) technology to FinFETs, FinFET SRAM design has become important. However, increasing leakage power consumption of FinFETs due to aggressive scaling, width quantization, read-write conflict, and process variations make FinFET SRAM design challenging. In this article, we show how Multiparameter Asymmetric (MPA) FinFETs can be used to design ultra-low-leakage and robust 6T SRAM cells. We combine multiple asymmetries, namely, asymmetry in gate work function, source/drain doping concentration, and gate underlap, to address various SRAM design issues all at once. We propose five novel MPA FinFET SRAM cell designs and compare them with symmetric and Single-Parameter Asymmetric (SPA) FinFET SRAM cells using dc and transient metrics. We show that the leakage current of MPA FinFET SRAM cells can be reduced by up to 58 × while ensuring reasonable read/write stability metric values. In addition, high stability metric values can be achieved with 22 × leakage current reduction compared to the traditional symmetric FinFET SRAM cell. There is no area overhead associated with MPA FinFET SRAM cells.
Abdullah Guler, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2016 Delay/Power Modeling and Optimization of FinFET Circuit Modules under PVT Variations: Observing the Trends between the 22nm and 14nm Technology Nodes
abstract
The semiconductor industry has moved to FinFETs because of their superior ability to mitigate short-channel effects relative to CMOS. Thus, good FinFET delay and power models are urgently needed to facilitate FinFET IC design at the upcoming technology nodes. Another urgent problem that needs to be addressed with continued technology scaling is how to analyze circuit performance and power consumption under process, voltage, and temperature (PVT) variations. Such variations arise due to limitations of lithography that lead to variations in the physical dimensions of the device or due to environmental variations. In this article, we propose a delay/power modeling framework for analysis of FinFET logic circuits under PVT variations. We present models for FinFET logic gates and three FinFET SRAM cells. We use GenFin, which is a genetic algorithm based statistical circuit-level delay/power optimizer, to produce the models for functional units (FUs) employed in a processor. We compare the impact of PVT variations at the 22nm and 14nm FinFET technology nodes. We evaluate cache performance for various cache capacities and temperatures as well as that of FUs. Our device simulation results show that the 3σ/μ spread for 14nm circuits is, on average, 38.5% higher in dynamic power and 21.4% higher in logarithm of leakage power relative to 22nm FinFET circuits. However, the delay spread depends on the circuit.
Aoxiang Tang, Lung-Yen Chen, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.4
2016 Compressed Signal Processing on Nyquist-Sampled Signals
abstract
Pattern-recognition algorithms from the domain of machine learning play a prominent role in embedded sensing systems, in order to derive inferences from sensor data. Very often, such systems face severe energy constraints. The focus of this work is to mitigate the computational energy by exploiting a form of compression which preserves a similarity metric widely used for pattern recognition. The form of compression is random projection, and the similarity metric is inner products between source vectors. Given the prominence of random projections within compressive sensing, previous research has explored this idea for application to compressively-sensed signals. In this work, we analyze the error sources faced by such approaches and show that the compressive-sensing setting itself introduces a significant source of feature-computation error ($\sim$30 percent). We show that random projections can be exploited more generally without compressive sensing, enabling significant reduction in computational energy, and avoiding a significant source of error. The approach is referred to ascompressed signal processing (CSP), and it applies to Nyquist-sampled signals. We validate the CSP approach through two case studies. The first focuses on seizure detection using spectral-energy features extracted from electroencephalograms. We show that at a 32$\times$compression ratio, the number of multiply-accumulate (MAC) and operand-access operations required is reduced by 21.2$\times$, while achieving a sensitivity of 100 percent, latency of 4.33 sec, and false alarm rate of 0.22/hr; this compares to a baseline performance of 100 percent, 4.37 sec, and 0.12/hr, respectively. The second case study focuses on neural prosthesis based on extracting wavelet features from a set of detected spikes. We show that at a 32$\times$compression ratio, the number of MAC and operand access computations required is reduced by 3.3$\times$, while spike sorting performance can be maintained within an average error of 4.89 percent for spike count, 3.42 percent for coefficient of variance, and 4.90 percent for firing rate; this compares with a baseline average error of 4.00, 2.75, and 4.00 percent for spike count, coefficient of variance, and firing rate, respectively.
Naveen Verma, Niraj K. Jha
IEEE Trans. Computers3
2016 TCAD-Assisted Capacitance Extraction of FinFET SRAM and Logic Arrays
abstract
Technology computer-aided design (TCAD)-assisted automation of structure synthesis followed by a transport-analysis-based capacitance extraction approach has been shown to have a significant potential in the recent past in terms of both accuracy and efficiency. In this brief, we first propose three methods to speedup TCAD-assisted capacitance extraction of an SRAM array layout, by partitioning the layout into several fragments. We demonstrate that the speedup of these partitioning methods, relative to that of TCAD-assisted full-array capacitance extraction method, can be as high as 17×, with an average of 3×. On the other hand, the error in capacitance derived using the partitioning methods can be as low as 1.01%, with an average of 4.93%. We apply the partitioning methods to several layout configurations and styles of FinFET SRAM array layouts and demonstrate their efficacy. We then apply the partitioning methods to a ring oscillator, register, and two-bit adder, and demonstrate similar results.
Debajit Bhattacharya, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2016 A 3-D CPU-FPGA-DRAM Hybrid Architecture for Low-Power Computation
abstract
The power budget is expected to limit the portion of the chip that we can power ON at the upcoming technology nodes. This problem, known as the utilization wall or dark silicon, is becoming increasingly serious. With the introduction of 3-D integrated circuits (ICs), it is likely to become more severe. Thus, how to take advantage of the extra transistors, made available by Moore's law and the onset of 3-D ICs, within the power budget poses a significant challenge to system designers. To address this challenge, we propose a 3-D hybrid architecture consisting of a CPU layer with multiple cores, a field-programmable gate array (FPGA) layer, and a DRAM layer. The architecture is designed for low power without sacrificing performance. The FPGA layer is capable of supporting a large number of accelerators. It is placed adjacent to the CPU layer, with a communication mechanism that allows it to access CPU data caches directly. This enables fast switches between these two layers. This architecture reduces the power and energy significantly, at better or similar performance. This then alleviates the dark silicon problem by letting us power ON more components to achieve higher performance. We evaluate the proposed architecture through a new framework we have developed. Relative to the out-of-order CPU, the accelerators on the FPGA layer can reduce function-level power by 6.9× and energy-delay product (EDP) by 7.2×, and application-level power by 1.9× and EDP by 2.2×, while delivering similar performance. For the entire system, this translates to a 47.5% power reduction relative to a baseline system that consists of a CPU layer and a DRAM layer. This also translates to a 72.9% power reduction relative to an alternative system that consists of a CPU layer, an L3 cache layer, and a DRAM layer.
Xianmin Chen, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Reducing Wire and Energy Overheads of the SMART NoC Using a Setup Request Network
abstract
Network-on-chip (NoC), due to its superior scalability, has become increasingly popular in multicore and many-core designs. A recent state-of-the-art NoC design, called single-cycle multihop asynchronous repeated traversal (SMART), enables ultralow latency performance by enabling a flit to bypass the pipelines of intermediate routers entirely. This enables a flit to traverse multiple routers within a single clock cycle. However, there are two concerns related to SMART: 1) it employs dedicated broadcast wires to transmit SMART-hop setup requests (SSRs), incurring large wire and energy overheads and 2) it needs a complex allocator to arbitrate multiple simultaneous SSRs. In this paper, we propose an SSR network to address these two concerns. This is a specialized network that replaces long and overlapping broadcast wires with shorter wires and switches. It reduces the wire and energy overheads significantly. It also eliminates lowpriority SSRs before they reach the allocator and, therefore, leads to a simplified allocator design. We evaluate SSR networks in various contexts. Our evaluation demonstrates that the SSR network can reduce the wire overhead by up to 12.2×. Moreover, it results in up to 63.7% area reduction, and up to 15.7% dynamic energy reduction. It also makes multiple physical networks and/or wider flits feasible. This paper shows that with fewer wires and the same amount of buffer space, multiple simple SMART-1D networks can deliver a higher bandwidth and lower latency than a single complex SMART-2D network, when our SSR network is used.
Xianmin Chen, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2016 GenFin: Genetic Algorithm-Based Multiobjective Statistical Logic Circuit Optimization Using Incremental Statistical Analysis
abstract
As the semiconductor technology node scales into the deep submicrometer regime, it has become very difficult to obtain high IC yields because the process-voltage-temperature variations induce large spreads in delay and power. In this paper, we propose a new framework, called GenFin, which is, as far as we know, the first to target the multiobjective yield optimization of logic circuits. Since FinFETs are a promising substitute for CMOS at 22-nm technology node and beyond, we evaluate the framework with a 22-nm FinFET logic library. By combining the power of genetic algorithm (GA) and adaptive multiobjective optimization, GenFin produces a set of nondominated logic circuits whose timing, leakage power, and dynamic power yields are simultaneously optimized. This can help designers make tradeoff decisions wisely and avoid suboptimal solutions. We also propose an incremental statistical circuit analyzer, called incremental FinPrin, that speeds up the statistical static timing analysis by up to 9.6× and the statistical power analysis by up to 2235.7×, while incurring errors of only up to 0.031% in mean and 0.74% in standard deviation relative to nonincremental analysis. We use heuristics based on the deterministic timing analysis and gate criticality to reduce the GA search space and also improve the quality of its solutions. We present extensive experimental results to demonstrate the efficacy of GenFin.
Aoxiang Tang, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Vibration-based secure side channel for medical devices
abstract
Implantable and wearable medical devices are used for monitoring, diagnosis, and treatment of an ever-increasing range of medical conditions, leading to an improved quality of life for patients. The addition of wireless connectivity to medical devices has enabled post-deployment tuning of therapy and access to device data virtually anytime and anywhere but, at the same time, has led to the emergence of security attacks as a critical concern. While cryptography and secure communication protocols may be used to address most known attacks, the lack of a viable secure connection establishment and key exchange mechanism is a fundamental challenge that needs to be addressed. We propose a vibration-based secure side channel between an external device (medical programmer or smartphone) and a medical device. Vibration is an intrinsically short-range, user-perceptible channel that is suitable for realizing physically secure communication at low energy and size/weight overheads. We identify and address key challenges associated with the vibration channel, and propose a vibration-based wakeup and key exchange scheme, named SecureVibe, that is resistant to battery drain attacks. We analyze the risk of acoustic eavesdropping attacks and propose an acoustic masking countermeasure. We demonstrate and evaluate vibration-based wakeup and key exchange between a smartphone and a prototype medical device in the context of a realistic human body model.
Younghyun Kim 0001, Woo Suk Lee, Vijay Raghunathan, Niraj K. Jha, Anand Raghunathan
DAC4
2015 Synthesis of Quantum Circuits for Dedicated Physical Machine Descriptions
Philipp Niemann 0001, Saikat Basu, Amlan Chakrabarti, Niraj K. Jha, Robert Wille
RC4
2015 gem5-PVT: A Framework for FinFET System Simulation under PVT Variations
abstract
FinFET has begun replacing CMOS at the 22nm technology node and beyond. Compared to planar CMOS, FinFET has a higher on-current and lower leakage due to its double-gate structure. A FinFET-based system simulation framework can be very helpful to system architects for early-stage design-space exploration using this new technology. However, such a simulator does not exist. We fill this gap by presenting the details of one such simulation framework, called gem5-PVT, that we have developed. Our simulation framework combines and extends existing lower-level FinFET simulators to support timing, power, and thermal studies of FinFET-based chip multiprocessor systems under process, voltage, and temperature (PVT) variations. It uses a bottom-up modeling approach based on logic/memory cell libraries that have been very accurately characterized using TCAD device simulation. This allows accuracy to bubble up to the system level. The framework is modular and automated, hence enables system designers the flexibility to evaluate various system implementations. It is currently targeted at the 22nm FinFET technology. We report results for two case studies to demonstrate its usefulness. One study shows that more than 62.1× system-level leakage reduction, at the same performance, is possible when using a particular FinFET logic style. Another study characterizes core-to-core frequency and power variations that result from underlying PVT variations and compares the effectiveness of variation-aware scheduling schemes.
Xianmin Chen, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2015 Systematic Poisoning Attacks on and Defenses for Machine Learning in Healthcare
abstract
Machine learning is being used in a wide range of application domains to discover patterns in large datasets. Increasingly, the results of machine learning drive critical decisions in applications related to healthcare and biomedicine. Such health-related applications are often sensitive, and thus, any security breach would be catastrophic. Naturally, the integrity of the results computed by machine learning is of great importance. Recent research has shown that some machine-learning algorithms can be compromised by augmenting their training datasets with malicious data, leading to a new class of attacks called poisoning attacks. Hindrance of a diagnosis may have life-threatening consequences and could cause distrust. On the other hand, not only may a false diagnosis prompt users to distrust the machine-learning algorithm and even abandon the entire system but also such a false positive classification may cause patient distress. In this paper, we present a systematic, algorithm-independent approach for mounting poisoning attacks across a wide range of machine-learning algorithms and healthcare datasets. The proposed attack procedure generates input data, which, when added to the training set, can either cause the results of machine learning to have targeted errors (e.g., increase the likelihood of classification into a specific class), or simply introduce arbitrary errors (incorrect classification). These attacks may be applied to both fixed and evolving datasets. They can be applied even when only statistics of the training dataset are available or, in some cases, even without access to the training dataset, although at a lower efficacy. We establish the effectiveness of the proposed attacks using a suite of six machine-learning algorithms and five healthcare datasets. Finally, we present countermeasures against the proposed generic attacks that are based on tracking and detecting deviations in various accuracy metrics, and benchmark their effectiveness.
Mehran Mozaffari Kermani, Susmita Sur-Kolay, Anand Raghunathan, Niraj K. Jha
IEEE J. Biomed. Health Informatics4
2015 Design of Efficient Content Addressable Memories in High-Performance FinFET Technology
abstract
Content addressable memories (CAMs) enable high-speed parallel search operations in table lookup-based applications, such as Internet routers and processor caches. Traditional CAM design has always suffered from the high dynamic power consumption associated with its large and active parallel hardware. However, deeply scaled technology nodes, with multigate devices replacing planar MOSFETs, are expected to bring new tradeoffs to CAM design. FinFET, a vertical-channel gate-wrap-around double-gate device, has emerged as the best alternative to planar MOSFET. In this brief, for the first time, we explore the design space of symmetric and asymmetric gate-workfunction FinFET CAMs. We propose several design alternatives and evaluate them in terms of their dc and transient metrics for different mismatch probabilities using technology computer-aided design simulations with 22-nm FinFET devices. We also propose two orthogonal layout styles for CAM design and show that one of them (vertical-search line) outperforms the other (vertical-match line) in terms of total power (22.3%) and search delay (5.8%).
Debajit Bhattacharya, Ajay N. Bhoj, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
2015 PAQCS: Physical Design-Aware Fault-Tolerant Quantum Circuit Synthesis
abstract
Quantum circuits consist of a cascade of quantum gates. In a physical design-unaware quantum logic circuit, a gate is assumed to operate on an arbitrary set of quantum bits (qubits), without considering the physical location of the qubits. However, in reality, physical qubits have to be placed on a grid. Each node of the grid represents a qubit. The grid implements the architecture of the quantum computer. A physical constraint often imposed is that quantum gates can only operate on adjacent qubits on the grid. Hence, a communication channel needs to be built if the qubits in the logical circuit are not adjacent. In this paper, we introduce a tool called the physical design-aware fault-tolerant quantum circuit synthesis (PAQCS). It contains two algorithms: one for physical qubit placement and another for routing of communications. With the help of these two algorithms, the overhead of converting a logical to a physical circuit is reduced by 30.1%, on an average, relative to previous work. The optimization algorithms in PAQCS are evaluated on circuits implemented using quantum operations supported by two different quantum physical machine descriptions and three quantum error-correcting codes. They reduce the number of primitive operations by 11.5%-68.6%, and the number of execution cycles by 16.9%-59.4%.
Chia-Chun Lin, Susmita Sur-Kolay, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
2015 FDR 2.0: A Low-Power Dynamically Reconfigurable Architecture and Its FinFET Implementation
abstract
Large area/delay/power overheads are required to support the reconfigurability of field-programmable gate arrays (FPGAs). We proposed a hybrid CMOS/nanotechnology dynamically reconfigurable architecture, called NATURE, earlier to address this challenge. It uses the concept of temporal logic folding and fine-grain (i.e., cycle-level) dynamic reconfiguration to increase logic density and save area. Because logic folding reduces area significantly, most of the on-chip communications become localized. To take full advantage of localized communications, we then presented a new CMOS-based fine-grain dynamically reconfigurable (FDR) architecture. It consists of an array of homogeneous logic elements (LEs), which can be configured into logic or interconnect or a combination of both. FDR eliminates most of the long-distance and global wires, which occupy a large amount of area in conventional FPGAs. FDR improves the area-delay product by an order of magnitude relative to conventional architectures. In this paper, we present an augmented FDR 2.0 architecture, where: 1) the LE is augmented with dedicated carry logic to facilitate arithmetic operations; 2) diagonal direct links are incorporated to improve the flexibility of local communication; and 3) coarse-grain blocks, including embedded memories and digital signal processing (DSP) blocks, are added to support fast data-intensive computations. Experimental results show that the coarse-grain design can improve circuit performance by 3.6× compared with the fine-grain FDR architecture. Incorporation of the DSP blocks in FDR 2.0 also enables more effective area-delay and power-delay tradeoffs, allowing the users to trade performance for smaller area or power consumption. We have implemented the design in the 22-nm FinFET technology, which enables more flexible and effective power management. Finally, different types of FinFETs and power management techniques have been explored in FDR 2.0 to optimize power.
Ting-Jung Lin, Wei Zhang 0012, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
2015 Signal Processing With Direct Computations on Compressively Sensed Data
abstract
Sparsity is characteristic of a signal that potentially allows us to represent information efficiently. We present an approach that enables efficient representations based on sparsity to be utilized throughout a signal processing system, with the aim of reducing the energy and/or resources required for computation, communication, and storage. The representation we focus on is compressive sensing. Its benefit is that compression is achieved with minimal computational cost through the use of random projections; however, a key drawback is that reconstruction is expensive. We focus on inference frameworks for signal analysis. We show that reconstruction can be avoided entirely by transforming signal processing operations (e.g., wavelet transforms, finite impulse response filters, etc.) such that they can be applied directly to the compressed representations. We present a methodology and a mathematical framework that achieve this goal and also enable significant computational-energy savings through operations over fewer input samples. This enables explicit energy-versus-accuracy tradeoffs that are under the control of the designer. We demonstrate the approach through two case studies. First, we consider a system for neural prosthesis that extracts wavelet features directly from compressively sensed spikes. Through simulations, we show that spike sorting can be achieved with 54× fewer samples, providing an accuracy of 98.63% in spike count, 98.56% in firing-rate estimation, and 96.51% in determining the coefficient of variation; this compares with a baseline Nyquist-domain detector with corresponding performance of 98.97%, 99.69%, and 97.09%, respectively. Second, we consider a system for detecting epileptic seizures by extracting spectral-energy features directly from compressively sensed electroencephalogram. Through simulations of the end-to-end algorithm, we show that detection can be achieved with 21× fewer samples, providing a sensitivity of 94.43%, false alarm rate of 0.1543/h, and latency of 4.70 s; this compares with a baseline Nyquist-domain detector with corresponding performance of 96.03%, 0.1471/h, and 4.59 s, respectively.
Mohammed Shoaib, Niraj K. Jha, Naveen Verma
IEEE Trans. Very Large Scale Integr. Syst.2
2015 McPAT-PVT: Delay and Power Modeling Framework for FinFET Processor Architectures Under PVT Variations
abstract
As technology has moved into the deep-submicrometer regime, the shrinking feature size has placed a considerable stress on CMOS fabrication due to short-channel effects (SCEs) and excessive leakage. Although many research efforts have been devoted to seeking system-level solutions, underlying transistor-level solutions are still urgently required to overcome these obstacles. FinFETs have emerged as promising substitutes for conventional CMOS due to their superior control of SCEs and process scalability. However, FinFETs still face lithographic and workfunction engineering challenges, in addition to those posed by supply voltage and temperature variations across the integrated circuit (IC). These lead to process, supply voltage, and temperature (PVT) variations in FinFET ICs, which, in turn, lead to large spreads in delay and leakage. In this paper, we present a multicore power, area, and timing (McPAT)-PVT, an integrated framework for the simulation of power, delay, as well as PVT variations of FinFET-based processors. McPAT-PVT uses a FinFET design library, consisting of logic and memory cells, to model circuit-level characteristics as well as their PVT variation trends. It is based on macromodels, derived from very accurate TCAD device simulations that characterize various functional units in a processor under PVT variations, making yield analysis for timing and power for processor components possible. McPAT-PVT can model both shorted-gate (SG) and asymmetric-workfunction shorted-gate (ASG) FinFET-based processors. Combining these macromodels with a FinFET-based CACTI-PVT cache model and an ORION-PVT on-chip network model, McPAT-PVT is able to simulate a delay and power consumption of all processor components under PVT variations. We present extensive simulation results to demonstrate its efficacy, including for an alpha-like processor and multicore simulations based on Princeton Application Repository for Shared-Memory Computers benchmarks. Results show that the ASG FinFET-based processor implementation has 73× lower leakage power and 2.6× lower total power relative to the SG FinFET-based processor implementation for the same performance, with <;1% area penalty.
Aoxiang Tang, Chun-Yi Lee, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.4
2014 3D vs. 2D Device Simulation of FinFET Logic Gates under PVT Variations
abstract
Recently, multigate transistors have been gaining attention as an alternative to conventional MOSFETs. Superior gate control over the channel, smaller subthreshold leakage, and reduced susceptibility to process variations are some of the key features that give multigate structures a competitive edge over MOSFETs. Among various multigate structures, silicon-on-insulator (SOI) FinFETs are promising, owing to their ease of fabrication. However, characterization of SOI FinFET devices/gates needs immediate attention in order for them to gain greater popularity in this decade. Ideally, 3D device simulation should be done for accurate circuit analysis. However, this is impractical due to the huge CPU time required. As a possible alternative, simulating a 2D crosssection of the device yields 10× to 100× reduction in CPU time. However, this introduces significant error in the range of 7% to 20% when evaluating the on/off current ( I ON /I OFF ) for a single device and leakage current or propagation delay ( I LEAK /t D ) for logic gates. In this work, we first present a methodology to obtain optimized 3D device simulation models for SOI FinFETs. Based on these 3D models, we develop adjusted 2D models to capture 3D simulation accuracy with 2D simulation efficiency. We report results for the 22nm SOI FinFET technology node. We adjust gate underlap ( L UN ) in the 2D cross section of the n/pFinFET devices in order to mimic 3D device behavior. When the adjusted 2D models are employed in mixed-mode simulation of FinFET logic gates, the error in the evaluation of I LEAK /t D is very small. To the best of our knowledge, this is the first such attempt. We show that 2D device models remain valid even under process, voltage, and temperature (PVT) variations. We target process variations in gate length ( L G ), fin thickness ( T SI ), gate oxide thickness ( T OX ), and gate workfunction ( Φ G ), which are the parameters that have been shown to have the most impact on leakage and delay.
Sourindra Chaudhuri, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2014 Accurate Leakage/Delay Estimation for FinFET Standard Cells under PVT Variations using the Response Surface Methodology
abstract
Among different multi-gate transistors, FinFETs and Trigate FETs have set themselves apart as the most promising candidates for the upcoming 22nm technology node and beyond owing to their superior device performance, lower leakage power consumption, and cost-effective fabrication process. Innovative circuit design and optimization techniques will be required to harness the power of multi-gate transistors, which in turn will depend on accurate leakage and timing characterization of these devices under spatial and environmental variations. Hence, in order to aid circuit designers, we present accurate analytical models using central composite rotatable design (CCRD) based on response surface methodology (RSM) to estimate the leakage current and delay of FinFET standard cells under the effect of variations in gate length ( L G ), fin thickness ( T SI ), gate-oxide thickness (T OX ), gate-workfunction (Φ G ), supply voltage ( V DD ), and temperature ( T ). To the best of our knowledge, this is the first such attempt to develop analytical models for leakage/delay estimation of FinFET logic gates. To derive these models, we employ TCAD device simulations of adjusted 2D device cross sections that have been shown to track TCAD device simulations of 3D device behavior within a 1--3% error range. This drastically reduces the CPU time of our modeling technique (by several orders of magnitude) without much loss in accuracy. We present analytical leakage and delay models for different sizes and logic styles (e.g., shorted-gate (SG) and independent-gate (IG) FinFETs at the 22nm technology node). Both leakage and delay estimates derived from the analytical models are in close agreement with quasi-Monte Carlo (QMC) simulation results (QMC simulations track the accuracy of Monte Carlo simulations, but are several orders of magnitude faster) obtained for different adjusted-2D logic gates with a root mean square error (RMSE) in the 0.23%--5.87% range.
Sourindra Chaudhuri, Prateek Mishra, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.3
2014 Ultra-low-leakage chip multiprocessor design with hybrid FinFET logic styles
abstract
FinFET has begun replacing CMOS at the 22nm technology node because of its enhanced ability to mitigate short-channel effects. Although leakage power of FinFET logic gates is lower than their CMOS counterparts, it still contributes to a large part of total power consumption. In this article, we show how ultra-low-leakage FinFET chip multiprocessors (CMPs) can be designed using a hybrid logic style. This hybrid style exploits the ultra-low-leakage feature of asymmetric-workfunction shorted-gate (ASG) FinFETs and the high-performance feature of shorted-gate (SG) FinFETs. We explore the impact of the hybrid style at both the module and CMP levels. To do this, we have developed FinFET logic libraries targeted at SG and ASG logic gates, suitably characterized for various parameters of interest. We have also modified existing tools and created a framework to evaluate the hybrid designs of SRAMs, caches, and CMPs. Using the design with SG FinFETs as the baseline for comparison, our experimental results show that the hybrid style can reduce leakage power of execution units to as low as 10.6% of the baseline without hurting performance, that of SRAMs to between 21.5% and 4.8% of the baseline with 0%-8.3% delay overhead, and that of CMPs to 10.0% of the baseline with negligible performance degradation.
Xianmin Chen, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2014 QLib: Quantum module library
abstract
Quantum algorithms are known for their ability to solve some problems much faster than classical algorithms. They are executed on quantum circuits, which consist of a cascade of quantum gates. However, synthesis of quantum circuits is not straightforward because of the complexity of quantum algorithms. Generally, quantum algorithms contain two parts: classical and quantum. Thus, synthesizing circuits for the two parts separately reduces overall synthesis complexity. In addition, many quantum algorithms use similar subroutines that can be implemented with similar circuit modules. Because of their frequent use, it is important to use automated scripts to generate such modules efficiently. These modules can then be subjected to further synthesis optimizations. This article proposes QLib, a quantum module library, which contains scripts to generate quantum modules of different sizes and specifications for well-known quantum algorithms. Thus, QLib can also serve as a suite of benchmarks for quantum logic and physical synthesis.
Chia-Chun Lin, Amlan Chakrabarti, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.3
2014 RMDDS: Reed-muller decision diagram synthesis of reversible logic circuits
abstract
In this article, we propose a flexible and efficient reversible logic synthesizer. It exploits the complementary advantages of two methods: Reed-Muller Reversible Logic Synthesis (RMRLS) and Decision Diagram Synthesis (DDS), and is thus called Reed-Muller Decision Diagram Synthesis (RMDDS). RMRLS does not scale to a large number of qubits (i.e., quantum bits). DDS tools, even though efficient, add a large number of ancillary qubits and typically incur much higher quantum cost than necessary. RMDDS overcomes these obstacles. It is flexible in the sense that users can either optimize the number of qubits or the quantum cost in the circuit implementation. It is also efficient because the circuits can be synthesized within user-defined CPU times. This combination of flexibility and efficiency has been missing from synthesizers presented earlier. When used to synthesize reversible functions, RMDDS reduces the number of qubits by up to 79.2% (average of 54.6%) when the synthesis objective is to minimize the number of qubits and the quantum cost by up to 71.5% (average of 35.7%) when the synthesis objective is to minimize quantum cost, relative to DDS methods. For irreversible functions (which are automatically embedded in reversible functions), the corresponding best (average) reductions in the number of qubits is 42.1% (22.5%) when minimizing the number of qubits, and in quantum cost, it is 63.0% (25.9%) when minimizing quantum cost.
Chia-Chun Lin, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2014 Trustworthiness of Medical Devices and Body Area Networks
abstract
Implantable and wearable medical devices (IWMDs) are commonly used for diagnosing, monitoring, and treating various medical conditions. A general trend in these medical devices is toward increased functional complexity, software programmability, and connectivity to body area networks (BANs). However, as IWMDs become more “intelligent,” they also become less trustworthy-less reliable and more prone to attacks. Various shortcomings-hardware failures, software errors, wireless attacks, malware and software exploits, and side-channel attacks-could undermine the trustworthiness of IWMDs and BANs. While these concerns have been recognized for some time, recent demonstrations of security attacks on commercial products, e.g., pacemakers and insulin pumps, have elevated medical device security from the realm of theoretical possibility to an immediate concern. The trustworthiness of IWMDs must be addressed aggressively and proactively due to the potential for catastrophic consequences. Conventional fault tolerance and information security solutions, e.g., redundancy and cryptography, that have been employed in general-purpose and embedded computing systems cannot be applied to many IWMDs due to their extreme size and power constraints and unique usage models. While several recent efforts address defense of IWMDs against specific security attacks, a holistic strategy that considers all concerns and types of threats is required. This paper discusses trustworthiness concerns in IWMDs and BANs through a comprehensive identification and analysis of potential threats and, for each threat, provides a discussion of the merits and inadequacies of current solutions.
Anand Raghunathan, Niraj K. Jha
Proc. IEEE3
2014 Parasitics-Aware Design of Symmetric and Asymmetric Gate-Workfunction FinFET SRAMs
abstract
Multigate FET technology is the most viable successor to planar CMOS technology at the 22-nm node and beyond. Prior research on multigate SRAMs is generally confined to the optimization of DC targets. However, on account of the nonplanar nature of multigate FETs, it is highly questionable whether multigate SRAM DC metrics can guide bitcell designers, as parasitic capacitances for two topologically equivalent bitcells can be very different - due to various issues such as fin pitches - resulting in widely varying transient characteristics. In this paper, we evaluate several known symmetric gate-workfunction (Symm- ΦG) 6T FinFET SRAMs and, for the first time, asymmetric gate-workfunction (Asymm-ΦG) 6T FinFET SRAMs, head-to-head in a 22-nm silicon-on-insulator process, from the perspective of transient behavior, using a unified 3-D/mixed-mode 2-D TCAD technology-circuit co-design methodology. We accomplish the latter by capturing bitcell parasitics accurately through transport analysis-based 3-D TCAD capacitance extractions that leverage automated layout-3-D TCAD structure synthesis algorithms. Mixed-mode transient device simulations (incorporating back-annotated 3-D TCAD parasitics) indicate that a design guided by DC metrics alone can lead to erroneous conclusions and suboptimal bitcell choices. Overall, from the perspective of area and performance, in single- ΦGprocesses, shorted-gate (or vanilla) configurations are superior to topologies employing independent-gate configurations, even though the latter often have better DC metrics. In a larger design space encompassing dual/Asymm-ΦGdevices, Asymm-ΦGFinFET SRAMs are very competitive with respect to vanilla topologies in terms of DC metrics and have better dynamic write-ability, even at low VDD.
Ajay N. Bhoj, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2014 FinCANON: A PVT-Aware Integrated Delay and Power Modeling Framework for FinFET-Based Caches and On-Chip Networks
abstract
Recently, FinFETs have emerged as promising substitutes for conventional CMOS because of their superior control of short-channel effects and processing scalability. Nevertheless, lithographic constraints, difficulties in workfunction engineering, supply voltage variations, and temperature nonuniformity across the FinFET integrated circuit may lead to process, supply voltage, and temperature (PVT) variations, which are manifested as large spreads in delay and leakage. In this paper, we present FinCANON, an integrated framework for the simulation of power, delay, as well as PVT variations of FinFET-based caches and on-chip networks. FinCANON consists of CACTI-PVT and ORION-PVT that model caches and on-chip networks, respectively. We have developed a FinFET design library to model the circuit-level characteristics as well as their variation trends with respect to various PVT parameters for FinFET logic gates and memory cells, using accurate device simulation. With a statistical static timing analysis technique and macromodel-based methodology, we have also derived PVT variation models for delay and leakage, considering spatial correlations, to characterize the impact of PVT variations on FinFET-based caches and networks-on-chip (NoCs). In addition, we incorporate voltage generators in the FinFET design library to model back-gate biasing of FinFETs. The cache and NoC models are significantly enhanced to be more modular and scalable. We present results for various FinFET design styles and show that mixing different styles may be a promising strategy for optimizing delay and leakage of caches and NoCs.
Chun-Yi Lee, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2014 FTQLS: Fault-Tolerant Quantum Logic Synthesis
abstract
Quantum computation can solve certain problems much faster than classical computation. However, quantum computations are more susceptible to errors than conventional digital computations. Thus, a fault-tolerant (FT) quantum circuit design is required for a practical implementation. Quantum circuits consist of a cascade of quantum gates. These gates are themselves realized using primitive quantum operations that are supported by the quantum physical machine description (PMD). As different quantum systems are associated with different Hamiltonians, they have different PMDs. In addition, the quantum cost for implementing a quantum operation may differ from one PMD to another. Thus, a quantum logic circuit needs to be realized with and optimized for the set of primitive quantum operations supported by the given PMD. In this paper, an FT quantum logic synthesis (FTQLS) methodology and tool is described for six different PMDs. A methodology, such as this, which can be targeted at multiple PMDs, has not been attempted before, to the best of our knowledge. The input to FTQLS is an unoptimized quantum circuit realized using a set of commonly used gates and its output is an optimized FT quantum circuit that only comprises of primitive quantum operations supported by the given PMD. FTQLS does technology mapping for different PMDs and then converts non-FT circuits to FT circuits. For technology mapping, it utilizes an optimized quantum gate library targeted at various PMDs that decomposes gates into primitive operations. Efficient conversion to FT circuits is done by integrating two quantum compilers and an FT cache table into FTQLS. For improving the synthesis results, an FT set of gates that is directly supported by each PMD is proposed. Quantum circuit optimization is done by utilizing quantum identity rules. The performance of FTQLS is evaluated using two cost metrics: number of primitive operations (#ops) and execution cycles on the critical path (#cycles). Experiment results show that the decrease in #ops varies between 58.1% and 87.0% and in #cycles between 42.8% and 76.4%, on an average, depending on the PMD.
Chia-Chun Lin, Amlan Chakrabarti, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
2014 A Fine-Grain Dynamically Reconfigurable Architecture Aimed at Reducing the FPGA-ASIC Gaps
abstract
Prior work has shown that due to the overhead incurred in enabling reconfigurability, field-programmable gate arrays (FPGAs) require 21× more silicon area, 3× larger delay, and 10× more dynamic power consumption compared with application-specific integrated circuits (ASICs). We have earlier presented a hybrid CMOS/nanotechnology reconfigurable architecture (NATURE). It uses the concept of temporal logic folding and fine-grain (i.e., cycle-level) dynamic reconfiguration to increase logic density by an order of magnitude. Since logic folding reduces area usage significantly, on-chip communications tend to become localized. To take full advantage of this fact, we propose a new architecture, called fine-grain dynamically reconfigurable (FDR), that consists of an array of homogeneous reconfigurable logic elements (LEs). Each LE can be arbitrarily configured into a lookup table (LUT) or interconnect or a combination of both. This significantly enhances the flexibility of allocating hardware resources between LUTs and interconnects based on application needs. The proposed FDR architecture eliminates most of the long-distance and global wires, which occupy most of the area in conventional FPGAs. Fine-grain dynamic reconfiguration is enabled by local embedded static RAM blocks. The experiments show that, on an average, area, delay, and power are improved by 9.14×, 1.11×, and 1.45×, compared with a conventional FPGA architecture that does not use the concept of logic folding. Compared with NATURE with deep logic folding, area, delay, and power are improved by 2.12×, 3.28×, and 1.74×, respectively. Although this does not eliminate the FPGA-ASIC area/delay/power gaps, it makes progress toward bridging these gaps.
Ting-Jung Lin, Wei Zhang 0012, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
2014 FinPrin: FinFET Logic Circuit Analysis and Optimization Under PVT Variations
abstract
Continued scaling of bulk CMOS technology is facing formidable challenges. FinFETs, with better control of short-channel effects, offer a promising alternative for the 22-nm technology node and beyond. However, FinFETs still suffer from process, voltage, and temperature (PVT) variations. Thus, to analyze the delay of FinFET logic circuits, statistical static timing analysis (SSTA) is more suitable than traditional static timing analysis. In this paper, using an existing SSTA algorithm as a foundation, we analyze silicon-on-insulator FinFET circuits in various new ways: basing the analysis on accurate device simulation of the logic library; considering voltage and temperature variations, in addition to process variations; deriving accurate timing and leakage macromodels; and investigating the impact of PVT variations on delay/leakage distributions at the circuit level. We propose a simplified timing model that greatly reduces the computational complexity without giving rise to any convergence issues. The timing model has an average absolute error of 3.4% and 4.4%, respectively, for gate output slope and gate delay over all logic gates and sizes, compared with accurate quasi-Monte Carlo simulations. We evaluate the performance of our SSTA algorithm with respect to Monte Carlo simulation, and extend the algorithm to enable statistical leakage and dynamic power analysis as well. We investigate the impact of PVT variations on delay/power distributions at the circuit level. We show that deterministic optimization methods can optimize both the mean and variance of circuit delay/power distributions, and even the ratio of standard deviation to mean in some cases. Finally, we show that FinFET circuits need to be carefully optimized with temperature taken into consideration, since the ratio between the leakage and dynamic power of a circuit can vary drastically depending on the operating temperature assumed.
Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Towards trustworthy medical devices and body area networks
abstract
Implantable and wearable medical devices (IWMDs) are commonly used for diagnosing, monitoring, and treating various medical conditions. A general trend in IWMDs is towards increased functional complexity, software programmability, and connectivity to body area networks (BANs). However, as medical devices become more "intelligent," they also become less trustworthy -- less reliable and more vulnerable to malicious attacks. Various shortcomings -- hardware failures, software errors, wireless attacks, malware and software exploits, and side-channel attacks -- could undermine the trustworthiness of IWMDs and BANs. The trustworthiness of IWMDs must be addressed aggressively and proactively due to the potential for catastrophic consequences. While some recent efforts address the defense of IWMDs against specific security attacks, a holistic strategy that considers all concerns and types of threats is required. This paper discusses trustworthiness concerns in IWMDs and BANs through a comprehensive identification and analysis of potential threats and, for each threat, provides a discussion of the merits and inadequacies of current solutions.
Anand Raghunathan, Niraj K. Jha
DAC3
2013 Thermal Characterization of Test Techniques for FinFET and 3D Integrated Circuits
abstract
Power consumption has become a very important consideration during integrated circuit (IC) design and test. During test, it can far exceed the values reached during normal operation and, thus, lead to temperatures above the allowed threshold. Without appropriate temperature reduction, permanent damage may be caused to the IC or invalid test results may be obtained. FinFET is a double-gate field-effect transistor (DG-FET) that was introduced commercially in 2012. Due to the vertical nature of FinFETs and, hence, weaker ability to dissipate heat, this problem is likely to get worse for FinFET circuits. Another technology rapidly gaining popularity is 3D IC integration. Unfortunately, the compact nature of a multidie 3D IC is likely to aggravate the temperature-during-test problem even further. Hence, before temperature-aware test methodologies can be developed, it is important to thermally analyze both FinFET and 3D circuits under test. In this article, we present a methodology for thermal characterization of various test techniques, such as scan and built-in self-test (BIST), for FinFET and 3D ICs. FinFET thermal characterization makes use of a FinFET standard cell library that is characterized with the help of the University of Florida double-gate (UFDG) SPICE model. Thermal profiles for circuits under test are produced by ISAC2 from University of Colorado for FinFET circuits and HotSpot from University of Virginia for 3D ICs. Experimental results indicate that high temperatures result under BIST and much less often under scan, and that both power consumption and test application time should be reduced to lower the temperature of circuits under test, just reducing the power consumption is not enough.
Aoxiang Tang, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2013 Design space exploration of FinFET cache
abstract
Integration of cache on-chip has significantly improved the performance of modern processors. The relentless demand for ever-increasing performance has led to the need to increase the cache capacity and number of cache levels. However, the performance improvement is accompanied by an increase in chip's power dissipation, requiring the use of more expensive cooling technologies to ensure chip reliability and long product life. The emergence of FinFETs as the technology of choice for high-performance computing poses new challenges to processor designers. With the introduction of new features in FinFETs, for example, independently controllable back gates, researchers have proposed several innovative memory cells that can reduce leakage power significantly, making the integration of a larger cache more practical. In this article, we comprehensively evaluate and compare the performance, power consumption (both dynamic and leakage), area, and temperature of different FinFET SRAM caches by exploring common configurations with varying cache size, block size, associativity, and number of banks. We evaluate caches based on four well-known FinFET SRAM cells: Pass-Gate FeedBack (PGFB), Row-based Back-Gate Biasing (RBGB), 8T, and 4T. We show how the caches can be simulated at self-consistent temperatures (at which leakage and temperature are in equilibrium). Drowsy and decay caches are two well-known leakage reduction techniques. We implement them in the context of FinFET caches to investigate their impact. We show that the RBGB cell-based cache is far superior in leakage and Power-Delay Product (PDP) to those based on the other three cells, sometimes by an order of magnitude. This superiority is maintained even when drowsy or decay leakage reduction techniques are applied to caches based on the other three cells, but not to the one based on the RBGB cell. This significantly diminishes the importance of drowsy or decay cache techniques, at least when the RBGB cell is used.
Aoxiang Tang, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2013 Efficient Methodologies for 3-D TCAD Modeling of Emerging Devices and Circuits
abstract
Over the past decade, 3-D process simulation, which is central to the 3-D Technology Computer-Aided Design (3-D TCAD) approach, has severely limited the scope and applicability of TCAD to circuits with a small number of field-effect transistors, owing to its prohibitively high computational costs for large layouts. Due to rapidly changing process recipes and shorter production cycles in the industry, design-time optimization and iterative layout-3-D TCAD exploration for yield-critical or yield-characterizing circuits, such as static random-access memories (SRAMs), ring oscillators, and others, is currently impossible in a practical time frame. In this paper, we architect a novel layout/process/device-independent TCAD methodology in the Sentaurus tool suite to overcome the process simulation barrier for accurate 3-D TCAD structure generation. We adopt an automated structure synthesis (SS) approach, thereby bypassing the need for repetitive 3-D process simulations for different layouts or different versions of the same layout. Results for 32-nm bulk process simulations versus SS and 32-nm silicon-on-insulator (SOI) hardware measurements versus corresponding synthesized structures indicate that the method is an excellent substitute to 3-D process simulation of large layouts, with extremely favorable time and memory scaling behavior. Finally, the robustness and scalability of the proposed abstractions are highlighted through the synthesis of 22-nm SOI 6T FinFET SRAMs and ring oscillator structures.
Ajay N. Bhoj, Rajiv V. Joshi, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2013 Design of Logic Gates and Flip-Flops in High-Performance FinFET Technology
abstract
With the emergence of nonplanar CMOS devices at the 22-nm node and beyond, it is highly likely that multigate device adoption will occur in a high-performance process technology, owing to the increased performance and area benefits. In this paper, for the first time, we evaluate symmetric (Symm-ΦG) and asymmetric (Asymm-ΦG) gate-workfunction FinFETs head to head in a high-performance process, using technology computer-aided design 3-D device simulations. We demonstrate that Asymm-ΦGshorted-gate (a-SG) n/p-FinFETs, which use both workfunctions corresponding to typical high-performance metal-gate n/p-FinFETs, are promising, as they can yield over two orders of magnitude lower leakage without excessive degradation in ON-state current, in comparison to Symm- ΦG shorted-gate (SG) FinFETs, placing them in a better position than back-gate biased independent-gate (IG) FinFETs for leakage reduction. Thereafter, we explore the design space of FinFET logic gates, latches, and flip-flops, for optimal tradeoffs in leakage versus delay and temperature, using mixed-mode 2-D device simulations. Elementary logic gates (such as INV, NAND2, NOR2, XOR2, and XNOR2) using Asymm-ΦGSG-mode FinFETs appear to be located optimally in the leakage-delay spectrum, in comparison to the most versatile configurations possible by mixing corresponding Symm-ΦGSG- and IG-mode FinFETs. Latches and flip-flops, however, require an astute combination of Symm-ΦGand Asymm-ΦGFinFETs to optimize leakage, delay, and setup time simultaneously.
Ajay N. Bhoj, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2013 3-D-TCAD-Based Parasitic Capacitance Extraction for Emerging Multigate Devices and Circuits
abstract
In recent years, the multigate field-effect transistor (FET) has emerged as the most viable contender for technology scaling down to the sub-10-nm nodes. The nonplanar nature of multigate devices, along with rapidly shrinking front-end-of-line (FEOL) and back-end-of-line (BEOL) features, has compounded the problem of parasitics extraction in future technology nodes. In this paper, for the first time, we address the above problem through a holistic 3-D-technology CAD (3-D-TCAD) flow for the extraction of FEOL/(FEOL+BEOL) capacitances in generic multigate circuit layouts, using a transport analysis-based approach. We investigate device-level parasitic capacitances in 3-D-process-simulated bulk and silicon-on-insulator FinFETs, and uncover capacitance scaling trends for candidate single/multifin multigate FETs along the 22-nm/14-nm/10-nm technology nodes. Leveraging automated structure synthesis algorithms, we synthesize 3-D multigate 6T SRAM structures using the process-simulated devices, and examine the effects of fin pitch, gate pitch, and fin count on circuit-level parasitics. Thereafter, we show that traditional segregated FEOL/BEOL modeling approaches fail to provide accurate estimates, by back-annotating 3-D-TCAD-extracted capacitances into mixed-mode write simulations of a 6T FinFET SRAM bitcell. Finally, using FinFET NAND2 logic gate delay simulations, we establish the fact that capturing parasitics accurately is as important as modeling device transport accurately, and that performance/dynamic behavior in multigate circuits is highly sensitive to both factors.
Ajay N. Bhoj, Rajiv V. Joshi, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
2013 Variable-Pipeline-Stage Router
abstract
On-chip interconnection networks are fast becoming significant power consumers in high-performance chip multiprocessors. Increased power consumption leads to more heat, degrades system reliability, and may increase the cost of cooling integrated circuit packages. This situation is becoming worse as bulk CMOS technology scales further into the nanometer regime because of the excessive leakage power caused by short-channel effects. In this paper, we explore the use of FinFETs, which are promising substitutes for bulk CMOS at the 22-nm node and beyond, to design on-chip network routers. We present a detailed design of a variable-pipeline-stage router (VPSR) targeted at FinFET technology. We employ a dynamic power management scheme, which we call adaptive back-gate biasing, for FinFET implementations. We propose enhanced token flow control (ETFC), a flow control mechanism that improves upon the energy/delay/throughput of the previous state-of-the-art token flow control mechanism. We evaluate VPSR and ETFC on a simulation platform specifically designed for power/performance simulations of FinFET-based interconnection networks. The results show that VPSR is able to successfully adapt its power consumption to incoming traffic, with a resultant 21.5% reduction in power with almost no impact on latency.
Chun-Yi Lee, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Optimized Quantum Gate Library for Various Physical Machine Descriptions
abstract
Quantum logic circuits consist of a cascade of quantum gates. These gates are realized using primitive quantum operations that are supported by a quantum physical machine description (PMD). Since different quantum systems are associated with different Hamiltonians, a specific quantum operation may be more easily realizable in one quantum system than another. Thus, different quantum systems have different PMDs. Also, the quantum cost for implementing a quantum operation may differ from one PMD to another. Thus, a quantum logic circuit needs to be realized with and optimized for only the set of primitive quantum operations supported by the given PMD. Quantum logic design that can be targeted at multiple PMDs has not been attempted before, to the best of our knowledge. In this paper, we target quantum logic design with respect to the set of primitive quantum operations that are supported by six different PMDs: quantum dot, superconducting, ion trap, neutral atom, and two photonics systems. Our aim is to build a quantum gate library that targets these PMDs. This is akin to a cell library in traditional logic design that enables logic gates to be mapped to cells realizable in an underlying technology. To make our quantum gate library efficient in terms of the number of primitive quantum operations involved and the associated delay, we explore one- and two-qubit quantum identity rules that can help remove redundancies in the quantum gate implementation. We show that, using these identities, each gate in the library can be efficiently mapped to just the set of primitive operations supported by each of the six PMDs. Each mapping results in a different circuit structure and quantum cost, which is measured in terms of the number of primitive quantum operations and the number of execution cycles required. Thus, such a library provides the foundation for quantum logic synthesis, just like a cell library provides the foundation for technology-dependent logic synthesis (i.e., technology mapping) in traditional synthesis.
Chia-Chun Lin, Amlan Chakrabarti, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
2013 Algorithm-Driven Architectural Design Space Exploration of Domain-Specific Medical-Sensor Processors
abstract
Data-driven machine-learning techniques enable the modeling and interpretation of complex physiological signals. The energy consumption of these techniques, however, can be excessive, due to the complexity of the models required. In this paper, we study the tradeoffs and limitations imposed by the energy consumption of high-order detection models implemented in devices designed for intelligent biomedical sensing. Based on the flexibility and efficiency needs at various processing stages in data-driven biomedical algorithms, we explore options for hardware specialization through architectures based on custom instruction and coprocessor computations. We identify the limitations in the former, and propose a coprocessor-based platform that exploits parallelism in computation as well as voltage scaling to operate at a subthreshold minimum-energy point. We present results from post-layout simulation of cardiac arrhythmia detection with patient data from the MIT-BIH database. After wavelet-based feature extraction, which consumes 12.28 μJ, we demonstrate classification computations in the 12.00-120.05 μJ range using 10000-100000 support vectors. This represents 1170× lower energy than that of a low-power processor with custom instructions alone. After morphological feature extraction, which consumes 8.65 μJ of energy, the corresponding energy numbers are 10.24-24.51 μJ, which is 1548× smaller than one based on a custom-instruction design. Results correspond to Vdd=0.4 V and a data precision of 8 b.
Mohammed Shoaib, Niraj K. Jha, Naveen Verma
IEEE Trans. Very Large Scale Integr. Syst.2
2012 Enabling advanced inference on sensor nodes through direct use of compressively-sensed signals
abstract
Nowadays, sensor networks are being used to monitor increasingly complex physical systems, necessitating advanced signal analysis capabilities as well as the ability to handle large amounts of network data. For the first time, we present a methodology to enable advanced decision support on a low-power sensor node through the direct use of compressively-sensed signals in a supervised-learning framework; such signals provide a highly efficient means of representing data in the network, and their direct use overcomes the need for energy-intensive signal reconstruction. Sensor networks for advanced patient monitoring are representative of the complexities involved. We demonstrate our technique on a patient-specific seizure detection algorithm based on electroencephalograph (EEG) sensing. Using data from 21 patients in the CHB-MIT database, our approach demonstrates an overall detection sensitivity, latency, and false alarm rate of 94.70%, 5.83 seconds, and 0.199 per hour, respectively, while achieving data compression by a factor of 10x. This compares well with the state-of-the-art baseline detector with corresponding results being 96.02%, 4.59 seconds, and 0.145 per hour, respectively.
Mohammed Shoaib, Niraj K. Jha, Naveen Verma
DATE2
2012 INVISIOS: A Lightweight, Minimally Intrusive Secure Execution Environment
abstract
Many information security attacks exploit vulnerabilities in “trusted” and privileged software executing on the system, such as the operating system (OS). On the other hand, most security mechanisms provide no immunity to security-critical user applications if vulnerabilities are present in the underlying OS. While technologies have been proposed that facilitate isolation of security-critical software, they require either significant computational resources and are hence not applicable to many resource-constrained embedded systems, or necessitate extensive redesign of the underlying processors and hardware. In this work, we propose INVISIOS: a lightweight, minimally intrusive hardware-software architecture to make the execution of security-critical software invisible to the OS, and hence protected from its vulnerabilities. The INVISIOS software architecture encapsulates the security-critical software into a self-contained software module. While this module is part of the kernel and is run with kernel-level privileges, its code, data, and execution are transparent to and protected from the rest of the kernel. The INVISIOS hardware architecture consists of simple add-on hardware components that are responsible for bootstrapping the secure core, ensuring that it is exercised by applications in only permitted ways, and enforcing the isolation of its code and data. We implemented INVISIOS by enhancing a full-system emulator and Linux to model the proposed software and hardware enhancements, and applied it to protect a commercial cryptographic library. Our experiments demonstrate that INVISIOS is capable of facilitating secure execution at very small overheads, making it suitable for resource-constrained embedded systems and systems-on-chip.
Divya Arora 0001, Najwa Aaraj, Anand Raghunathan, Niraj K. Jha
ACM Trans. Embed. Comput. Syst.4
2012 Secure reconfiguration of software-defined radio
abstract
Software-defined radio (SDR) implements a radio system in software that executes on a programmable processor. The components of SDR, such as the filters, amplifiers, and modulators, can be easily reconfigured to adapt to the operating environment and user preferences. However, the flexibility of radio reconfiguration brings along the serious security concern of malicious modification of software in the SDR system, leading to radio malfunction and interference with other users' communications. Both the SDR device and the network need to be protected from such malicious radio reconfiguration. In this article, a new architecture targeted at protecting SDR devices from malicious reconfiguration is proposed. The architecture is based on robust separation of the radio operation environment and user application environment, through the use of virtualization. A new radio middleware layer is designed to securely intercept all attempts to reconfigure the radio, and a security policy monitor checks the target configuration against security policies that represent the interests of various parties. Even if the operating system in the user application environment is compromised, the proposed architecture can ensure secure reconfiguration in the radio operation environment. We have prototyped the proposed secure SDR architecture using VMware and the GNU Radio toolkit and demonstrate that overheads incurred by the architecture are small and tolerable. Therefore, we believe that the proposed solution could be applied to address secure SDR reconfiguration in both general-purpose and embedded computing systems.
Chunxiao Li 0007, Niraj K. Jha, Anand Raghunathan
ACM Trans. Embed. Comput. Syst.2
2012 A Trusted Virtual Machine in an Untrusted Management Environment
abstract
Virtualization is a rapidly evolving technology that can be used to provide a range of benefits to computing systems, including improved resource utilization, software portability, and reliability. Virtualization also has the potential to enhance security by providing isolated execution environments for different applications that require different levels of security. For security-critical applications, it is highly desirable to have a small trusted computing base (TCB), since it minimizes the surface of attacks that could jeopardize the security of the entire system. In traditional virtualization architectures, the TCB for an application includes not only the hardware and the virtual machine monitor (VMM), but also the whole management operating system (OS) that contains the device drivers and virtual machine (VM) management functionality. For many applications, it is not acceptable to trust this management OS, due to its large code base and abundance of vulnerabilities. For example, consider the "computing-as-a-service” scenario where remote users execute a guest OS and applications inside a VM on a remote computing platform. It would be preferable for many users to utilize such a computing service without being forced to trust the management OS on the remote platform. In this paper, we address the problem of providing a secure execution environment on a virtualized computing platform under the assumption of an untrusted management OS. We propose a secure virtualization architecture that provides a secure runtime environment, network interface, and secondary storage for a guest VM. The proposed architecture significantly reduces the TCB of security-critical guest VMs, leading to improved security in an untrusted management environment. We have implemented a prototype of the proposed approach using the Xen virtualization system, and demonstrated how it can be used to facilitate secure remote computing services. We evaluate the performance penalties incurred by the proposed architecture, and demonstrate that the penalties are minimal.
Chunxiao Li 0007, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Serv. Comput.3
2012 SRAM-Based NATURE: A Dynamically Reconfigurable FPGA Based on 10T Low-Power SRAMs
abstract
We presented a hybrid CMOS/nanotechnology reconfigurable architecture (NATURE), earlier. It was based on CMOS logic and nano RAMs. It used the concept of temporal logic folding and fine-grain (e.g., cycle-level) dynamic reconfiguration to increase logic density by an order of magnitude. This dynamic reconfiguration is done intra-circuit rather than inter-circuit. However, the previous design of NATURE required fine-grained distribution of nano RAMs throughout the field-programmable gate array (FPGA) architecture. Since the fabrication process of nano RAMs is not mature yet, this prevents immediate exploitation of NATURE. In this paper, we present a NATURE architecture that is based on CMOS logic and CMOS SRAMs that are used for on-chip dynamic reconfiguration. We use fast and low-power SRAM blocks that are based on 10T SRAM cells. We have also laid out the various FPGA components in a 65-nm technology to evaluate the FPGA performance. We hide the dynamic reconfiguration delay behind the computation delay through the use of shadow SRAM cells. Experimental results show more than an order of magnitude improvement in logic density and improvement in the area-delay product relative to a traditional baseline FPGA architecture that does not use the concept of logic folding.
Ting-Jung Lin, Wei Zhang 0012, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
2011 CACTI-FinFET: an integrated delay and power modeling framework for FinFET-based caches under process variations
abstract
We present CACTI-FinFET, an integrated framework for simulation of power, delay, temperature, as well as process variations of FinFET-based caches. We have developed a FinFET design library and process variation models to characterize the delay and leakage spreads of such caches. We present results for various FinFET design styles and show that mixing different design styles may be a promising strategy for optimizing cache delay and leakage.
Chun-Yi Lee, Niraj K. Jha
DAC2
2011 A low-energy computation platform for data-driven biomedical monitoring algorithms
abstract
A key challenge in closed-loop chronic biomedical systems is the ability to detect complex physiological states from patient signals within a constrained power budget. Data-driven machine-learning techniques are major enablers for the modeling and interpretation of such states. Their computational energy, however, scales with the complexity of the required models. In this paper, we propose a low-energy, biomedical computation platform optimized through the use of an accelerator for data-driven classification. The accelerator retains selective flexibility through hardware reconfiguration and exploits voltage scaling and parallelism to operate at a sub-threshold minimum-energy point. Using cardiac arrhythmia detection algorithms with patient data from the MIT-BIH database, classification is achieved in 2.96 μJ (at Vdd = 0.4 V), over four orders of magnitude smaller than that on a low-power general-purpose processor. The energy of feature extraction is 148 μJ while retaining flexibility for a range of possible biomarkers.
Mohammed Shoaib, Niraj K. Jha, Naveen Verma
DAC2
2011 3D vs. 2D analysis of FinFET logic gates under process variations
abstract
Among various multi-gate structures, FinFETs have emerged dominant owing to their ease of fabrication. Thus, characterization of FinFET devices/gates needs immediate attention for them to become the industry driver in this decade. Ideally, 3D device simulation should be done to enable accurate circuit synthesis. However, this is impractical due to the huge CPU times required. Simulating a 2D cross-section of the device yields 100-1000× reduction in CPU time. However, this introduces significant error, in the range of 10% to 50%, while evaluating the on/off current (ION/IOFF) for a single device and leakage current or propagation delay (ILEAK/tD) for logic gates. In this work, we develop accurate 2D models of FinFET devices to capture 3D simulation accuracy with 2D simulation efficiency. We report results for the 22nm FinFET technology node. As far as we know, this is the first such attempt. We establish the validity of the model even under process variations. We target variations in gate length (LG), workfunction (ΦG) and fin thickness (TSI) that are known to have the most impact on leakage and delay. We adjust their values in the 2D model in order to mimic the actual 3D device behavior. When the 2D models are employed in mixed-mode simulation of FinFET logic gates, the error in the evaluation of ILEAK/tDis quite small.
Sourindra Chaudhuri, Niraj K. Jha
ICCD2
2011 FinFET-Based Power Management for Improved DPA Resistance with Low Overhead
abstract
Differential power analysis (DPA) is a side-channel attack that statistically analyzes the power consumption of a cryptographic system to obtain secret information. This type of attack is well known as a major threat to information security. Effective solutions with low energy and area cost for improved DPA resistance are urgently needed, especially for energy-constrained modern devices that are often in the physical proximity of attackers. This article presents a novel countermeasure against DPA attacks on smart cards and other digital ICs based on FinFETs, an emerging substitute for bulk CMOS at the 22nm technology node and beyond. We exploit the adaptive power management characteristic of FinFETs to generate a high level of noise at critical moments in the execution of a cryptosystem to thwart DPA attacks. We demonstrate the effectiveness of the proposed countermeasure by developing a simple power model for estimating DPA spikes. We then validate the model by carrying out DPA attacks on an ASIC implementation of the advanced encryption standard system via gate-level simulation. Both modeling and simulation-based experiment indicate that with the proposed countermeasure, even 8,000,000 power acquisitions are not sufficient to reveal the secret key. As opposed to other countermeasures presented in the literature, the proposed hardware design requires less than 1% increase in area and 15% increase in total energy consumption without any extra delay in the critical path. The proposed method is generic and can be applied to other encryption algorithms as well.
Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2011 A framework for defending embedded systems against software attacks
abstract
The incidence of malicious code and software vulnerability exploits on embedded platforms is constantly on the rise. Yet, little effort is being devoted to combating such threats to embedded systems. Moreover, adapting security approaches designed for general-purpose systems generally fails because of the limited processing capabilities of their embedded counterparts. In this work, we evaluate a malware and software vulnerability exploit defense framework for embedded systems. The proposed framework extends our prior work, which defines two isolated execution environments: a testing environment, wherein an untrusted application is first tested using dynamic binary instrumentation (DBI), and a real environment, wherein a program is monitored at runtime using an extracted behavioral model, along with a continuous learning process. We present a suite of software and hardware optimizations to reduce the overheads induced by the defense framework on embedded systems. Software optimizations include the usage of static analysis, complemented with DBI in the testing environment (i.e., a hybrid software analysis approach is used). Hardware optimizations exploit parallel processing capabilities of multiprocessor systems-on-chip. We have evaluated the defense framework and proposed optimizations on the ARM-Linux operating system. Experiments demonstrate that our framework achieves a high coverage of considered security threats, with acceptable performance penalties (the average execution time of applications goes up to 1.68X, considering all optimizations, which is much smaller than the 2.72X performance penalty when no optimizations are used).
Najwa Aaraj, Anand Raghunathan, Niraj K. Jha
ACM Trans. Embed. Comput. Syst.3
2011 Editorial Announcing a New Editor-in-Chief
Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.1
2010 Secure Virtual Machine Execution under an Untrusted Management OS
abstract
Virtualization is a rapidly evolving technology that can be used to provide a range of benefits to computing systems, including improved resource utilization, software portability, and reliability. For security-critical applications, it is highly desirable to have a small trusted computing base (TCB), since it minimizes the surface of attacks that could jeopardize the security of the entire system. In traditional virtualization architectures, the TCB for an application includes not only the hardware and the virtual machine monitor (VMM), but also the whole management operating system (OS) that contains the device drivers and virtual machine (VM) management functionality. For many applications, it is not acceptable to trust this management OS, due to its large code base and abundance of vulnerabilities. In this paper, we address the problem of providing a secure execution environment on a virtualized computing platform under the assumption of an untrusted management OS. We propose a secure virtualization architecture that provides a secure run-time environment, network interface, and secondary storage for a guest VM. The proposed architecture significantly reduces the TCB of security-critical guest VMs, leading to improved security in an untrusted management environment. We have implemented a prototype of the proposed approach using the Xen virtualization system, and demonstrated how it can be used to facilitate secure remote computing services. We evaluate the performance penalties incurred by the proposed architecture, and demonstrate that the penalties are minimal.
Chunxiao Li 0007, Anand Raghunathan, Niraj K. Jha
IEEE CLOUD3
2010 Low-power FinFET circuit synthesis using surface orientation optimization
abstract
FinFETs with channel surface along theplane can be easily fabricated by rotating the fins by 45° from theplane. By designing logic gates, which have pFinFETs in theplane and nFinFETs in theplane, the gate delay can be reduced by as much as 14%, compared to the conventionallogic gates. The reduction in delay can be traded off for reduced power in FinFET circuits. In this paper, we propose a low-power FinFET-based circuit synthesis methodology based on surface orientation optimization. We study various logic design styles, which depend on different FinFET channel orientations, for synthesizing low-power circuits. We use BSIM, a process/physics based double-gate model in HSPICE, to derive accurate delay and power estimates. We design layouts of standard library cells containing FinFETs in different orientations to obtain an accurate area estimate for the low-power synthesized netlists after place-and-route. We use a linear programming based optimization methodology that gives power-optimized netlists, consisting of oriented gates, at tight delay constraints. Experimental results demonstrate the efficacy of our scheme.
Prateek Mishra, Niraj K. Jha
DATE2
2010 Gated-diode FinFET DRAMs: Device and circuit design-considerations
abstract
Scaling bulk CMOS SRAM technology for on-chip caches beyond the 22nm node is questionable, on account of high leakage power consumption, performance degradation, and instability due to process variations. Recently, two-three transistor one gated-diode (2T/3T1D) DRAMs were proposed as alternatives to address the SRAM variability problem, with an emphasis on high-activity embedded cache applications. They are highly competitive to SRAM in terms of performance, while having a smaller power and area footprint at lower technology nodes. The current evolutionary trend in transistor structures is moving toward an era of multigate devices, which makes it necessary to identify design issues and advantages of gated-diode DRAMs implemented in a multigate technology. In this work, we address gated-diode DRAM design in FinFET technology using mixed-mode 2D-device simulations. We revisit the model of internal voltage gain in bulk gated diodes and extend it to provide quantitative insight into designing Fin gated diodes, that is, gated diodes in FinFET technology. To this effect, we propose FinFET variants of the bulk gated-diode configuration and identify parameters that are critical to enhancing the retention time and read current in 2T/3T1D FinFET DRAMs. Additionally, we show the superiority of 2T1D FinFET DRAM over 6T FinFET SRAM having pass-gate feedback (6T PGFB) and 2T1D bulk DRAM under the effect of physical parameter process variations using a quasi--Monte Carlo method implemented in FinE, an environment we have developed for double-gate circuit design that integrates Sentaurus TCAD from Synopsys with the Spice3-UFDG double-gate compact model from University of Florida, under a single framework. Finally, we present a new tunable threshold gated diode FinFET amplifier which uses an n-type gated diode for voltage-boosting, along with a p-type gated diode for zero-suppression.
Ajay N. Bhoj, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2010 FinFET-based power simulator for interconnection networks
abstract
Double-gate FETs, specifically FinFETs, are emerging as promising substitutes for bulk CMOS at the 32nm technology node and beyond because of the various obstacles to scaling faced by CMOS, such as short-channel effects, leakage power, and process variations. Another trend in chip multiprocessor design is incorporation of sophisticated on-chip interconnection networks. However, such networks are significant power-consumers. In this article, we address these two trends by presenting a power simulator for FinFET-based on-chip interconnection networks. It estimates both dynamic and leakage power. We present results for various FinFET design styles and temperatures (since leakage power changes drastically with temperature), and show that one FinFET design style may be much superior to another from the power consumption point of view.
Chun-Yi Lee, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2010 Low-power 3D nano/CMOS hybrid dynamically reconfigurable architecture
abstract
In order to continue technology scaling beyond CMOS, diverse nanoarchitectures have been proposed in recent years based on emerging nanodevices, such as nanotubes, nanowires, etc. Among them, some hybrid nano/CMOS reconfigurable architectures enjoy the advantage that they can be fabricated using photolithography. NATURE is one such architecture that we have proposed recently. It comprises CMOS reconfigurable logic and CMOS fabrication-compatible nano RAMs. It uses distributed high-density and fast nano RAMs as on-chip storage for storing multiple reconfiguration copies, enabling fine-grain cycle-by-cycle reconfiguration. It supports a highly efficient computational model, called temporal logic folding, which makes possible more than an order of magnitude improvement in logic density and area-delay product, significant power reduction, and significant design flexibility in performing area-delay trade-offs. In this article, we extend NATURE in various dimensions, evaluating various FPGA approaches in the context of today's emerging technologies. First, we explore the introduction of embedded coarse-grain modules in the fine-grain NATURE architecture and present a unified dynamically reconfigurable architecture, which can significantly enhance NATURE's computation power for data-dominated applications. Second, we explore a 3D architecture for NATURE in which the nano RAM for reconfiguration storage is on one layer and the rest of the CMOS logic on another layer. This leads to further improvements in logic density and performance. Finally, we explore the possibility of using FinFETs, an emerging double-gate CMOS technology, to implement NATURE. Since power consumption is an important consideration in the deep nanometer regime, especially for FPGAs, we present a back-gate biasing methodology for flexible threshold voltage adjustment in FinFETs to significantly reduce NATURE's power consumption. Simulation results demonstrate the efficacy of the proposed methods.
Wei Zhang 0012, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2010 Editorial: New Associate Editor Appointments
Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.1
2009 An architecture for secure software defined radio
abstract
Software defined radio (SDR) is a rapidly evolving technology which implements some functional modules of a radio system in software executing on a programmable processor. SDR provides a flexible mechanism to reconfigure the radio, enabling networked devices to easily adapt to user preferences and the operating environment. However, the very mechanisms that provide the ability to reconfigure the radio through software also give rise to serious security concerns such as unauthorized modification of the software, leading to radio malfunction and interference with other users' communications. Both the SDR device and the network need to be protected from such malicious radio reconfiguration. In this paper, we propose a new architecture to protect SDR devices from malicious reconfiguration. The proposed architecture is based on robust separation of the radio operation environment and user application environment through the use of virtualization. A secure radio middleware layer is used to intercept all attempts to reconfigure the radio, and a security policy monitor checks the target configuration against security policies that represent the interests of various parties. Therefore, secure reconfiguration can be ensured in the radio operation environment even if the operating system in the user application environment is compromised. We have prototyped the proposed secure SDR architecture using VMware and the GNU Radio toolkit, and demonstrate that the overheads incurred by the architecture are small and tolerable. Therefore, we believe that the proposed solution could be applied to address SDR security concerns in a wide range of both general-purpose and embedded computing systems.
Chunxiao Li 0007, Anand Raghunathan, Niraj K. Jha
DATE3
2009 In-Network Snoop Ordering (INSO): Snoopy coherence on unordered interconnects
abstract
Realizing scalable cache coherence in the many-core era comes with a whole new set of constraints and opportunities. It is widely believed that multi-hop, unordered on-chip networks would be needed in many-core chip multiprocessors (CMPs) to provide scalable on-chip communication. However, providing ordering among coherence transactions on unordered interconnects is a challenge. Traditional approaches for tackling coherence either have to use ordered interconnects (snoop-based protocols) which lead to scalability problems, or rely on an ordering point (directory-based protocols) which adds indirection latency. In this paper, we propose In-Network Snoop Ordering (INSO), in which coherence requests from a snoop-based protocol are inserted into the interconnect fabric and the network orders the requests in a distributed manner, creating a global ordering among requests. Essentially, when coherence requests enter the network, they grab snoop-orders at the injection router before being broadcasted. A snoop-order specifies the global ordering of the particular request with respect to other requests. Before requests reach their destinations, they get ordered along the way, at intermediate routers and destination network interfaces. Our logical ordering scheme can be mapped onto any unordered interconnect. This enables a cache coherence protocol which exploits the low-latency nature of unordered interconnects without adding indirection to coherence transactions. Our full-system evaluations compare INSO against a directory protocol and a broadcast based Token Coherence protocol. INSO outperforms these protocols by up to 30% and 8.5%, respectively, on a wide range of scientific and emerging applications.
Niket Agarwal, Li-Shiuan Peh, Niraj K. Jha
HPCA3
2009 Pragmatic design of gated-diode FinFET DRAMs
abstract
Scaling bulk CMOS SRAM technology for on-chip caches beyond the 22 nm node is questionable, on account of high leakage power consumption, performance degradation, and instability due to process variations. Recently, two/three transistor one gated-diode (2T/3T1D) DRAMs were proposed as alternatives to address the SRAM variability problem, with an emphasis on high-activity embedded cache applications. They are highly competitive with an SRAM in terms of performance, while having a smaller power and area footprint at lower technology nodes. The current evolutionary trend in transistor structures is toward an era of multi-gate devices, which makes it necessary to identify design issues and advantages of gated-diode DRAMs implemented in a multi-gate technology. In this work, we address gated-diode DRAM design in FinFET technology using mixed-mode 2D-device simulations. We revisit the model of internal voltage gain in bulk gated-diodes and extend it to provide quantitative insight into designing Fin gated-diodes, i.e., gated-diodes in FinFET technology. To this effect, we propose FinFET variants of the bulk gated-diode configuration and identify parameters that are critical to enhancing the retention time and read current in 2T/3T1D FinFET DRAMs. Additionally, we show the superiority of 2T1D FinFET DRAM over 6T FinFET SRAM having pass-gate feedback (6T PGFB) and 2T1D bulk DRAM under the effect of variations using a quasi-Monte Carlo method implemented in FinE, an environment we have developed for double-gate circuit design that integrates Sentaurus TCAD from Synopsys with the Spice3-UFDG double-gate compact model from University of Florida under a single framework. Finally, we present a new tunable threshold gated-diode FinFET amplifier which uses an n-type gated-diode for voltage-boosting, along with a p-type gated-diode for zero-suppression.
Ajay N. Bhoj, Niraj K. Jha
ICCD2
2009 FinFET-based dynamic power management of on-chip interconnection networks through adaptive back-gate biasing
abstract
On-chip interconnection networks are fast becoming significant power-consumers in high-performance chip multiprocessors (CMPs). Increased power consumption leads to more heat, adversely degrades system reliability, and may increase the cost of cooling IC packages. This situation becomes even worse as bulk CMOS scales further into the nanometer regime because of excessive leakage power due to short-channel effects. In this paper, we explore the use of FinFETs, which are promising substitutes for bulk CMOS at the 32 nm node and beyond, to design on-chip network routers. We present a detailed design of a variable pipeline stage router (VPSR) targeted at FinFET technology. We employ a dynamic power management scheme, which we call adaptive back-gate biasing (ABGB), for FinFET implementations. We evaluate VPSR and ABGB on a simulation platform specifically designed for power and performance simulations for FinFET-based interconnection networks. The results show that VPSR is able to successfully adapt its power consumption to incoming traffic, with a resultant 20% reduction in power at almost no impact on latency.
Chun-Yi Lee, Niraj K. Jha
ICCD2
2009 GARNET: A detailed on-chip network model inside a full-system simulator
abstract
Until very recently, microprocessor designs were computation-centric. On-chip communication was frequently ignored. This was because of fast, single-cycle on-chip communication. The interconnect power was also insignificant compared to the transistor power. With uniprocessor designs providing diminishing returns and the advent of chip multiprocessors (CMPs) in mainstream systems, the on-chip network that connects different processing cores has become a critical part of the design. Transistor miniaturization has led to high global wire delay, and interconnect power comparable to transistor power. CMP design proposals can no longer ignore the interaction between the memory hierarchy and the interconnection network that connects various elements. This necessitates a detailed and accurate interconnection network model within a full-system evaluation framework. Ignoring the interconnect details might lead to inaccurate results when simulating a CMP architecture. It also becomes important to analyze the impact of interconnection network optimization techniques on full system behavior. In this light, we developed a detailed cycle-accurate interconnection network model (GARNET), inside the GEMS full-system simulation framework. GARNET models a classic five-stage pipelined router with virtual channel (VC) flow control. Microarchitectural details, such as flit-level input buffers, routing logic, allocators and the crossbar switch, are modeled. GARNET, along with GEMS, provides a detailed and accurate memory system timing model. To demonstrate the importance and potential impact of GARNET, we evaluate a shared and private L2 CMP with a realistic state-of-the-art interconnection network against the original GEMS simple network. The objective of the evaluation was to figure out which configuration is better for a particular workload. We show that not modeling the interconnect in detail might lead to an incorrect outcome. We also evaluate Express Virtual Channels (EVCs), an on-chip network flow control proposal, in a full-system fashion. We show that in improving on-chip network latency-throughput, EVCs do lead to better overall system runtime, however, the impact varies widely across applications.
Niket Agarwal, Tushar Krishna, Li-Shiuan Peh, Niraj K. Jha
ISPASS4
2009 Thermal characterization of BIST, scan design and sequential test methodologies
abstract
It is a well known fact that during testing of a complex integrated circuit (IC), power consumption can far exceed the values reached during its normal operation. High power consumption, combined with limited cooling support, leads to overheating of ICs. This can cause permanent damage to the chip or can invalidate test results due to changes in the path delay. Therefore, even good chips can fail the test. To prevent this problem, a methodology to generate the thermal profile of chips during test is needed. If such profiles are provided beforehand, temperature-aware testing techniques can be devised. In this paper, we address this problem by presenting a methodology for thermally characterizing circuits under test. In our methodology, first, the test sequences for each targeted test strategy, namely, built-in self-test (BIST), scan design and sequential test generation, are generated automatically. Then, power profiles are extracted by using the switching activity information obtained from simulations. Finally, a very fast thermal profiling tool is used to produce the final thermal profiles. To the best of our knowledge, this is the first work on characterizing the thermal effects of different test methods. Such a thermal characterization can be leveraged for temperature-aware system-on-chip (SoC) test scheduling. Our experimental results present the maximum temperature values attained when using different testing techniques on several benchmarks. Results also demonstrate that low power testing techniques are not necessarily temperature-aware.
Muzaffer O. Simsir, Niraj K. Jha
ITC2
2009 In-network coherence filtering: snoopy coherence without broadcasts
abstract
With transistor miniaturization leading to an abundance of on-chip resources and uniprocessor designs providing diminishing returns, the industry has moved beyond single-core microprocessors and embraced the many-core wave. Scalable cache coherence protocol implementations are necessary to allow fast sharing of data among various cores and drive the many-core revolution forward. Snoopy coherence protocols, if realizable, have the desirable property of having low storage overhead and not adding indirection delay to cache-to-cache accesses. There are various proposals, like Token Coherence (TokenB), Uncorq, Intel QPI, INSO and Timestamp Snooping, that tackle the ordering of requests in snoopy protocols and make them realizable on unordered networks. However, snoopy protocols still have the broadcast overhead because each coherence request goes to all cores in the system. This has substantial network bandwidth and power implications. In this work, we propose embedding small in-network coherence filters inside on-chip routers that dynamically track sharing patterns among various cores. This sharing information is used to filter away redundant snoop requests that are traveling towards unshared cores. Filtering these useless messages saves network bandwidth and power and makes snoopy protocols on many-core systems truly scalable. Our in-network coherence filters are able to reduce the total number of snoops in the system on an average by 41.9%, thereby reducing total network traffic by 25.4% on 16-processor chip multiprocessor (CMP) systems running parallel applications. For 64-processor CMP systems, our filtering technique on an average achieves 46.5% reduction in total number of snoops that ends up reducing the total network traffic by 27.3%, on an average.
Niket Agarwal, Li-Shiuan Peh, Niraj K. Jha
MICRO3
2009 Low-power FinFET circuit synthesis using multiple supply and threshold voltages
abstract
According to Moore's law, the number of transistors in a chip doubles every 18 months. The increased transistor-count leads to increased power density. Thus, in modern circuits, power efficiency is a central determinant of circuit efficiency. With scaling, leakage power accounts for an increasingly larger portion of the total power consumption in deep submicron technologies (>40%). FinFET technology has been proposed as a promising alternative to deep submicron bulk CMOS technology, because of its better scalability, short-channel characteristics, and ability to suppress leakage current and mitigate device-to-device variability when compared to bulk CMOS. The subthreshold slope of a FinFET is approximately 60mV which is close to ideal. In this article, we propose a methodology for low-power FinFET based circuit synthesis. A mechanism called TCMS (Threshold Control through Multiple Supply Voltages) was previously proposed for improving the power efficiency of FinFET based global interconnects. We propose a significant generalization of TCMS to the design of any logic circuit. This scheme represents a significant divergence from the conventional multiple supply voltage schemes considered in the past. It also obviates the need for voltage level-converters. We employ accurate delay and power estimates using table look-up methods based on HSPICE simulations for supply voltage and threshold voltage optimization. Experimental results demonstrate that TCMS can provide power savings of 67.6% and device area savings of 65.2% under relaxed delay constraints. Two other variants of TCMS are also proposed that yield similar benefits. We compare our scheme to extended cluster voltage scaling (ECVS), a popular dual- V dd scheme presented in the literature. ECVS makes use of voltage level-converters. Even when it is assumed that these level-converters have zero delay, thus significantly favoring ECVS in time-constrained power optimization, TCMS still outperforms ECVS.
Prateek Mishra, Anish Muttreja, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.3
2009 A hybrid nano-CMOS architecture for defect and fault tolerance
abstract
As the end of the semiconductor roadmap for CMOS approaches, architectures based on nanoscale molecular devices are attracting attention. Among several alternatives, silicon nanowires and carbon nanotubes are the two most promising nanotechnologies according to the ITRS. These technologies may enable scaling deep into the nanometer regime. However, they suffer from very defect-prone manufacturing processes. Although the reconfigurability property of the nanoscale devices can be used to tolerate high defect rates, it may not be possible to locate all defects. With very high device densities, testing each component may not be possible because of time or technology restrictions. This points to a scenario in which even though the devices are tested, the tests are not very comprehensive at locating defects, and hence the shipped chips are still defective. Moreover, the devices in the nanometer range will be susceptible to transient faults which can produce arbitrary soft errors. Despite these drawbacks, it is possible to make nanoscale architectures practical and realistic by introducing defect and fault tolerance. In this article, we propose and evaluate a hybrid nanowire-CMOS architecture that addresses all three problems—namely high defect rates, unlocated defects, and transient faults—at the same time. This goal is achieved by using multiple levels of redundancy and majority voters. A key aspect of the architecture is that it contains a judicious balance of both nanoscale and traditional CMOS components. A companion to the architecture is a compiler with heuristics to quickly determine if logic can be mapped onto partially defective nanoscale elements. The heuristics make it possible to introduce defect-awareness in placement and routing. The architecture and compiler are evaluated by applying the complete design flow to several benchmarks.
Muzaffer O. Simsir, Srihari Cadambi, Franjo Ivancic, Martin Rötteler, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.5
2009 A hybrid Nano/CMOS dynamically reconfigurable system - Part II: Design optimization flow
abstract
In Part I of this work, a hybrid nano/CMOS reconfigurable architecture, called NATURE, was described. It is composed of CMOS reconfigurable logic and interconnect fabric, and nonvolatile nano on-chip memory. Through its support for cycle-by-cycle runtime reconfiguration and a highly-efficient computation model, temporal logic folding, NATURE improves logic density and area-delay product by more than an order of magnitude compared to existing CMOS-based field-programmable gate arrays (FPGAs). NATURE can be fabricated using mainstream photo-lithography fabrication techniques. Thus, it offers a currently commercially feasible architecture with high performance, superior logic density, and excellent runtime design flexibility. In Part II of this work, we present an integrated design and optimization flow for NATURE, called NanoMap. Given an input design specified in register-transfer level (RTL) and/or gate-level VHDL, NanoMap optimizes and implements the design on NATURE through logic mapping, temporal clustering, temporal placement, and routing. As opposed to other design tools for traditional FPGAs, NanoMap supports and leverages temporal logic folding by integrating novel mapping techniques. It can automatically explore and identify the best temporal logic folding configuration, targeting area, delay or area-delay product optimization. A force-directed scheduling technique is used to optimize and balance resource usage across different folding cycles. By supporting logic folding, NanoMap can provide significant design flexibility in performing area-delay trade-offs under various user-specified constraints. We present details of the mapping procedure and results for different architectural instances. Experimental results demonstrate that NanoMap can judiciously trade off area and delay targeting different optimization goals, and effectively exploit the advantages of NATURE. Part I of this work will appear in JETC Vol. 5, No. 4.
Wei Zhang 0012, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2009 A hybrid nano/CMOS dynamically reconfigurable system - Part I: Architecture
abstract
Rapid progress on nanodevices points to a promising direction for future circuit design. However, since nanofabrication techniques are not yet mature, implementation of nanocircuits, at least on a large scale, in the near future is infeasible. To ease fabrication and overcome the problem of high defect levels in nanotechnology, hybrid nano/CMOS reconfigurable architectures are attractive choices. Moreover, if the current photolithography fabrication process can be used to manufacture the hybrid chips, the benefits of nanotechnologies can be realized today. Traditional reconfigurable architectures can only support partial or coarse-grain runtime reconfiguration due to their limited on-chip storage and long off-chip reconfiguration latency. Recent progress on nano Random Access Memories (RAMs), such as carbon nanotube-based RAM (NRAM), Phase-Change Memory (PCM), magnetoresistive RAM (MRAM), etc., provides us with a chance to realize on-chip fine-grain runtime reconfiguration. These nano RAMs have good compatibility with the current fabrication process. By utilizing them in the hybrid design, we can take advantage of both CMOS and nanotechnology, and greatly improve the logic density, resource utilization, and performance of our design. In this article, we propose a high-performance reconfigurable architecture, called NATURE, that utilizes CMOS logic and nano RAMs. An automatic design flow for NATURE is presented in Part II of the article. In NATURE, the highly dense nonvolatile nano RAMs are distributed throughout the chip to allow large embedded on-chip configuration storage, which enables fast reading and hence supports fine-grain runtime reconfiguration and temporal logic folding of a circuit before being mapped to the architecture. Temporal logic folding can significantly increase the logic density of NATURE (by over an order of magnitude for large circuits) while remaining competitive in performance and power consumption. For ease of exposition, we use NRAMs to illustrate various concepts in this article due to the excellent properties of NRAMs. However, other nano RAMs can also be used instead. Experimental results based on NRAMs establish the efficacy of NATURE.
Wei Zhang 0012, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2009 Design space exploration and data memory architecture design for a hybrid nano/CMOS dynamically reconfigurable architecture
abstract
In recent years, research on nanotechnology has advanced rapidly. Novel nanodevices have been developed, such as those based on carbon nanotubes, nanowires, etc. Using these emerging nanodevices, diverse nanoarchitectures have been proposed. Among them, hybrid nano/CMOS reconfigurable architectures have attracted attention because of their advantages in performance, integration density, and fault tolerance. Recently, a high-performance hybrid nano/CMOS reconfigurable architecture, called NATURE, was presented. NATURE comprises CMOS reconfigurable logic and interconnect fabric, and CMOS-fabrication-compatible nanomemory. High-density, fast nano RAMs are distributed in NATURE as on-chip storage to store multiple reconfiguration copies for each reconfigurable element. It enables cycle-by-cycle runtime reconfiguration and a highly efficient computational model, called temporal logic folding. Through logic folding, NATURE provides more than an order of magnitude improvement in logic density and area-delay product, and significant design flexibility in performing area-delay trade-offs, at the same technology node. Moreover, NATURE can be fabricated using mainstream photolithography fabrication techniques. Hence, it offers a currently commercially viable reconfigurable architecture with high performance, superior logic density, and outstanding design flexibility, which is very attractive for deployment in cost-conscious embedded systems. In order to fully explore the potential of NATURE and further improve its performance, in this article, a thorough design space exploration is conducted to optimize its architecture. Investigations in terms of different logic element architectures, interconnect designs, and various technologies for nano RAMs are presented. Nano RAMs can not only be used as storage for configuration bits, but the high density of nano RAMs also makes them excellent candidates for large-capacity on-chip data storage in NATURE. Many logic- and memory-intensive applications, such as video and image processing, require large storage of temporal results. To enhance the capability of NATURE for implementing such applications, we investigate the design of nano data memory structures in NATURE and explore the impact of memory density. Experimental results demonstrate significant throughput improvements due to area saving from logic folding and parallel data processing.
Wei Zhang 0012, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2009 Editorial Appointments for the 2009-2010 Term
Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.1
2009 Fast Enhancement of Validation Test Sets for Improving the Stuck-at Fault Coverage of RTL Circuits
abstract
A digital circuit usually comprises a controller and datapath. The time spent for determining a valid controller behavior to detect a fault usually dominates test generation time. A validation test set is used to verify controller behavior and, hence, it activates various controller behaviors. In this paper, we present a novel methodology wherein the controller behaviors exercised by test sequences in a validation test set are reused for detecting faults in the datapath. A heuristic is used to identify controller behaviors that can justify/propagate pre-computed test vectors/responses of datapath register-transfer level (RTL) modules. Such controller behaviors are said to becompatiblewith the corresponding precomputed test vectors/responses. The heuristic is fairly accurate, resulting in the detection of a majority of stuck-at faults in the datapath RTL modules. Also, since test generation is performed at the RTL and the controller behavior is predetermined, test generation time is reduced. For microprocessors, if the validation test set consists of instruction sequences then the proposed methodology also generates instruction-level test sequences.
Loganathan Lingappan, Vijay Gangaram, Niraj K. Jha, Sreejit Chakravarty
IEEE Trans. Very Large Scale Integr. Syst.3
2008 Dynamic Binary Instrumentation-Based Framework for Malware Defense
Najwa Aaraj, Anand Raghunathan, Niraj K. Jha
DIMVA3
2008 A system-level perspective for efficient NoC design
abstract
With the advent of chip multiprocessors (CMPs) in mainstream systems, the on-chip network that connects different processing cores becomes a critical part of the design. There has been significant work in the recent past on designing these networks for efficiency and scalability. However, most network design evaluations use a stand-alone network simulator which fails to capture the system-level implications of the design. New design innovations, which might yield promising results when evaluated using such stand-alone models, may not look that attractive when evaluated in a full-system simulation framework. In this work, we present GARNET, a detailed network model incorporated inside a full-system simulator which enables system-level performance and power modeling of network-level techniques. GARNET also facilitates accurate evaluation of techniques that simultaneously leverage the memory hierarchy as well as the interconnection network. We also discuss express virtual channels, a novel flow control technique which improves network energy/delay by creating virtual lanes in the network along which packets can bypass intermediate routers.
Amit Kumar 0002, Niket Agarwal, Li-Shiuan Peh, Niraj K. Jha
IPDPS4
2008 Token flow control
abstract
As companies move towards many-core chips, an efficient on-chip communication fabric to connect these cores assumes critical importance. To address limitations to wire delay scalability and increasing bandwidth demands, state-of-the-art on-chip networks use a modular packet-switched design with routers at every hop which allow sharing of network channels over multiple packet flows. This, however, leads to packets going through a complex router pipeline at every hop, resulting in the overall communication energy/delay being dominated by the router overhead, as opposed to just wire energy/delay.
Amit Kumar 0002, Li-Shiuan Peh, Niraj K. Jha
MICRO3
2008 Reversible logic synthesis with Fredkin and Peres gates
abstract
Reversible logic has applications in low-power computing and quantum computing. Most reversible logic synthesis methods are tied to particular gate types, and cannot synthesize large functions. This article extends RMRLS, a reversible logic synthesis tool, to include additional gate types. While classic RMRLS can synthesize functions using NOT, CNOT, and n -bit Toffoli gates, our work details the inclusion of n -bit Fredkin and Peres gates. We find that these additional gates reduce the average gate count for three-variable functions from 6.10 to 4.56, and improve the synthesis results of many larger functions, both in terms of gate count and quantum cost.
James Donald, Niraj K. Jha
ACM J. Emerg. Technol. Comput. Syst.2
2008 System-Level Dynamic Thermal Management for High-Performance Microprocessors
abstract
Thermal issues are fast becoming major design constraints in high-performance systems. Temperature variations adversely affect system reliability and prompt worst-case design. In recent history, researchers have proposed dynamic thermal-management (DTM) techniques targeting average-case design and tackling the temperature issue at runtime. While past work on DTM has focused on different techniques in isolation, it fails to consider a system-level approach which uses both hardware and software support in a synergistic fashion and hence leads to a significant execution-time overhead. In this paper, we propose HybDTM, a system-level framework for doing fine-grained coordinated thermal management using a hybrid of hardware techniques (like clock gating) and software techniques (like thermal-aware process scheduling), leveraging the advantages of both approaches in a synergistic fashion. We show that while hardware techniques can be used reactively to manage the overall temperature in case of thermal emergencies, proactive use of software techniques can build on top of it to balance the overall thermal profile with minimal overhead using the operating system (OS) support. In order to evaluate our proposed hybrid-DTM policy, we develop a novel regression-based thermal model, providing fast and accurate temperature estimates to do runtime thermal characterization of all applications running on the system, using hardware performance counters available in modern high-performance processors alongside thermal sensors for training the model at runtime. Our model is validated against actual temperature measurements from online thermal sensors, with the average estimation error found to be less than 5%. We also study system-level DTM issues, jointly considering both the processor and memory, and show how a unified DTM approach can benefit from global knowledge of individual system components. We evaluate our proposed methodology on a desktop system with an Intel Pentium-4 processor and a modified Linux OS, running a number of SPEC2000 benchmarks, in both uniprocessor and simultaneous multithreaded environments and show that our proposed technique is able to successfully manage the overall temperature with an average execution-time overhead of only 10.4% (20.1% maximum) compared to the case without any DTM, as opposed to 23.9% (46% maximum) overhead for purely hardware-based DTM. Our system, including the thermal-aware OS, built-in runtime thermal-characterization model, and interface to the underlying hardware using the Pentium-4 processor, is ready for release.
Amit Kumar 0002, Li-Shiuan Peh, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2008 Analysis and design of a hardware/software trusted platform module for embedded systems
abstract
Trusted platforms have been proposed as a promising approach to enhance the security of general-purpose computing systems. However, for many resource-constrained embedded systems, the size and cost overheads of a separate Trusted Platform Module (TPM) chip are not acceptable. One alternative is to use a software-based TPM, which implements TPM functions using software that executes in a protected execution domain on the embedded processor itself. However, since many embedded systems have limited processing capabilities and are battery-powered, it is also important to ensure that the computational and energy requirements for SW-TPMs are acceptable. In this article, we perform an evaluation of the energy and execution time overheads for a SW-TPM implementation on a handheld appliance (Sharp Zaurus PDA). We characterize the execution time and energy required by each TPM command through actual measurements on the target platform. We observe that for most commands, overheads are primarily due to the use of 2,048-bit RSA operations that are performed within the SW-TPM. In order to alleviate SW-TPM overheads, we evaluate the use of Elliptic Curve Cryptography (ECC) as a replacement for the RSA algorithm specified in the Trusted Computing Group (TCG) standards. In addition, we also evaluate the overheads of using the SW-TPM in the context of various end applications, including trusted boot of the Linux operating system (OS), a secure VoIP client, and a secure Web browser. Furthermore, we analyze the computational workload involved in running SW-TPM commands using ECC. We then present a suite of hardware and software enhancements to accelerate these commands—generic custom instructions and exploitation of parallel processing capabilities in multiprocessor systems-on-chip (SoCs). We report results of evaluating the proposed architectures on a commercial embedded processor (Xtensa from Tensilica). Through uniprocessor and multiprocessor optimizations, we could achieve speed-ups of up to 5.71X for individual TPM commands.
Najwa Aaraj, Anand Raghunathan, Niraj K. Jha
ACM Trans. Embed. Comput. Syst.3
2008 An energy-aware framework for dynamic software management in mobile computing systems
abstract
Energy efficiency has become a very important and challenging issue for resource-constrained mobile computers. In this article, we propose a novel dynamic software management (DSOM) framework to improve battery utilization. We have designed and implemented a DSOM module in user space, independent of the operating system (OS), which explores quality-of-service (QoS) adaptation to reduce system energy and employs a priority-based preemption policy for multiple applications to avoid competition for limited energy resources. Software energy macromodels for mobile applications are employed to predict energy demand at each QoS level, so that the DSOM module is able to select the best possible trade-off between energy conservation and application QoS; it also honors the priority desired by the user. Our experimental results for some mobile applications (video player, speech recognizer, voice-over-IP) show that this approach can meet user-specified task-oriented goals and significantly improve battery utilization.
Yunsi Fei, Lin Zhong 0001, Niraj K. Jha
ACM Trans. Embed. Comput. Syst.3
2008 Automatic Test Generation for Combinational Threshold Logic Networks
abstract
We propose an automatic test pattern generation (ATPG) framework for combinational threshold networks. The motivation behind this work lies in the fact that many emerging nanotechnologies, such as resonant tunneling diodes (RTDs), single electron transistor (SET), and quantum cellular automata (QCA), implement threshold logic. Consequently, there is a need to develop an ATPG methodology for this type of logic. We have built the first automatic test pattern generator and fault simulator for threshold logic which has been integrated on top of an existing computer-aided design (CAD) tool. These exploit new fault collapsing techniques we have developed for threshold networks. We perform fault modeling, backed by HSPICE simulations, to show that many cuts and shorts in RTD-based threshold gates are equivalent to stuck-at faults at the inputs and output of the gate. Experimental results with the MCNC benchmarks indicate that test vectors were found for all testable stuck-at faults in their threshold network implementations.
Pallav Gupta, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
2007 NanoMap: An Integrated Design Optimization Flow for a Hybrid Nanotube/CMOS Dynamically Reconfigurable Architecture
abstract
NATURE is a recently developed hybrid nano/CMOS reconfigurable architecture. It consists of complementary metal-oxide semiconductor (CMOS) reconfigurable logic and interconnect fabric, and carbon nanotube-based non-volatile onchip configuration memory. Compared to existing CMOS-based field-programmable gate arrays (FPGAs), NATURE increases logic density by more than an order of magnitude and offers cycle-by-cycle run-time reconfiguration capability. As opposed to some other recently proposed hybrid nano/CMOS designs, which mostly rely on the not-yet-mature self-assembly fabrication process, NATURE is compatible with mainstream photolithography fabrication techniques. Thus, NATURE offers a commercially feasible technology with high performance, superior integration density, and excellent run-time flexibility.
Wei Zhang 0012, Niraj K. Jha
DAC3
2007 Energy and execution time analysis of a software-based trusted platform module
Najwa Aaraj, Anand Raghunathan, Srivaths Ravi 0001, Niraj K. Jha
DATE4
2007 A 4.6Tbits/s 3.6GHz single-cycle NoC router with a novel switch allocator in 65nm CMOS
abstract
As chip multiprocessors (CMPs) become the only viable way to scale up and utilize the abundant transistors made available in current microprocessors, the design of on-chip networks is becoming critically important. These networks face unique design constraints and are required to provide extremely fast and high bandwidth communication, yet meet tight power and area budgets. In this paper, we present a detailed design of our on-chip network router targeted at a 36-core shared-memory CMP system in 65nm technology. Our design targets an aggressive clock frequency of 3.6GHz, thus posing tough design challenges that led to several unique circuit and microarchitectural innovations and design choices, including a novel high throughput and low latency switch allocation mechanism, a non-speculative single-cycle router pipeline which uses advanced bundles to remove control setup overhead, a low-complexity virtual channel allocator and a dynamically-managed shared buffer design which uses prefetching to minimize critical path delay. Our router takes up 1.19mm2area and expends 551 mW power at 10% activity, delivering a single-cycle no-load latency at 3.6GHz clock frequency while achieving a peak switching data rate in excess of 4.6Tbits/s per router node.
Amit Kumar 0002, Partha Kundu, Arvind P. Singh, Li-Shiuan Peh, Niraj K. Jha
ICCD5
2007 CMOS logic design with independent-gate FinFETs
abstract
Fin-type field-effect transistors (FinFETs) are promising substitutes for bulk CMOS in nano-scale circuits. In this paper, it is observed that in spite of improved device characteristics, high active leakage may remain a problem for FinFET logic circuits. Leakage is found to contribute 31.3% of total power consumption in power-optimized FinFET logic circuits. Various FinFET logic design styles, based on independent control of FinFET gates, are studied. A new low-leakage logic style is presented. Leakage (total) power savings of 64.7% (14.5%) under tight delay constraints and 91.2% (37.2%) under relaxed delay constraints, through the judicious use of FinFET logic styles, are demonstrated.
Anish Muttreja, Niket Agarwal, Niraj K. Jha
ICCD3
2007 Express virtual channels: towards the ideal interconnection fabric
abstract
Due to wire delay scalability and bandwidth limitations inherent in shared buses and dedicated links, packet-switched on-chip interconnection networks are fast emerging as the pervasive communication fabric to connect different processing elements in many-core chips. However, current state-of-the-art packet-switched networks rely on complex routers which increases the communication overhead and energy consumption as compared to the ideal interconnection fabric.
Amit Kumar 0002, Li-Shiuan Peh, Partha Kundu, Niraj K. Jha
ISCA4
2007 Efficient Design for Testability Solution Based on Unsatisfiability for Register-Transfer Level Circuits
abstract
In this paper, we present a novel and accurate method for identifying design for testability (DFT) solutions for register-transfer level (RTL) circuits. Test generation proceeds by abstracting the circuit components using input/output propagation rules so that any justification/propagation event can be captured as a Boolean implication. Consequently, the RTL test generation problem is reduced to a satisfiability (SAT) instance. If a given SAT instance is not satisfiable, then we identify Boolean implications (also known as the unsatisfiable segment) that are responsible for unsatisfiability. We show that adding DFT elements is equivalent to modifying these clauses such that the unsatisfiable segment becomes satisfiable. The proposed DFT technique is both fast and accurate as it is applicable to RTL and mixed gate-level/RTL circuits and uses exact unsatisfiability conditions to identify the DFT solutions.
Loganathan Lingappan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Power-Efficient Scheduling for Heterogeneous Distributed Real-Time Embedded Systems
abstract
This paper addresses the problem of variable-voltage scheduling of multirate periodic task graphs (i.e., tasks with precedence relationships) in heterogeneous distributed real-time embedded systems. Such an embedded system may contain general-purpose processors, field-programmable gate arrays, and application-specific integrated circuits. First, we discuss the implications of the distribution of power consumption, i.e., power profile, of tasks and characteristics of voltage-scalable processing elements (PEs) on variable-voltage scaling. Then, we present a power-efficient variable-voltage scheduling algorithm to address these implications. The scheduling algorithm performs execution order optimization of scheduled events to increase the chances of scaling down voltages and frequencies of these voltage-scalable PEs in the distributed embedded system. It also performs power-profile and timing-constraint driven slack allocation to maximize power reduction via voltage scaling, based on the observation that the energy consumption of a task on a voltage-scalable PE is normally a convex function of the clock speed. The scheduling algorithm is also effective in the case where the variations in power consumption of different tasks can be ignored. It can be included in the inner loop of a system-level synthesis tool for design space exploration of real-time heterogeneous embedded systems, since it is very fast. We show its efficacy by comparing it to other approaches from the literature
Jiong Luo, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Automated Energy/Performance Macromodeling of Embedded Software
abstract
This paper presents an automatic methodology to perform characterization-based high-level software macromodeling. Macromodeling-based estimation can be used to speed up simulation-based software-performance/energy estimation. High-level software macromodels, which are significantly faster to evaluate than detailed models of the target hardware platform, can be used instead of the latter during simulation, resulting in orders-of-magnitude simulation speedup. However, in order to realize this potential, significant challenges need to be overcome in both the generation and use of macromodels-including how to identify the parameters to be used in the macromodel, how to define the template function to which the macromodel is fitted, etc. The authors' methodology attempts to address the aforementioned issues. Given a subprogram to be macromodeled for execution time and/or energy consumption, the proposed methodology automates the steps of parameter identification, data collection through detailed simulation, macromodel template selection, and fitting. The authors propose a novel technique to identify potential macromodel parameters and perform data collection, which draws from the concept of data-structure serialization used in distributed programming. They utilize symbolic-regression techniques to concurrently filter out irrelevant macromodel parameters, construct a macromodel function, and derive the optimal coefficient values to minimize fitting error. Experiments with several realistic benchmarks suggest that the proposed methodology improves estimation accuracy and enables wide applicability of macromodeling to complex embedded software, while realizing its potential for estimation speedup
Anish Muttreja, Anand Raghunathan, Srivaths Ravi 0001, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2007 Hybrid Simulation for Energy Estimation of Embedded Software
abstract
Software energy estimation is a critical step in the design of energy-efficient embedded systems. Instruction-level simulation techniques, despite several advances, remain too slow for iterative use in system-level exploration and for embedded systems with high software complexity. In this paper, we propose a methodology called hybrid simulation, which combines instruction set simulation with selective native execution (execution of some parts of the program directly on the simulation host computer). Hybrid simulation attempts to overcome the disadvantages of instruction-level simulation (low speed) and pure native execution (estimation accuracy and inapplicability to target-dependent code) while exploiting their advantages. In order to perform the energy estimation for natively executed subprograms, hybrid simulation leverages previously developed techniques for software energy macromodeling. This paper identifies and addresses the main challenges involved in hybrid simulation, including control/data transfer and memory synchronization between instruction set simulation and native execution domains, estimation errors due to macromodeling, and simulation overheads for switching between domains. We present an automatic tool flow for hybrid simulation, which analyzes a given program and selects functions for native execution in order to achieve maximum estimation efficiency while limiting estimation error. Our tool generates native ldquodrop-inrdquo modules that are linked into the instruction set simulator (ISS) and simulation ldquostubsrdquo that are linked with the application to facilitate hybrid simulation. We have applied the proposed hybrid simulation methodology to a variety of embedded software programs, resulting in average speedups of 25.2x (maximum of 124) and estimation error of only 3% (maximum of 6%), as compared to one of the fastest publicly available ISSs.
Anish Muttreja, Anand Raghunathan, Srivaths Ravi 0001, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2007 SLOPES: Hardware-Software Cosynthesis of Low-Power Real-Time Distributed Embedded Systems With Dynamically Reconfigurable FPGAs
abstract
In this paper, we present a multiobjective hardware-software cosynthesis system, called SLOPES, for multirate low-power real-time distributed embedded systems consisting of dynamically reconfigurable field-programmable gate arrays (FPGAs), processors, and heterogeneous communication resources. This cosynthesis algorithm simultaneously optimizes system price and average power consumption. First, we present an evolutionary algorithm that automatically determines the quantities and types of system resources, assigns tasks to different potentially reconfigurable processing elements, and assigns communication events to communication resources. Second, we propose a dynamic priority multirate scheduling algorithm to determine the times at which all the tasks and communication events in the system occur. This two-dimensional scheduling algorithm determines task priorities based on real-time constraints and detailed frame-by-frame FPGA reconfiguration overhead information. Experimental results indicate that the proposed method reduces schedule length by an average of 34.3% and reconfiguration energy by an average of 40.4%, compared to a method that does not consider the effect of partial reconfiguration during synthesis. SLOPES yields multiple system architectures that tradeoff system price and average power consumption under real-time constraints
Robert P. Dick, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2007 A Synthesis Methodology for Hybrid Custom Instruction and Coprocessor Generation for Extensible Processors
abstract
Systems-on-chip often use hardware accelerators or coprocessors to provide efficient implementations of application-specific functions. The emergence of extensible processor cores with supporting design tools has given designers with another viable alternative, namely, the use of application-specific custom instructions. Coprocessors and custom instructions can be viewed as two different forms of hardware acceleration that are applicable at different levels of granularity and offer differing tradeoffs. Classical hardware/software-partitioning techniques and application-specific instruction-set design tools address the individual problems of coprocessor generation and custom-instruction addition. However, given a complex application, it is not clear which design choice (coprocessors or custom instructions or a combination) will result in better performance, area, or power consumption. We demonstrate that a combination of custom instructions and coprocessors is often the best solution in many applications, making the case for a hybrid custom-instruction and coprocessor-synthesis methodology. We propose such a methodology that builds upon the basic observations that coprocessors are usually good for coarse-grained tasks and require minimal intervention or support from the processor, while custom instructions are usually suited to fine-grained operations that are best integrated into a processor pipeline. Our methodology uses a hierarchical task-graph representation in order to support both coarse-and fine-grained views of an application, which are necessary to make meaningful tradeoffs. We propose a hierarchical synthesis algorithm that incorporates multiobjective evolutionary optimization in order to handle different design dimensions, such as area and performance, and provide a wide range of nondominated solutions. We have implemented the proposed methodology in the context of a commercial extensible processor-based platform (Xtensa from Tensilica). Our design flow uses a commercial behavioral-synthesis tool and an existing automatic-custom-instruction-generation tool. Our experiments with several applications show that simultaneous custom-instruction and coprocessor synthesis can achieve significantly better area/performance tradeoffs than using only one of them.
Fei Sun 0002, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2007 Majority and Minority Network Synthesis With Application to QCA-, SET-, and TPL-Based Nanotechnologies
abstract
In this paper, we present a methodology for efficient majority/minority network synthesis of arbitrary multiout- put Boolean functions. Many emerging nanoscale technologies, such as quantum cellular automata (QCA), single electron tunneling (SET), and tunneling phase logic (TPL), are capable of implementing majority or minority logic very efficiently. The main purpose of this paper is to lay the foundation for research on the development of synthesis methodologies and tools to generate optimized majority/minority networks for these emergent technologies. Functionally correct QCA-, SET-, and TPL-based majority/ minority gates have been successfully demonstrated. However, there exists no comprehensive methodology or design automation tool for general multilevel majority/minority network synthesis. We have built the first such tool, majority logic synthesizer, on top of an existing Boolean logic synthesis tool. Experiments with 40 Microelectronics Center of North Carolina benchmarks were performed. They indicate that up to 68.0% reduction in gate count is possible when utilizing majority/minority logic, with the average reduction being 21.9%, compared to traditional logic synthesis, in which two-input and/or gates in the circuit are converted to majority/minority gates.
Pallav Gupta, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2007 Energy-optimizing source code transformations for operating system-driven embedded software
abstract
This paper proposes four types of source code transformations for operating system (OS)-driven embedded software programs to reduce their energy consumption. Their key features include spanning of process boundaries and minimization of the energy consumed in the execution of OS services—opportunities which are beyond the reach of conventional compiler optimizations and source code transformations. We have applied the proposed transformations to several multiprocess benchmark programs in the context of an embedded Linux OS running on an Intel StrongARM processor. They achieve up to 37.9% (23.8%, on average) energy reduction compared to highly compiler-optimized implementations.
Yunsi Fei, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
ACM Trans. Embed. Comput. Syst.4
2007 Hybrid Architectures for Efficient and Secure Face Authentication in Embedded Systems
abstract
In this paper, we propose an efficient and secure embedded processing architecture that addresses various challenges involved in using face-based biometrics for authenticating a user to an embedded system. Our paper considers the use of robust face verifiers (PCA-LDA, Bayesian), and analyzes the computational workload involved in running their software implementations on an embedded processor. We then present a suite of hardware and software enhancements to accelerate these algorithms-fixed-point arithmetic, various code optimizations, generic custom instructions and dedicated coprocessors, and exploitation of parallel processing capabilities in multiprocessor systems-on-chip (SoCs). We also identify attacks targeted against the authentication process, and develop security measures to ensure the integrity of biometric code/data. We evaluated the proposed architectures in the context of popular open-source software implementations of face authentication algorithms running on a commercial embedded processor (Xtensa from Tensilica). Our paper shows that fast, in-system verification is possible even in the context of many resource-constrained embedded systems. We also demonstrate that the security of the authentication process for the given attack model can be achieved with minimum hardware overheads
Najwa Aaraj, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.4
2007 Architectural Support for Run-Time Validation of Program Data Properties
abstract
As computer systems penetrate deeper into our lives and handle private data, safety-critical applications, and transactions of high monetary value, efforts to breach their security also assume significant dimensions way beyond an amateur hacker's play. Until now, security was always an afterthought. This is evident in regular updates to antivirus software, patches issued by vendors after software bugs are discovered, etc. However, increasingly, we are realizing the need to incorporate security during the design of a system, be it software or hardware. We invoke this philosophy in the design of a hardware-based system to enable protection of a program's data during execution. In this paper, we develop a general framework that provides security assurance against a wide class of security attacks. Our work is based on the observation that a program's normal or permissible behavior with respect to data accesses can be characterized by various properties. We present a hardware/software approach wherein such properties can be encoded as data attributes and enforced as security policies during program execution. These policies may be application- specific (e.g., access control for certain data structures), compiler- generated (e.g., enforcing that variables are accessed only within their scope), or universally applicable to all programs (e.g., disallowing WRITES to unallocated memory). We show how an embedded system architecture can support such policies by: 1) enhancing the memory hierarchy to represent the attributes of each datum as security tags that are linked to it throughout its lifetime and 2) adding a configurable hardware checker that interprets the semantics of the tags and enforces the desired security policies. We evaluated the effectiveness of the proposed architecture in enforcing various security policies for several embedded benchmark applications. Our experiments in the context of the Simplescalar framework demonstrate that the proposed solution ensures run-time validation of application-defined data properties with minimal execution time overheads.
Divya Arora 0001, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.4
2007 Exploring Software Partitions for Fast Security Processing on a Multiprocessor Mobile SoC
abstract
The functionality of mobile devices, such as cell phones and personal digital assistants (PDAs), has evolved to include various applications where security is a critical concern (secure web transactions, mobile commerce, download and playback of protected audio/video content, connection to corporate private networks, etc.). Security mechanisms (e.g., secure communication protocols) involve cryptographic algorithms, and are often quite computationally intensive, challenging the constrained processing and battery resources of mobile devices. Extensive design effort and aggressive hardware and software optimizations are required to address this challenge. Previous work has addressed the design of hardware architectures (custom accelerators, domain-specific processors, etc.) to accelerate security processing, and many emerging systems-on-chip (SoCs) feature some form of hardware support for security. In this paper, we address the complementary problem of mapping a complex security software library to an SoC platform with security hardware enhancements. We present a systematic methodology for exploring the software architecture for security processing for a commercial heterogeneous multiprocessor SoC for mobile devices. The SoC contains multiple host processors executing applications and a dedicated programmable security processing engine. We developed an exploration methodology to map the code and data of security software libraries onto the platform, with the objective of maximizing the overall application-visible performance. The salient features of the methodology include: 1) the use of real performance measurements from a prototyping board, which contains the target platform, to drive the exploration; 2) a new data structure access profiling framework that allows us to accurately model the communication overheads involved in off loading a given set of functions to the security processor; and 3) an exact branch-and-bound-based design space exploration algorithm that determines the best mapping of security library functions and data structures to the host and security processors. We used the proposed framework to map a commercial security library to the target mobile application SoC. The resulting optimized software architecture outperformed several manually designed software architectures, resulting in up to 12.5 times speed-up for individual cryptographic operations (encryption, hashing) and 2.2-6.2 times speed-up for applications such as a digital rights management (DRM) agent and secure sockets layer (SSL) client. We also demonstrate the applicability of our framework to software architecture exploration in other multiprocessor scenarios.
Divya Arora 0001, Anand Raghunathan, Srivaths Ravi 0001, Murugan Sankaradass, Niraj K. Jha, Srimat T. Chakradhar
IEEE Trans. Very Large Scale Integr. Syst.5
2007 A Test Generation Framework for Quantum Cellular Automata Circuits
abstract
In this paper, we present a test generation framework for quantum cellular automata (QCA) circuits. QCA is a nanotechnology that has attracted recent significant attention and shows promise as a viable future technology. This work is motivated by the fact that the stuck-at fault test set of a circuit is not guaranteed to detect all defects that can occur in its QCA implementation. We show how to generate additional test vectors to supplement the stuck-at fault test set to guarantee that all simulated defects in the QCA gates get detected. Since nanotechnologies will be dominated by interconnects, we also target bridging faults on QCA interconnects. The efficacy of our framework is established through its application to QCA implementations of MCNC and ISCAS'85 benchmarks that use majority gates as primitives
Pallav Gupta, Niraj K. Jha, Loganathan Lingappan
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Generation of Heterogeneous Distributed Architectures for Memory-Intensive Applications Through High-Level Synthesis
abstract
Memory-intensive applications present unique challenges to an application-specific integrated circuit (ASIC) designer in terms of the choice of memory organization, memory size requirements, bandwidth and access latencies, etc. The high potential of single-chip distributed logic-memory architectures in addressing many of these issues has been recognized in general-purpose computing, and more recently, in ASIC design. The high-level synthesis (HLS) techniques presented in this paper are motivated by the fact that many memory-intensive applications exhibit irregular array data access patterns. Synthesis should, therefore, be capable of determining a partitioned architecture, wherein array data and computations may have to be heterogeneously distributed for achieving the best performance speed-up. We use a combination of clustering and min-cut style partitioning techniques to yield distributed architectures, based on simulation profiling while considering various factors including data access locality, balanced workloads, inter-partition communication, etc. Our experiments with several benchmark applications show that the proposed techniques yielded two-way partitioned architectures that can achieve upto 2.1 x (average of 1.9 x) performance speed-up over conventional HLS solutions, while achieving upto 1.5 x (average of 1.4 x) performance speed-up over the best homogeneous partitioning solution feasible. At the same time, the reduction in the energy-delay product over conventional single-memory designs is upto 2.7 x (average of 2.0 x). A larger amount of partitioning makes further system performance improvement achievable at the cost of chip area.
Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.4
2007 Editorial
Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.1
2007 Satisfiability-Based Automatic Test Program Generation and Design for Testability for Microprocessors
abstract
In this paper, we present a satisfiability (SAT)-based framework for automatically generating test programs that target gate-level stuck-at faults in microprocessors. The microarchitectural description of a processor is first translated into a unified register-transfer level (RTL) circuit description, called assignment decision diagram (ADD), for test analysis. Test generation involves extraction of justification/propagation paths in the unified circuit representation from an embedded module's input-output (I/O) ports to primary I/O ports, abstraction of RTL modules in the justification/propagation paths, and translation of these paths into Boolean clauses in conjunctive normal form (CNF). Additional clauses are added that capture precomputed test vectors/responses at the embedded module's I/O ports. An SAT solver is then invoked to find valid paths that justify the precomputed vectors to primary input ports and propagate the good/faulty responses to primary output ports. Since the ADD is derived directly from a microarchitectural description, the generated test sequences correspond to a test program. If a given SAT instance is not satisfiable, then Boolean implications (also known as the unsatisfiable segment) that are responsible for unsatisfiability are efficiently and accurately identified. We show that adding design for testability (DFT) elements is equivalent to modifying these clauses such that the unsatisfiable segment becomes satisfiable. Test generation at the RTL also imposes a large number of initial conditions that need to be satisfied for successful detection of targeted stuck-at faults. We demonstrate that application of the Boolean constraint propagation (BCP) engine in SAT solvers propagates these conditions leading to significant pruning of the sequential search space which in turn leads to a reduction in test generation time. Experimental results demonstrate an 11.1X speedup in test generation time for test generation at the RTL over a state-of-the-art gate-level sequential generator called MIX, at comparable fault coverages. An unsatisifiability-based DFT approach at the RTL improves this fault coverage to near 100% and incurs very low area overhead (3.1%). Unlike previous approaches that either generate a test program consisting of random instruction sequences or assume the existence of test program templates, the proposed approach constructs test programs in a deterministic fashion from the microarchitectural description of a processor
Loganathan Lingappan, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Simultaneous Dynamic Voltage Scaling of Processors and Communication Links in Real-Time Distributed Embedded Systems
abstract
Dynamic voltage scaling has been widely acknowledged as a powerful technique for trading off power consumption and delay for processors. Recently, variable-frequency (and variable-voltage) parallel and serial links have also been proposed, which can save link power consumption by exploiting variations in the bandwidth requirement. This provides a new dimension for power optimization in a distributed embedded system connected by a voltage-scalable interconnection network. At the same time, it imposes new challenges for variable-voltage scheduling as well as flow control. First, the variable-voltage scheduling algorithm should be able to trade off the power consumption and delay jointly for both processors and links. Second, for the variable-frequency network, the scheduling algorithm should not only consider the real-time constraints, but should also be consistent with the underlying flow control techniques. In this paper, we address joint dynamic voltage scaling for variable-voltage processors and communication links in such systems. We propose a scheduling algorithm for real-time applications that captures both data flow and control flow information. It performs efficient routing of communication events through multihops, as well as efficient slack allocation among heterogeneous processors and communication links to maximize energy savings, while meeting all real-time constraints. Our experimental study shows that on an average, joint voltage scaling on processors and links can achieve 32% less power compared with voltage scaling on processors alone
Jiong Luo, Niraj K. Jha, Li-Shiuan Peh
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Aiding Side-Channel Attacks on Cryptographic Software With Satisfiability-Based Analysis
abstract
Cryptographic algorithms, irrespective of their theoretical strength, can be broken through weaknesses in their implementations. The most successful of these attacks are side-channel attacks which exploit unintended information leakage, e.g., timing information, power consumption, etc., from the implementation to extract the secret key. We propose a novel framework for implementing side-channel attacks where the attack is modeled as a search problem which takes the leaked information as its input, and deduces the secret key by using a satisfiability solver, a powerful Boolean reasoning technique. This approach can substantially enhance the scope of side-channel attacks by allowing a potentially wide range of internal variables to be exploited (not just those that are trivially related to the key). The proposed technique is particularly suited for attacking cryptographic software implementations which may inadvertently expose the values of intermediate variables in their computations (even though, they are very careful in protecting secret keys through the use of on-chip key generation and storage). We demonstrate our attack on standard software implementations of three popular cryptographic algorithms: DES, 3DES, and AES. Our attack technique is automated and does not require mathematical expertise on the part of the attacker
Nachiketh R. Potlapally, Anand Raghunathan, Srivaths Ravi 0001, Niraj K. Jha, Ruby B. Lee
IEEE Trans. Very Large Scale Integr. Syst.4
2007 Configuration and Extension of Embedded Processors to Optimize IPSec Protocol Execution
abstract
Security protocols, such as IPSec and SSL, are being increasingly deployed in the context of networked embedded systems. The resource-constrained nature of embedded systems and, in particular, the modest capabilities of embedded processors make it challenging to achieve satisfactory performance while executing security protocols. A promising approach for improving performance in embedded systems is to use application-specific instruction set processors that are designed based on configurable and extensible processors. In this paper, we perform a comprehensive performance analysis of the IPSec protocol on a state-of-the-art configurable and extensible embedded processor (Xtensa from Tensilica Inc.). We present performance profiles of a lightweight embedded IPSec implementation running on the Xtensa processor, and examine in detail the various factors that contribute to the processing latencies, including cryptographic and protocol processing. In order to improve the efficiency of IPSec processing on embedded devices, we then study the impact of customizing an embedded processor by synergistically 1) configuring architectural parameters, such as instruction and data cache sizes, processor-memory interface width, write buffers, etc., and 2) extending the base instruction set of the processor using custom instructions for both cryptographic and protocol processing. Our experimental results demonstrate that upto 3.2times speedup in IPSec processing is possible over a popular embedded IPSec software implementation
Nachiketh R. Potlapally, Srivaths Ravi 0001, Anand Raghunathan, Ruby B. Lee, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.5
2006 Software architecture exploration for high-performance security processing on a multiprocessor mobile SoC
abstract
We present a systematic methodology for exploring the security processing software architecture for a commercial heterogeneous multiprocessor system-on-chip (SoC) for mobile devices. The SoC contains multiple host processors executing applications and a dedicated programmable security processing engine. We developed an exploration methodology to map the code and data of security software libraries onto the platform, with the objective of maximizing the overall application-visible performance. The salient features of the methodology include (i) the use of real performance measurements from a prototyping board that contains the target platform to drive the exploration, (ii) a new data structure access profiling framework that allows us to accurately model the communication overheads involved in offloading a given set of functions to the security processor, and (iii) an exact branch-and-bound based design space exploration algorithm that determines the best mapping of security library functions and data structures to the host and security processors.We used the proposed framework to map a commercial security library to the target mobile application SoC. The resulting optimized software architecture outperformed several manually-designed software architectures, resulting in upto 12.5X speedup for individual cryptographic operations (encryption, hashing) and 2.2X-6.2X speedup for applications such as a Digital Rights Management (DRM) agent and Secure Sockets Layer (SSL) client. We also demonstrate the applicability of our framework to software architecture exploration in other multiprocessor scenarios.
Divya Arora 0001, Anand Raghunathan, Srivaths Ravi 0001, Murugan Sankaradass, Niraj K. Jha, Srimat T. Chakradhar
DAC5
2006 HybDTM: a coordinated hardware-software approach for dynamic thermal management
abstract
With ever-increasing power density and cooling costs in modern high-performance systems, dynamic thermal management (DTM) has emerged as an effective technique for guaranteeing thermal safety at run-time. While past works on DTM have focused on different techniques in isolation, they fail to consider a synergistic mechanism using both hardware and software support and hence lead to a significant execution time overhead.In this paper, we propose HybDTM, a methodology for fine-grained, coordinated thermal management using a hybrid of hardware techniques, such as clock gating, and software techniques, such as thermal-aware process scheduling, synergistically leveraging the advantages of both approaches. We show that while hardware techniques can be used reactively to manage thermal emergencies, proactive use of low-overhead software techniques can rely on application-specific thermal profiles to lower system temperature. Our technique involves a novel regression-based thermal model which provides fast and accurate temperature estimates for run-time thermal characterization of applications running on the system, using hardware performance counters, while considering system-level thermal issues. We evaluate HybDTM on an actual desktop system running a number of SPEC2000 benchmarks, in both uniprocessor and simultaneous multi-threading (SMT) environments, and show that it is able to successfully manage the overall temperature with an average execution time overhead of only 9.9% (16.3% maximum) compared to the case without any DTM, as opposed to 20.4% (29.5% maximum) overhead for purely hardware-based DTM.
Amit Kumar 0002, Li-Shiuan Peh, Niraj K. Jha
DAC4
2006 NATURE: a hybrid nanotube/CMOS dynamically reconfigurable architecture
abstract
Recent progress on nanodevices, such as carbon nanotubes and nanowires, points to promising directions for future circuit design. However, nanofabrication techniques are not yet mature, making implementation of such circuits, at least on a large scale, in the near future infeasible. However, if photo-lithography could be used to implement circuits using these nanodevices, then hybrid nano/CMOS chips could be fabricated and the benefits of nanotechnology could be utilized immediately. A startup company, called Nantero, has developed and implemented a non-volatile nanotube random-access memory (NRAM) using photo-lithography that is considerably faster and denser than DRAM, has much lower power consumption than DRAM or flash, has similar speed to SRAM and is highly resistant to environmental forces (temperature, magnetism). In this paper, we propose a novel high performance reconfigurable architecture, called NATURE, that utilizes CMOS logic and NRAMs. Use of the highly-dense NRAMs allows large on-chip configuration storage, enabling fine-grain run-time reconfiguration and temporal logic folding of a circuit before being mapped to the architecture. This can significantly increase the logic density of NATURE (by over an order of magnitude for larger circuits) while remaining competitive in performance. Compared to traditional reconfigurable architectures, NATURE also allows the designer the flexibility to adjust the level of logic folding in order to improve performance or perform area-performance trade-offs. Experimental results establish its efficacy and give comparisons with today's mainstream FPGA technology which does not allow logic folding.
Wei Zhang 0012, Niraj K. Jha
DAC2
2006 Test generation for combinational quantum cellular automata (QCA) circuits
abstract
In this paper, we present a test generation framework for testing of quantum cellular automata (QCA) circuits. QCA is a nanotechnology that has attracted significant recent attention and shows immense promise as a viable future technology. This work is motivated by the fact that the stuck-at fault test set of a circuit is not guaranteed to detect all defects that can occur in its QCA implementation. We show how to generate additional test vectors to supplement the stuck-at fault test set to guarantee that all simulated defects in the QCA gates get detected. Since nanotechnologies will be dominated by interconnects, we also target bridging faults on QCA interconnects. The efficacy of our framework is established through its application to QCA implementations of MCNC benchmarks that use majority gates as primitives
Pallav Gupta, Niraj K. Jha, Loganathan Lingappan
DATE2
2006 Threshold/majority logic synthesis and concurrent error detection targeting nanoelectronic implementations
abstract
Many nanometer-scale devices have been proposed and fabricated recently. Several can implement threshold and majority logic efficiently. Research has also begun on design methodologies to keep pace with the development of these devices. Specifically, a threshold logic synthesis tool (TELS) and a majority/minority logic synthesis tool (MALS) have been developed recently. In this paper, we discuss several factorization methods to enhance the efficacy of these two tools significantly. We then augment the design methodology to allow the tools to produce totally self-checking (TSC) circuits which can efficiently implement concurrent error detection. Such circuits can be used to detect run-time errors. Two schemes are used to synthesize TSC circuits - one based on the Berger code and the other on the parity code. We compare and contrast these two schemes. Experimental results establish the effectiveness of the proposed approaches.
Niraj K. Jha
ACM Great Lakes Symposium on VLSI2
2006 Active Learning Driven Data Acquisition for Sensor Networks
abstract
Online monitoring of a physical phenomenon over a geographical area is a popular application of sensor networks. Networks representative of this class of applications are typically operated in one of two modes, viz. an always-on mode where every sensor reading is streamed to a base station, possibly after in-network aggregation, and a snapshot mode where a user queries the network for an instantaneous summary of the observed field. However, a continuum of data acquisition policies exists between these two extreme modes, depending upon the rate and manner in which each sensor node is queried. In this work, we explore this continuum to improve network energy efficiency. We present a data acquisition framework that models the evolution of the observed data field at each sensor location as a function of time and uses an active learning based criterion to intelligently sample each sensor. Sensor nodes in our framework are organized in a clustered hierarchy. Time-dependent models of sensor readings are maintained at cluster-head nodes, which sample nodes in their cluster in a way that minimizes total energy consumption while maintaining confidence bounds on the overall model. We use sparse Gaussian processes to model sensor readings and variance minimization based active learning to intelligently select sensor nodes for querying. Finally, we present simulation results demonstrating up to 70% savings in total network energy, compared to the base case, in which sensors are sampled according to a cyclic schedule.
Anish Muttreja, Anand Raghunathan, Srivaths Ravi 0001, Niraj K. Jha
ISCC4
2006 An Algorithm for Synthesis of Reversible Logic Circuits
abstract
Reversible logic finds many applications, especially in the area of quantum computing. A completely specified n-input, n-output Boolean function is called reversible if it maps each input assignment to a unique output assignment and vice versa. Logic synthesis for reversible functions differs substantially from traditional logic synthesis and is currently an active area of research. The authors present an algorithm and tool for the synthesis of reversible functions. The algorithm uses the positive-polarity Reed-Muller expansion of a reversible function to synthesize the function as a network of Toffoli gates. At each stage, candidate factors, which represent subexpressions common between the Reed-Muller expansions of multiple outputs, are explored in the order of their attractiveness. The algorithm utilizes a priority-based search tree, and heuristics are used to rapidly prune the search space. The synthesis algorithm currently targets the generalized n-bit Toffoli gate library. However, other algorithms exist that can convert an n-bit Toffoli gate into a cascade of smaller Toffoli gates. Experimental results indicate that the authors' algorithm quickly synthesizes circuits when tested on the set of all reversible functions of three variables. Furthermore, it is able to quickly synthesize all four-variable and most five-variable reversible functions that were in the test suite. The authors also present results for some benchmark functions widely discussed in literature and some new benchmarks that the authors have developed. The algorithm is shown to synthesize many, but not all, randomly generated reversible functions of as many as 16 variables with a maximum gate count of 25
Pallav Gupta, Abhinav Agrawal 0002, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2006 Use of Computation-Unit Integrated Memories in High-Level Synthesis
abstract
High-level synthesis (HLS) of memory-intensive applications has featured several innovations in terms of enhancements made to the basic memory organization and data layout. However, increasing performance and energy demands faced by application-specific integrated circuits (ASICs) are forcing designers to alter the fundamental architectural template of the HLS output, namely, a controller datapath associated with a memory subsystem (monolithic, partitioned, etc.). An architectural template for the HLS output that consists of a controller-datapath circuit associated with a memory subsystem into which computation units have been integrated is proposed. The enhanced memory subsystem is called computation-unit integrated memory (CIM). A CIM offers higher memory bandwidth (relative to what is offered through the system bus) to computation units present locally within it and reduces the overall communication between the memory subsystem and the controller datapath, thus providing a template highly suitable for deriving efficient implementations of memory-intensive applications. This paper addresses the challenge of providing a systematic synthesis framework for a CIM-based architecture. This framework can analyze the various tradeoffs involved in selecting suitable operations in a behavior for execution using a CIM and generate a high-performance low-overhead implementation. Efficient data reuse of register files have also been fully exploited to further improve system performance. Experiments with several behaviors indicate that an average performance improvement of 2.02times(a maximum of 2.70times) is possible with very low area overheads. The energy-delay product improves by an average of 2.5times(maximum of 3.8times)
Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2006 Satisfiability-based test generation for nonseparable RTL controller-datapath circuits
abstract
In this paper, we present a satisfiability (SAT)-based algorithm for automatically generating test sequences that target gate-level stuck-at faults in a circuit by using its register-transfer level (RTL) description. Our methodology uses a unified RTL circuit representation, called assignment-decision diagrams (ADDs), for test analysis. Test generation proceeds by abstracting the components in this unified representation using input/output propagation rules, so that any justification/propagation event can be captured as a Boolean implication. Consequently, we reduce RTL test generation to an SAT instance that has a significantly lower complexity than the equivalent problem at the gate level. Our algorithm is tailored to overcome the disadvantages of several existing RTL precomputed test-set-based approaches, such as the need for an explicit controller/datapath separation, the use of all test vectors or none from the precomputed test set for any given module, a dependence on symbolic justification (observability) paths from (to) circuit inputs (outputs) for a module, and a lack of applicability to mixed gate-level/RTL designs. Using the state-of-the-art SAT solver Zchaff, we show that our RTL test generator can outperform gate-level sequential automatic test-pattern generation (ATPG), in terms of both fault coverage and test-generation time (two-to-three orders of magnitude speedup), in comparable test-application times. Furthermore, we show that in a bilevel testing scenario, in which RTL ATPG is followed by gate-level sequential ATPG on the remaining faults, we improve the fault coverage even further, while maintaining a high speedup in test-generation time (nearly 32/spl times/) over pure gate-level sequential ATPG, at comparable test-application times.
Loganathan Lingappan, Srivaths Ravi 0001, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2006 Test-Volume Reduction in Systems-on-a-Chip Using Heterogeneous and Multilevel Compression Techniques
abstract
In this paper, the authors present compression techniques for effectively reducing the test-data-volume requirements of modern systems-on-a-chip (SOC). Their techniques are based on the following observations: 1) Conventional test compression schemes, which are designed to satisfy various constraints including low hardware overheads and decompression times, cannot fully exploit compression opportunities present in test data and 2) due to the diversity of components used in SOCs (and consequently in their test strategies and test-data characteristics), a single compression strategy may not be best suited to handle them. The authors propose the use of multilevel and heterogeneous test compression schemes to address the above issues and demonstrate that they can provide significant reductions in the test volume above currently known state-of-the-art test compression techniques. An architecture that reuses infrastructure components already present in SOCs (programmable processors, on-chip communication architecture, memory, etc.) for an efficient implementation of their techniques is proposed. Finally, the authors suggest various architectural-customization techniques, such as partitioning of the decompression functionality between the hardware and software and the addition of custom instructions, to improve decompression times and reduce hardware overheads. Experiments with several designs, including an industrial media-processing SOC, demonstrate the efficacy of the proposed techniques in achieving test-data-volume reductions with low overheads
Loganathan Lingappan, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha, Srimat T. Chakradhar
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2006 PowerHerd: a distributed scheme for dynamically satisfying peak-power constraints in interconnection networks
abstract
As interconnection networks proliferate to a wide range of high-performance systems, power consumption is becoming a significant architectural issue. In interconnection networks, the peak-power consumption directly affects the solution for package cooling and power-delivery design. Off-line worst-case power analysis is typically used to estimate the network peak-power consumption and guarantee safe online operation, which not only increases system cost, but also constrains network performance. In this paper, we present an online mechanism, called PowerHerd, to efficiently manage network power resources at runtime, and guarantee that network peak-power constraints are not exceeded. PowerHerd is a distributed approach-within the interconnection network, each router dynamically maintains a local power budget, controls its local power dissipation, and exchanges spare power resources with its neighboring routers to optimize network performance. Experiments demonstrate that PowerHerd can effectively regulate network power consumption, meeting peak-power constraints with negligible network-performance penalty. Armed with PowerHerd, network designers can focus on system performance and power optimization for the average case, rather than the worst-case, thus making it possible to employ a more powerful interconnection network in the system.
Li-Shiuan Peh, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2006 Application-specific heterogeneous multiprocessor synthesis using extensible processors
abstract
The complexity of many embedded software applications and their stringent performance requirements have reached a point where they can no longer be supported by conventional embedded-system architectures based on a single general-purpose processor. Nanometer fabrication technologies have made it feasible to integrate several embedded processors on a single chip to create multiprocessor systems-on-chip (MPSoCs) for a wide range of applications. It is already well known that, for embedded systems, customizing the processor to the application can significantly increase performance and energy efficiency. Single-chip heterogeneous multiprocessors, where the different processors are customized to the tasks they perform, are thus attracting attention. Despite their significant potential, a primary bottleneck to their adoption remains the development of programming paradigms and tools to alleviate the design complexity. Some advances have been made in the areas of custom multiprocessor architecture, tool support, and design methodology in order to reduce the design turnaround time. However, because of the heterogeneity of the custom processors, the designers are still required to manually map the application tasks to the different processors and, at the same time, customize each processor so that the overall performance requirements are satisfied. A multilevel custom multiprocessor-synthesis methodology to perform the assignment and scheduling of the application tasks on the various processors together with processor customization in an integrated manner is proposed. This paper focuses on extensible processors, which provide a good tradeoff between efficiency, flexibility, and design turnaround time, by combining a configurable base processor with hardware that implements custom instructions to extend its instruction set. This paper motivates the need for such an integrated approach by demonstrating that custom-instruction selection has complex interdependencies with task assignment and scheduling, and performing these steps independently often results in significant degradation in the quality of the synthesized multiprocessor architecture. The methodology presented here uses an iterative improvement algorithm to first assign and schedule tasks on processors and then select custom instructions along the critical path. It utilizes the concept of "expected execution time" to connect these two steps. It not only considers the currently selected custom instructions for the current task assignment and schedule, but also the possibility of better custom instructions being selected in future iterations. The methodology presented here is also enhanced to integrate task-level software pipelining to further increase the parallelism in the task graph and provide opportunities for multiprocessing. In this paper, the proposed heterogeneous multiprocessor-synthesis methodology is implemented in the context of a commercial extensible-processor design flow, using the Xtensa platform from Tensilica Inc. The presented tool was evaluated by automatically generating custom multiprocessor architectures for several complex embedded software benchmarks. The results show that architectures synthesized by the proposed methodology demonstrate an average speedup of 2.0/spl times/ (up to 2.9/spl times/) compared to symmetric multiprocessor architectures in which the processors have not been augmented with custom instructions. To the best of the authors' knowledge, this is the first tool for the synthesis of custom MPSoCs using extensible processors.
Fei Sun 0002, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2006 RTL-Aware Cycle-Accurate Functional Power Estimation
abstract
Most methods for hardware power estimation operate at the register-transfer level (RTL) or lower levels of design abstraction. Since cycle-accurate functional descriptions (CAFDs) are being widely adopted in integrated circuit (IC) design flows, power estimation can potentially benefit from the inherent increase in the efficiency of a cycle-based functional simulation. However, attempts at power estimation for functional descriptions have suffered from a poor accuracy because the design decisions performed during their synthesis lead to an unavoidable large uncertainty in any power estimate that is based solely on the functional description. The authors propose a methodology for a CAFD power estimation that combines the accuracy achieved by structural RTL power estimation with the efficiency of cycle-accurate functional simulation. This goal is achieved by viewing a CAFD as an abstraction of a specific known RTL implementation that is synthesized from it. The authors identify correlations between a CAFD and its RTL implementation and "back-annotate" information into the CAFD solely for power estimation. The resulting RTL-aware CAFD contains a layer of code that instantiates virtual placeholders for RTL components and maps values of CAFD variables into the RTL components' inputs/outputs, thus enabling efficient and accurate power estimation. Power estimation is performed in the proposed methodology by simply cosimulating the RTL-aware CAFD with a simulatable power-model library that contains power macromodels for each RTL component. Techniques to further improve the speed of CAFD power estimation through the use of control-state-based adaptive power sampling are presented. The authors have implemented and evaluated the proposed techniques in the context of a commercial C-based hardware design flow. Experiments with a number of large industrial designs (up to 1 000 000 gates) demonstrate that the proposed methodology achieves an accuracy very close to RTL power estimation with two-three orders of magnitude speedup in estimation times
Lin Zhong 0001, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2006 A Study of the Energy Consumption Characteristics of Cryptographic Algorithms and Security Protocols
abstract
Security is becoming an everyday concern for a wide range of electronic systems that manipulate, communicate, and store sensitive data. An important and emerging category of such electronic systems are battery-powered mobile appliances, such as personal digital assistants (PDAs) and cell phones, which are severely constrained in the resources they possess, namely, processor, battery, and memory. This work focuses on one important constraint of such devices-battery life-and examines how it is impacted by the use of various security mechanisms. In this paper, we first present a comprehensive analysis of the energy requirements of a wide range of cryptographic algorithms that form the building blocks of security mechanisms such as security protocols. We then study the energy consumption requirements of the most popular transport-layer security protocol: Secure Sockets Layer (SSL). We investigate the impact of various parameters at the protocol level (such as cipher suites, authentication mechanisms, and transaction sizes, etc.) and the cryptographic algorithm level (cipher modes, strength) on the overall energy consumption for secure data transactions. To our knowledge, this is the first comprehensive analysis of the energy requirements of SSL. For our studies, we have developed a measurement-based experimental testbed that consists of an iPAQ PDA connected to a wireless local area network (LAN) and running Linux, a PC-based data acquisition system for real-time current measurement, the OpenSSL implementation of the SSL protocol, and parameterizable SSL client and server test programs. Based on our results, we also discuss various opportunities for realizing energy-efficient implementations of security protocols. We believe such investigations to be an important first step toward addressing the challenges of energy-efficient security for battery-constrained systems.
Nachiketh R. Potlapally, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Mob. Comput.4
2006 Energy-Efficient Graphical User Interface Design
abstract
Mobile computers, such as cell phones and personal digital assistants (PDAs), have dramatically increased in sophistication. At the same time, the desire of consumers for portability limits batters size. As a result, many researchers have targeted hardware and software energy optimization. However, most of these techniques focus on compute-intensive applications rather than interactive applications, which are dominant in mobile computers. These systems frequently use graphical user interfaces (GUIs) to handle human-computer interaction. This paper is the first to explore how GUIs can be designed to improve system energy efficiency. We investigate how GUI design approaches should be changed to improve system. Energy efficiency and provide specific suggestions to mobile computer designers to enable them to develop more energy-efficient systems. We demonstrate that energy-efficient GUI (E/sup 2/GUI) design techniques can improve the average system energy of three benchmarks (text-viewer, personnel viewer, and calculator) by 26.9, 45.2, and 16.4 percent, respectively. Average performance is simultaneously improved by 23.7, 34.6, and 19.3 percent, respectively.
Keith S. Vallerio, Lin Zhong 0001, Niraj K. Jha
IEEE Trans. Mob. Comput.3
2006 Dynamic Power Optimization Targeting User Delays in Interactive Systems
abstract
Power has become a major concern for mobile computing systems such as laptops and handhelds, on which a significant fraction of software usage is interactive instead of compute-intensive. For interactive systems, an analysis shows that more than 90 percent of system energy and time is spent waiting for user input. Such idle periods provide vast opportunities for dynamic power management (DPM) and voltage scaling (DVS) techniques to reduce system energy. In this work, we propose to utilize user interface information to predict user delays based on human-computer interaction history and theories from the field of psychology. We show that such a delay prediction can be combined with DPM/DVS for aggressive power optimization. We verify the effectiveness of our methodologies with usage traces collected on a personal digital assistant (PDA) and a system power model based on accurate measurements. Experiments show that using predicted user delays for DPM/DVS achieves an average of 21.9 percent system energy reduction with little sacrifice in user productivity or satisfaction.
Lin Zhong 0001, Niraj K. Jha
IEEE Trans. Mob. Comput.2
2006 Hardware-Assisted Run-Time Monitoring for Secure Program Execution on Embedded Processors
abstract
Embedded system security is often compromised when "trusted" software is subverted to result in unintended behavior, such as leakage of sensitive data or execution of malicious code. Several countermeasures have been proposed in the literature to counteract these intrusions. A common underlying theme in most of them is to define security policies at the system level in an application-independent manner and check for security violations either statically or at run time. In this paper, we present a methodology that addresses this issue from a different perspective. It defines correct execution as synonymous with the way the program was intended to run and employs a dedicated hardware monitor to detect and prevent unintended program behavior. Specifically, we extract properties of an embedded program through static program analysis and use them as the bases for enforcing permissible program behavior at run time. The processor architecture is augmented with a hardware monitor that observes the program's dynamic execution trace, checks whether it falls within the allowed program behavior, and flags any deviations from expected behavior to trigger appropriate response mechanisms. We present properties that capture permissible program behavior at different levels of granularity, namely inter-procedural control flow, intra-procedural control flow, and instruction-stream integrity. We outline a systematic methodology to design application-specific hardware monitors for any given embedded program. Hardware implementations using a commercial design flow, and cycle-accurate performance simulations indicate that the proposed technique can thwart several common software and physical attacks, facilitating secure program execution with minimal overheads
Divya Arora 0001, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.4
2006 A Scalable Synthesis Methodology for Application-Specific Processors
abstract
Custom processors based on application-specific or domain-specific instruction sets are gaining popularity, and are often used to implement critical architectural blocks in complex systems-on-chip. While several advances have been made in the area of custom processor architectures, tools, and design methodologies, designers are still required to manually perform some critical tasks, such as selection of the custom instructions best suited to the given application and design constraints. We present a scalable methodology for the synthesis of a custom processor from an embedded software program. A key feature of the proposed methodology is its scalability, which is achieved by exploiting the structured, hierarchical nature of large software programs. We motivate the need for such a methodology, and describe the algorithms used for the critical steps, including hardware resource budgeting, local optimizations, and global exploration. Our methodology utilizes the concept of "soft" instruction templates, which can be adapted by adding operations to them or deleting operations from them at any time during the design space exploration process, allowing for global design decisions to be interleaved with fine-grained optimizations. To the best of our knowledge, this is the first work that uses the program hierarchy to derive soft instruction templates to synthesize application-specific processors for scalable applications. We have integrated our methodology in an open-source compiler, and verified it using a commercial extensible processor. Experiments with several benchmarks indicate that our methodology can effectively tackle large programs. It results in the synthesis of high-quality custom processors that demonstrate an average speedup of 2.82times and a maximum speedup of 6.07times. As a side-effect, the processor energy is also reduced. The average and maximum reduction in the energy-delay product for the benchmarks are 7.64times and 18.85times, respectively. The CPU times required for custom processor synthesis are quite small, indicating that the proposed techniques can be applied to embedded software programs of significant complexity
Fei Sun 0002, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.4
2005 Efficient fingerprint-based user authentication for embedded systems
abstract
User authentication, which refers to the process of verifying the identity of a user, is becoming an important security requirement in various embedded systems. While conventional solutions for user authentication have relied on password-based mechanisms, they are increasingly being replaced by biometric technologies such as fingerprint, face, and voice recognition, which are known to provide higher levels of security for user authentication. This paper investigates the problem of supporting efficient fingerprint-based user authentication in embedded systems. For improving the performance of fingerprint-based authentication, we propose hardware/software enhancements that include a generic set of custom instruction extensions to an embedded processor's instruction set architecture, a memory-aware software re-design, and fixed-point arithmetic. We believe that the custom instruction set extensions proposed in this work are generic enough to speed up many fingerprint matching algorithms and even other biometric algorithms. Our experiments with an open-source, high-fidelity fingerprint authentication algorithm and a testbed featuring a commercial extensible processor show that performance is improved by a factor of 10.4X when using the proposed enhancements, while incurring modest overheads.
Pallav Gupta, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
DAC4
2005 Hybrid simulation for embedded software energy estimation
abstract
Software energy estimation is a critical step in the design of energyefficient embedded systems. Instruction-level simulation techniques, despite several advances, remain too slow for iterative use in systemlevel exploration. In this paper, we propose a methodology called hybrid simulation, which combines instruction set simulation with selective native execution (execution of some parts of the program directly on the simulation host computer), thereby overcoming the disadvantages of instruction-level simulation (low speed) and pure native execution (estimation accuracy, inapplicability to targetdependent code), while exploiting their advantages. Previously developed techniques for software energy macromodeling are utilized to estimate energy consumption for natively executed sub-programs. We identify and address the main challenges involved in hybrid simulation, and present an automatic tool flow for it, which analyzes a given program and selects functions for native execution in order to achieve maximum estimation efficiency while limiting estimation error. We have applied the proposed hybrid simulation methodology to a variety of embedded software programs, resulting in an average speed-up of 70% and estimation error of at most 6%, compared to one of the fastest publicly-available instruction set simulators.
Anish Muttreja, Anand Raghunathan, Srivaths Ravi 0001, Niraj K. Jha
DAC4
2005 User-perceived latency driven voltage scaling for interactive applications
abstract
Power has become a critical concern for battery-driven computing systems, on which many applications that are run are interactive. System-level voltage scaling techniques, such as dynamic voltage scaling (DVS) and adaptive body biasing (ABB), have been shown to reduce energy consumption effectively. Previous works on DVS and ABB exploit low CPU utilization of the processor to drive voltage scaling. This has become inadequate for modern interactive applications involving high CPU usage. In this work, we target computer responsiveness during voltage scaling to exploit more opportunities for energy reduction. Instead of CPU utilization, we use the user-perceived latency, the delay between user input and computer response, to drive voltage scaling. Considering the tradeoff between energy consumption and computer responsiveness during voltage scaling not only reduces energy consumption effectively, but also ensures good computer responsiveness for interactive applications. Experimental results show that for the 70nm technology, during the execution of seven commonly-used interactive applications, the energy consumption of the processor using user-perceived latency driven DVS is reduced by an average of 37.3%, and the user-perceived latency by an average of 18.3%, compared to CPU utilization driven DVS. If both DVS and ABB are performed simultaneously based on the user-perceived latency, then the energy consumption is reduced by another 38.9% compared to when DVS is performed alone, while maintaining a similar computer responsiveness level. We have implemented user-perceived latency driven voltage scaling under Linux with X Window system. However, the methodology is extensible to other operating systems as well.
Le Yan 0003, Lin Zhong 0001, Niraj K. Jha
DAC3
2005 Secure Embedded Processing through Hardware-Assisted Run-Time Monitoring
abstract
Security is emerging as an important concern in embedded system design. The security of embedded systems is often compromised due to vulnerabilities in "trusted" software that they execute. Security attacks exploit these vulnerabilities to trigger unintended program behavior, such as the leakage of sensitive data or the execution of malicious code. In this work, we present a hardware-assisted paradigm to enhance embedded system security by detecting and preventing unintended program behavior. Specifically, we extract properties of an embedded program through static program analysis, and use them as the bases for enforcing permissible program behavior in real-time as the program executes. We present an architecture for hardware-assisted run-time monitoring, wherein the embedded processor is augmented with a hardware monitor that observes the processor's dynamic execution trace, checks whether the execution trace falls within the allowed program behavior, and flags any deviations from the expected behavior to trigger appropriate response mechanisms. We present properties that can be used to capture permissible program behavior at different levels of granularity within a program, namely inter-procedural control flow, intra-procedural control flow, and instruction stream integrity. We also present a systematic methodology to design application-specific hardware monitors for any given embedded program. We have evaluated the hardware requirements and performance of the proposed architecture for several embedded software benchmarks. Hardware implementations using a commercial design flow, and architectural simulations using the SimpleScalar framework, indicate that the proposed technique can thwart several common software and physical attacks, facilitating secure program execution with minimal overheads.
Divya Arora 0001, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
DATE4
2005 Nanotechnology in the Service of Embedded and Ubiquitous Computing
Niraj K. Jha
EUC1
2005 ALLCN: An Automatic Logic-to-Layout Tool for Carbon Nanotube Based Nanotechnology
abstract
Since rapid progress has been made in device improvement and integration of small carbon nanotube field-effect transistors (CNFETs) circuits, the time has come for developing computer-aided design (CAD) methodologies and tools for the design of larger CNFET circuits. In this paper, we present the first automatic logic-to-layout (ALLCN) tool for CNFET circuits. The main purpose of this work is to bridge the wide gap that currently exists between research on the development of nanoscale devices and design tools for such devices. ALLCN is built on top of existing CAD tools including Magic, TimberWolf and YACR. It can automatically generate a CNFET circuit layout from a logic implementation and then perform circuit extraction from the physical layout for SPICE simulation. Experiments were performed with various MCNC benchmarks and logic blocks. Their performance, area and power are reported.
Wei Zhang 0012, Niraj K. Jha
ICCD2
2005 Towards a Responsive, Yet Power-ef.cient, Operating System: A Holistic Approach
abstract
Although computing hardware has become increasingly more powerful, computer responsiveness is still an important issue due to multi-tasking and software bloat. We propose a holistic approach for improving computer responsiveness through user focus-aware resource management for CPU, memory, disk I/O, and graphics processing. Previous approaches only address one or two of these problems simultaneously. To the best of our knowledge, our work is the first to address disk I/O scheduling for better responsiveness. It also offers better solutions for the other problems. We also exploit the user-perceived latency to perform dynamic voltage scaling of the CPU to reduce power consumption at run-time without sacrificing responsiveness. We implemented our approach in the Linux/X Window system (henceforth referred to as Linux/X) on an IBM Thinkpad R32 laptop with mobile Pentium 4-M processor, which has two performance levels with different frequency/supply voltage settings: high (30.0W at 1.8GHz/1.3V) and low (20.8W at 1.2GHz/1.2V). Experimental results show that user focus-aware resource management achieves a significant improvement in computer responsiveness and some in energy efficiency. For example, for TuxRacer, a video racing game, with GpsDrive, a navigation system, running in the background, it provides a reduction of 42.0% in user-perceived latency and 7.5% in energy consumption with respect to the Linux/X system.
Le Yan 0003, Lin Zhong 0001, Niraj K. Jha
MASCOTS3
2005 A personal-area network of low-power wireless interfacing devices for handhelds: system and hardware design
abstract
Handhelds, such as smart-phones and Pocket PCs, have the potential to become the computing, storage, and connectivity hub, or Digital Hub, for pervasive computing. However, their current interfacing paradigms fall short of achieving this goal. To meet this challenge, we present the system and hardware design for a Bluetooth-based personal-area network (PAN) of low-power wireless interfacing devices. The network consists of a wrist-watch, single-hand single-tap multi-finger keypad, smart speech portal, and GPS receiver. These devices serve a handheld in a synergistic fashion, collectively providing the user with immediate and more natural access to computing power and enabling more and better services.
Lin Zhong 0001, Mike Sinclair, Niraj K. Jha
Mobile HCI3
2005 Energy efficiency of handheld computer interfaces: limits, characterization and practice
abstract
Energy efficiency has become a critical issue for battery-driven computers. Significant work has been devoted to improving it through better software and hardware. However, the human factors and user interfaces have often been ignored. Realizing their extreme importance, we devote this work to a comprehensive treatment of their role in determining and improving energy efficiency. We analyze the minimal energy requirements and overheads imposed by known human sensory/speed limits. We then characterize energy efficiency for state-of-the-art interfaces available on two commercial handheld computers. Based on the characterization, we offer a comparative study for them.Even with a perfect user interface, computers will still spend most of their time and energy waiting for user responses due to an increasingly large speed gap between users and computers in their interactions. Such a speed gap leads to a bottleneck in system energy efficiency. We propose a low-power low-cost cache device, to which the host computer can outsource simple tasks, as an interface solution to overcome the bottleneck. We discuss the design and prototype implementation of a low-power wireless wrist-watch for use as a cache device for interfacing.With this work, we wish to engender more interest in the mobile system design community in investigating the impact of user interfaces on system energy efficiency and to harvest the opportunities thus exposed.
Lin Zhong 0001, Niraj K. Jha
MobiSys2
2005 Unsatisfiability Based Efficient Design for Testability Solution for Register-Transfer Level Circuits
abstract
In this paper, we present a novel and accurate method for identifying design for testability (DFT) solutions for register-transfer level (RTL) circuits. In this technique, clauses are generated using a satisfiability (SAT) based automatic test pattern generation (ATPG) tool to represent the control and data flow for a module under test in the given RTL circuit. RTL test generation makes use of the concept of pre-computed test sets for different RTL modules. The generated clauses corresponding to different pre-computed test vectors are then resolved by a SAT solver to obtain the test sequences for that module. In case of an unsatisfiable (UNSAT) solution, recent advances in the field of satisfiability enable us to accurately and efficiently identify clauses that are responsible for unsatisfiability (also known as the unsatisfiable segment). We show that adding DFT elements is equivalent to modifying clauses such that the unsatisfiable segment becomes satisfiable. In order to minimize the number of DFT elements added to a circuit, a greedy algorithm is used to select circuit variables for DFT such that all the unsatisfiable segments become satisfiable. Unlike existing DFT techniques that are either inefficient in terms of the amount of test hardware added or take significant time to identify an efficient solution, the proposed DFT technique is both fast and accurate as it is applicable to RTL and mixed gate-level/RTL circuits and uses UNSAT to identify the DFT solutions. Experimental results on benchmarks show that for RTL circuits, the CPU time required to identify pre-computed test vectors for which the SAT ATPG fails to generate test sequences and to select DFT solutions for such cases is two orders of magnitude smaller than the time required for a single run of a gate-level sequential test generator. The DFT solution has very low area overhead (an average of 1.7%) and results in near-100% fault coverage.
Loganathan Lingappan, Niraj K. Jha
VTS2
2005 Generation of distributed logic-memory architectures through high-level synthesis
abstract
With the increasing cost of on-chip global communication, high-performance designs for data-intensive applications require architectures that distribute hardware resources (computing logic, memories, interconnect, etc.) throughout the chip, while restricting computations and communications to geographic proximities. In this paper, we present a methodology for high-level synthesis (HLS) of distributed logic-memory architectures, i.e., architectures that have logic and memory distributed across several partitions in a chip. Conventional HLS tools are capable of extracting parallelism from a behavior for architectures that assume a monolithic controller/datapath communicating with a memory or memory hierarchy. This paper provides techniques to extend the synthesis frontier to more general architectures that can extract both coarse and fine-grained parallelism from data accesses and computations in a synergistic manner. Our methodology selects many possible ways of organizing data and computations, carefully examines the tradeoffs (i.e., communication overheads, synchronization costs, area overheads) in choosing one solution over another, and utilizes conventional HLS techniques for intermediate steps. We have evaluated the proposed framework on several benchmarks by generating register-transfer level (RTL) implementations using an existing commercial HLS tool with and without our enhancements, and by subjecting the resulting RTL circuits to logic synthesis and layout. The results show that circuits designed as distributed logic-memory architectures using our framework achieve significant (up to 5.3/spl times/, average of 3.5/spl times/) performance improvements over well-optimized conventional designs with small area overheads (up to 19.3%, 15.1% on average). At the same time, the reduction in the energy-delay product is by an average of 5.9/spl times/ (up to 11.0/spl times/).
Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2005 Input space-adaptive optimization for embedded-software synthesis
abstract
This paper presents a technique for exploiting input statistics for energy and performance optimization of embedded software. The proposed technique is based on the fact that the computational complexities of programs or subprograms are often highly dependent on the values assumed by input and intermediate program variables during execution. This observation is exploited in the proposed software synthesis technique by augmenting the program with optimized versions of one or more subprograms that are specialized to, and executed under, specific input subspaces. We propose a methodology for input space-adaptive software synthesis that consists of the following steps: 1) control and value profiling of the input program; 2) application of compiler transformations in a preprocessing step; 3) identification of subprograms and corresponding input subspaces that hold the highest potential for optimization; and 4) an iterative application of known compiler transformations to realize performance and energy savings. We propose novel metrics based on the entropies of program variables to characterize subprograms and input subspaces that hold significant potential for optimization. The chosen subprograms are optimized by translating the input subspaces into value constraints on their variables, and iteratively applying known compiler transformations (that were not applicable in the context of the original program). We have evaluated input space-adaptive software synthesis by compiling the resulting optimized programs to two commercial embedded systems: an embedded system based on the Fujitsu SPARClite processor, and the Compaq iPAQ personal digital assistant (PDA) [64 MB memory, 206 MHz Intel StrongARM central processing unit (CPU)]. The energy and execution-time savings were calculated using energy-aware instruction-level simulators, as well as through direct-current measurement on the iPAQ. Our results demonstrate that the proposed technique can reduce energy by up to 54.5% (average of 30.6% and 25.6% for the SPARClite-based system and the iPAQ, respectively) while simultaneously improving performance by up to 59.6% (average of 31.3% and 31.5% for the SPARClite-based system and the iPAQ, respectively). In effect, improvements in the energy-delay product of up to 81.1% (average of 51.0% and 47.7% for the SPARClite-based system and the iPAQ, respectively) were observed. The energy savings resulting from our technique are fairly processor independent, and complementary to conventional compiler optimizations.
Anand Raghunathan, Ganesh Lakshminarayana, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2005 Joint dynamic voltage scaling and adaptive body biasing for heterogeneous distributed real-time embedded systems
abstract
While dynamic power consumption has traditionally been the primary source of power consumption, leakage power is becoming an increasingly important concern as technology feature size continues to shrink. Previous system-level approaches focus on reducing power consumption without considering leakage power consumption. To overcome this limitation, we propose a two-phase approach to combine dynamic voltage scaling (DVS) and adaptive body biasing (ABB) for distributed real-time embedded systems. DVS is a powerful technique for reducing dynamic power consumption quadratically. However, DVS often requires a reduction in the threshold voltage that increases subthreshold leakage current exponentially and, hence, subthreshold leakage power consumption. ABB, which exploits the exponential dependence of subthreshold leakage power on the threshold voltage, is effective in managing leakage power consumption. We first derive an energy consumption model to determine the optimal supply voltage and body bias voltage under a given clock frequency. Then, we analyze the tradeoff between energy consumption and clock period to allocate slack to a set of tasks with precedence relationships and real-time constraints. Based on this two-phase approach, we propose a new system-level scheduling algorithm that can optimize both dynamic power and leakage power consumption by performing DVS and ABB simultaneously for distributed real-time embedded systems. Experimental results show that the average power reduction of our technique with respect to DVS alone is 37.4% for the 70-nm technology.
Le Yan 0003, Jiong Luo, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2005 Threshold network synthesis and optimization and its application to nanotechnologies
abstract
We propose an algorithm for efficient threshold network synthesis of arbitrary multioutput Boolean functions. Many nanotechnologies, such as resonant tunneling diodes, quantum cellular automata, and single electron tunneling, are capable of implementing threshold logic efficiently. The main purpose of this work is to bridge the current wide gap between research on nanoscale devices and research on synthesis methodologies for generating optimized networks utilizing these devices. While functionally-correct threshold gates and circuits based on nanotechnologies have been successfully demonstrated, there exists no methodology or design automation tool for general multilevel threshold network synthesis. We have built the first such tool, threshold logic synthesizer (TELS), on top of an existing Boolean logic synthesis tool. Experiments with 56 multioutput benchmarks indicate that, compared to traditional logic synthesis, up to 80.0% and 70.6% reduction in gate count and interconnect count, respectively, is possible with the average being 22.7% and 12.6%, respectively. Furthermore, the synthesized networks are well-balanced structurally. The novelty of this work lies in the introduction of the first comprehensive synthesis methodology and tool for general multilevel threshold logic design.
Pallav Gupta, Lin Zhong 0001, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2005 Interconnect-aware low-power high-level synthesis
abstract
Interconnects (wires, buffers, clock distribution networks, multiplexers, and busses) consume a significant fraction of total circuit power. In this paper, we demonstrate the importance of optimizing on-chip interconnects for power during high-level synthesis. We present a methodology to integrate interconnect power optimization into high-level synthesis. It not only reduces datapath unit power consumption in the resultant register-transfer level architecture, but also optimizes interconnects for power. We take into account physical design information and coupling capacitance to estimate interconnect power consumption accurately for deep submicron technologies. We show that there is significant spurious (i.e., unnecessary) switching activity in the interconnects and propose techniques to reduce it. Compared with interconnect-unaware power-optimized circuits, interconnect power can be reduced by 53.1% on average, while overall power is reduced by an average of 26.8%, with negligible area overhead. Compared with area-optimized circuits, the interconnect power reduction is 72.9% and overall power reduction is 56.0%, with 44.4% area overhead. The power reductions are obtained solely through switched capacitance reduction (no voltage scaling is assumed).
Lin Zhong 0001, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 Energy macromodeling of embedded operating systems
abstract
As embedded systems get more complex, deployment of embedded operating systems (OSs) as software run-time engines has become common. In particular, this trend is true even for battery-powered embedded systems, where maximizing battery life is a primary concern. In such OS-driven embedded software, the overall energy consumption depends very much on which OS is used and how the OS is used. Therefore, the energy effects of the OS need to be studied in order to design low-energy systems effectively.In this paper, we discuss the motivation for performing OS energy characterization and propose a methodology to perform the characterization systematically. The methodology consists of two parts. The first part is analysis , which is concerned with identifying a set of components that can be used to characterize the OS energy consumption, called energy characteristics . The second part is macromodeling , which is concerned with obtaining quantitative macromodels for the energy characteristics. It involves the process of experiment design, data collection, and macromodel fitting. The OS energy macromodels can be used conveniently as OS energy estimators in high-level or architectural optimization of embedded systems for low-energy consumption.As far as we know, this work is the first attempt to systematically tackle energy macromodeling of an embedded OS. To demonstrate our approach, we present experimental results for two well-known embedded OSs, namely, μC/OS and embedded Linux OS.
Tat Kee Tan, Anand Raghunathan, Niraj K. Jha
ACM Trans. Embed. Comput. Syst.3
2005 Memory binding for performance optimization of control-flow intensive behavioral descriptions
abstract
This paper presents a memory binding algorithm for behaviors, used in application-specific integrated circuits (ASICs), that are characterized by the presence of conditionals and deeply nested loops that access memory extensively through arrays. Unlike previous works, this algorithm examines the effects of branch probabilities and allocation constraints. First, we demonstrate, through examples, the importance of incorporating branch probabilities and allocation constraint information when searching for a performance-efficient memory binding. We also show the interdependence of these two factors and how varying one without considering the other may greatly affect the performance of the behavior. Second, we introduce a memory binding algorithm that has the ability to examine numerous bindings by employing an efficient performance estimation procedure. The estimation procedure exploits locality of execution, which is an inherent characteristic of target behaviors. This enables the performance estimation technique to look at the global impact of the different bindings, given the allocation constraints. We tested our algorithm using a number of benchmarks from the parallel computing domain. A series of experiments demonstrates the algorithm's ability to produce bindings that optimize performance, meet memory allocation constraints, and adapt to different resource constraints and branch probabilities. One limitation of our algorithm is that, in its current form, it is not well suited for system-on-a-chip synthesis where there is complex communication between general-purpose microprocessors that use custom-designed arrays. Results show that the algorithm requires 41% fewer memories with a performance loss of only 0.2% when compared to a parallel memory architecture. When compared to the best of a series of random memory bindings, the algorithm improves schedule performance by 22%.
Kamal S. Khouri, Ganesh Lakshminarayana, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
2004 Automated energy/performance macromodeling of embedded software
abstract
Efficient energy and performance estimation of embedded software is a critical part of any system-level design flow. Macromodeling based estimation is an attempt to speed up estimation by exploiting reuse that is inherent in the design process. Macromodeling involves pre-characterizing reusable software components to construct high-level models, which express the execution time or energy consumption of a sub-program as a function of suitable parameters. During simulation, macromodels can be used instead of detailed hardware models, resulting in orders of magnitude simulation speedup. However, in order to realize this potential, significant challenges need to be overcome in both the generation and use of macromodels--- including how to identify the parameters to be used in the macromodel, how to define the template function to which the macromodel is fitted, em etc. This paper presents an automatic methodology to perform characterization-based high-level software macromodeling, which addresses the aforementioned issues. Given a sub-program to be macromodeled for execution time and/or energy consumption, the proposed methodology automates the steps of parameter identification, data collection through detailed simulation, macromodel template selection, and fitting. We propose a novel technique to identify potential macromodel parameters and perform data collection, which draws from the concept of bf data structure serialization used in distributed programming. We utilize bf symbolic regression techniques to concurrently filter out irrelevant macromodel parameters, construct a macromodel function, and derive the optimal coefficient values to minimize fitting error. Experiments with several realistic benchmarks suggest that the proposed methodology improves estimation accuracy and enables wide applicability of macromodeling to complex embedded software, while realizing its potential for estimation speedup. We describe a case study of how macromodeling can be used to rapidly explore algorithm-level energy tradeoffs, for the tt zlib data compression library.
Anish Muttreja, Anand Raghunathan, Srivaths Ravi 0001, Niraj K. Jha
DAC4
2004 Synthesis of Reversible Logic
abstract
A function is reversible if each input vector produces a unique output vector. Reversible functions find applications in low power design, quantum computing, and nanotechnology. Logic synthesis for reversible circuits differs substantially from traditional logic synthesis. In this paper, we present the first practical synthesis algorithm and tool for reversible functions with a large number of inputs. It uses positive-polarity Reed-Muller decomposition at each stage to synthesize the function as a network of Toffoli gates. The heuristic uses a priority queue based search tree and explores candidate factors at each stage in order of attractiveness. The algorithm produces near-optimal results for the examples discussed in the literature. The key contribution of the work is that the heuristic finds very good solutions for reversible functions with a large number of inputs.
Abhinav Agrawal 0002, Niraj K. Jha
DATE2
2004 An Algorithm for Nano-Pipelining of Circuits and Architectures for a Nanotechnology
abstract
In this paper, we describe an algorithm to post-process a register-transfer level (RTL) architecture to enable gate-level pipelining or nano-pipelining for the nanotechnology based on resonant tunneling diodes (RTDs). Nano-pipelining offers the opportunity to obtain massive throughput and, therefore, has applications in data-intensive algorithms such as digital signal processing (DSP). Since RTDs are a self-latching nanotechnology, nano-pipelining is an implicit property that should be exploited for this technology. The novelty of this work lies in exploring and demonstrating the benefits of nano-pipelining and presenting an algorithm for architectural nano-pipelining.
Pallav Gupta, Niraj K. Jha
DATE2
2004 Synthesis and Optimization of Threshold Logic Networks with Application to Nanotechnologies
abstract
We propose an algorithm for efficient threshold network synthesis of arbitrary multi-output Boolean functions. The main purpose of this work is to bridge the wide gap that currently exists between research on the development of nanoscale devices and research on the development of synthesis methodologies to generate optimized networks utilizing these devices. Many nanotechnologies, such as resonant tunneling diodes (RTD) and quantum cellular automata (QCA), are capable of implementing threshold logic. While functionally correct threshold gates have been successfully demonstrated, there exists no methodology or design automation tool, threshold logic synthesizer (TELS), on top of an existing Boolean logic synthesis tool. Experiments with about 60 multi-output benchmarks were performed, though the results of only 10 of them are reported in this paper because of space restriction. They indicate that up to 77% reduction in gate count is possible when utilizing threshold logic, with an average reduction being 52%, compared to traditional logic synthesis. Furthermore, the synthesized networks are well-balanced, and hence delay-optimized.
Pallav Gupta, Lin Zhong 0001, Niraj K. Jha
DATE4
2004 High-level synthesis using computation-unit integrated memories
abstract
High-level synthesis (HLS) of memory-intensive applications has featured several innovations in terms of enhancements made to the basic memory organization and data layout. However, increasing performance and energy demands faced by application-specific integrated circuits (ASIC) are forcing designers to alter the fundamental architectural template of the HLS output, namely, a controller-datapath associated with a memory subsystem (monolithic, banked, etc.). We propose an architectural template for the HLS output that consists of a controller-datapath circuit associated with a memory subsystem into which computation units have been integrated. The enhanced memory subsystem is called computation-unit integrated memory (CIM). A CIM offers higher memory bandwidth (relative to what is offered through the system bus) to computation units present locally within it and reduces the overall communication between the memory subsystem and the controller-datapath, thus providing a template highly suitable for deriving efficient implementations of memory-intensive applications. This work addresses the challenge of providing an automatic synthesis framework for a CIM-based architecture. Our framework can analyze the various trade-offs involved in selecting suitable operations in a behavior for execution using a CIM and generate a high-performance, low-overhead implementation. Experiments with several behaviors indicate that an average performance improvement of 1.88/spl times/ (a maximum of 2.63/spl times/) is possible with very low area overheads. The energy-delay product improves by an average of 2.1/spl times/ (maximum of 3.4/spl times/).
Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
ICCAD4
2004 Power estimation for cycle-accurate functional descriptions of hardware
abstract
Cycle-accurate functional descriptions (CAFD) are being widely adopted in integrated circuit (IC) design flows. Power estimation can potentially benefit from the inherent increase in simulation efficiency of cycle-based functional simulation. Currently, most approaches to hardware power estimation operate at the register-transfer level (RTL), or lower levels of design abstraction. Attempts at power estimation for functional descriptions have suffered from poor accuracy because the design decisions performed during their synthesis lead to an unavoidable, large uncertainty in any power estimate that is based solely on the functional description. We propose a methodology for CAFD power estimation that combines the accuracy achieved by power estimation at the structural RTL with the efficiency of cycle-accurate functional simulation. We achieve this goal by viewing a CAFD as an abstraction of a specific, known RTL implementation that is synthesized from it. We identify correlations between a CAFD and its RTL implementation, and "back-annotate" information into the CAFD solely for the purpose of power estimation. The resulting RTL-aware CAFD contains a layer of code that instantiates virtual placeholders for RTL components, and maps values of CAFD variables into the RTL components' inputs/outputs, thus enabling efficient and accurate power estimation. Power estimation is performed in our methodology by simply co-simulating the RTL-aware CAFD with a simulatable power model library that contains power macro-models for each RTL component. We present techniques to further improve the speed of CAFD power estimation, through the use of control state-based adaptive power sampling. We have implemented and evaluated the proposed techniques in the context of a commercial C-based hardware design flow. Experiments with a number of large industrial designs (up to 1 million gates) demonstrate that the proposed methodology achieves accuracy very close to RTL power estimation with two-to-three orders of magnitude speedup in estimation times.
Lin Zhong 0001, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
ICCAD4
2004 An Automatic Test Pattern Generation Framework for Combinational Threshold Logic Networks
abstract
We propose an automatic test pattern generation (ATPG) framework for combinational threshold networks. The motivation behind this work lies in the fact that many emerging nanotechnologies, such as resonant tunneling diodes (RTDs) and quantum cellular automata (QCA), implement threshold logic. Consequently, there is a need to develop an ATPG methodology for this type of logic. We have built the first automatic test pattern generator and fault simulator for threshold logic, which has been integrated on top of an existing computer-aided design (CAD) tool. These exploit new fault collapsing techniques we have developed for threshold networks. We perform fault modeling to show that many cuts and shorts in RTD-based threshold gates are equivalent to stuck-at faults at the inputs and output of the gate. Experimental results with the MCNC benchmarks indicate that test vectors were found for all testable stuck-at faults in their threshold network implementations.
Pallav Gupta, Niraj K. Jha
ICCD3
2004 Thermal Modeling, Characterization and Management of On-Chip Networks
abstract
Due to the wire delay constraints in deep submicron technology and increasing demand for on-chip bandwidth, networks are becoming the pervasive interconnect fabric to connect processing elements on chip. With ever-increasing power density and cooling costs, the thermal impact of on-chip networks needs to be urgently addressed. In this work, we first characterize the thermal profile of the MIT Raw chip. Our study shows networks having comparable thermal impact as the processing elements and contributing significantly to overall chip temperature, thus motivating the need for network thermal management. The characterization is based on an architectural thermal model we developed for on-chip networks that takes into account the thermal correlation between routers across the chip and factors in the thermal contribution of on-chip interconnects. Our thermal model is validated against finite-element based simulators for an actual chip with associated power measurements, with an average error of 5.3%. We next propose ThermalHerd, a distributed, collaborative run-time thermal management scheme for on-chip networks that uses distributed throttling and thermal-correlation based routing to tackle thermal emergencies. Our simulations show ThermalHerd effectively ensuring thermal safety with little performance impact. With Raw as our platform, we further show how our work can be extended to the analysis and management of entire on-chip systems, jointly considering both processors and networks.
Li-Shiuan Peh, Amit Kumar 0002, Niraj K. Jha
MICRO4
2004 COWLS: hardware-software cosynthesis of wireless low-power distributed embedded client-server systems
abstract
In this paper, we present COWLS, a hardware-software cosynthesis algorithm that targets embedded systems composed of servers and low-power clients that communicate with each other through a channel of limited bandwidth, e.g., a wireless link. A novel scheduling algorithm is used to pipeline the execution of tasks that serve multiple clients associated with a given server. COWLS simultaneously optimizes the price of the client-server system, the power consumption of the clients, and the response times of tasks that have only soft deadlines, while meeting all of the hard deadlines. It produces numerous solutions that trade off different architectural features, e.g., price, power consumption, and response time, of an embedded client-server system. As far as we know, this is the first synthesis algorithm of its kind. We present the experimental results for numerous pseudorandom examples, a low-power client-server camera system, as well as the rest of the benchmarks within a publicly released embedded system synthesis benchmark suite.
Robert P. Dick, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2004 A hybrid energy-estimation technique for extensible processors
abstract
In this paper, we present an efficient and accurate methodology for estimating the energy consumption of application programs running on extensible processors. Extensible processors, which are getting increasingly popular in embedded system design, allow a designer to customize a base processor core through instruction set extensions. Existing processor energy macromodeling techniques are not applicable to extensible processors, since they assume that the instruction set architecture as well as the underlying structural description of the micro-architecture remain fixed. Our solution to the above problem is a hybrid energy macromodel suitably parameterized to estimate the energy consumption of an application running on the corresponding application-specific extended processor instance, which incorporates any custom instruction extension. Such a characterization is facilitated by careful selection of macromodel parameters/variables that can capture both the functional and structural aspects of the execution of a program on an extensible processor. Another feature of the proposed energy characterization flow is the use of regression analysis to build the macromodel. Regression analysis allows for in-situ characterization, thus allowing arbitrary test programs to be used during macromodel construction. We validated the proposed methodology by characterizing the energy consumption of a state-of-the-art extensible processor (Tensilica's Xtensa). We used the macromodel to analyze the energy consumption of several benchmark applications with custom instructions. The mean absolute error in the macromodel estimates is only 3.3%, when compared to the energy values obtained by a commercial tool operating on the synthesized register-transfer level (RTL) description of the custom processor. Our approach achieves an average speedup of three orders of magnitude over the commercial RTL energy estimator. Our experiments show that the proposed methodology also achieves good relative accuracy, which is essential in energy optimization studies. Hence, our technique is both efficient and accurate.
Yunsi Fei, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2004 Common-case computation: a high-level energy and performance optimization technique
abstract
This paper proposes a novel circuit design methodology, called common-case computation (CCC)-based design, and new design automation algorithms for optimizing energy consumption and performance. The proposed techniques are applicable in conjunction with any high-level design methodology, where a structural register-transfer level (RTL) description and its corresponding scheduled behavioral (cycle-accurate functional) description are available. It is a well-known fact that in behavioral descriptions of hardware circuits (and also in software programs), a small set of computations often account for most of the computational complexity. However, in the hardware implementations (structural RTL or lower level), the common cases and the remaining computations are typically treated alike. This paper shows that identifying and exploiting common cases during the design process can lead to implementations that are much more efficient in terms of energy consumption and performance.
Ganesh Lakshminarayana, Anand Raghunathan, Kamal S. Khouri, Niraj K. Jha, Sujit Dey
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2004 Register binding-based RTL power management for control-flow intensive designs
abstract
One important way to reduce power consumption is to reduce the spurious switching activity in a circuit or circuit component, i.e., activity that is not required by its specified functionality. Given a scheduled behavior and functional unit binding, we show that spurious switching activity can be reduced through proper register binding using retentive multiplexers. Retentive multiplexers can preserve their previous select signal values in the control steps in which the select signals are don't cares. A functional unit, in which spurious switching activity is completely eliminated, is called perfectly power managed. We present a general sufficient condition for register binding to ensure a set of functional units to be perfectly power managed. This condition not only applies to data-flow intensive behaviors, but also to control-flow intensive behaviors. It leads to a straightforward power-managed (PM) register-binding algorithm, which uses this condition to preserve the previous values in the input registers of a functional unit during the states in which the unit is idle. The proposed algorithm is general and independent of the functional unit binding and scheduling algorithms. Hence, it can be easily incorporated into existing high-level synthesis systems. For the benchmarks we experimented with, an average 40.7% power reduction was achieved by our method at the cost of 6.9% average area overhead, compared to power-optimized register-transfer level circuits, which did not use PM register binding.
Jiong Luo, Lin Zhong 0001, Yunsi Fei, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2004 Custom-instruction synthesis for extensible-processor platforms
abstract
Efficiency and flexibility are critical, but often conflicting, design goals in embedded system design. The recent emergence of extensible processors promises a favorable tradeoff between efficiency and flexibility, while keeping design turnaround times short. Current extensible processor design flows automate several tedious tasks, but typically require designers to manually select the parts of the program that are to be implemented as custom instructions. In this work, we describe an automatic methodology to select custom instructions to augment an extensible processor, in order to maximize its efficiency for a given application program. We demonstrate that the number of custom instruction candidates grows rapidly with program size, leading to a large design space, and that the quality (speedup) of custom instructions varies significantly across this space, motivating the need for the proposed flow. Our methodology features cost functions to guide the custom instruction selection process, as well as static and dynamic pruning techniques to eliminate inferior parts of the design space from consideration. Furthermore, we employ a two-stage process, wherein a limited number of promising instruction candidates are first short-listed using efficient selection criteria, and then evaluated in more detail through cycle-accurate instruction set simulation and synthesis of the corresponding hardware, to identify the custom instruction combinations that result in the highest program speedup or maximize speedup under a given area constraint. We have evaluated the proposed techniques using a state-of-the-art extensible processor platform, in the context of a commercial design flow. Experiments with several benchmark programs indicate that custom processors synthesized using automatic custom instruction selection can result in large improvements in performance (up to 5.4/spl times/, an average of 3.4/spl times/), energy (up to 4.5/spl times/, an average of 3.2/spl times/), and energy-delay products (up to 24.2/spl times/, an average of 12.6/spl times/), while speeding up the design process significantly.
Fei Sun 0002, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2004 Resource budgeting for Multiprocess High-level synthesis
abstract
This paper presents a new high-level synthesis methodology to generate optimized register-transfer level (RTL) implementations for multiprocess behavioral descriptions. The concurrent communicating processes specification paradigm is widely used in digital circuit and system design, and is employed in all popular hardware description languages. It has been shown that interprocess communication and synchronization can result in complex timing interdependencies, which significantly affect the performance of a multiprocess system. In this paper, we demonstrate that state-of-the-art high-level synthesis tools can generate significantly suboptimal implementations for behaviors that contain concurrent communicating processes. We present an analysis of how interprocess communication impacts high-level synthesis steps, and describe a new methodology to adapt existing high-level synthesis tools to optimize multiprocess descriptions. Our methodology is based on executing multiprocess performance analysis and process-by-process scheduling in an iterative manner. We present algorithms for key steps in the proposed methodology. We have performed extensive experiments in the context of a commercial high-level design flow to evaluate the proposed techniques. The results clearly demonstrate the utility of our techniques in synthesizing implementations with superior area, performance, and energy consumption. For example, up to 40% performance improvement (average of 35.6%) was achieved with little or no area overhead (average of 4.8%). In effect, the proposed techniques lead to a shift of the entire area-delay tradeoff curve for a design, to include superior designs that were hitherto infeasible. Our techniques also simultaneously result in up to 50% (average of 33.5%) improvement in energy and up to 69% (average of 58.3%) improvement in the energy-delay product.
Anand Raghunathan, Niraj K. Jha, Sujit Dey
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2004 DESP: A Distributed Economics-Based Subcontracting Protocol for Computation Distribution in Power-Aware Mobile Ad Hoc Networks
abstract
In this paper, we present a new economics-based power-aware protocol, called the distributed economic subcontracting protocol (DESP) that dynamically distributes task computation among mobile devices in an ad hoc wireless network. Mobile computation devices may be energy buyers, contractors, or subcontractors. Tasks are transferred between devices via distributed bargaining and transactions. When additional energy is required, buyers and contractors negotiate energy prices within their local markets. Contractors and subcontractors spend communication and computation energy to relay or execute buyers' tasks. Buyers pay the negotiated price for this energy. Decision-making algorithms are proposed for buyers, contractors, and subcontractors, each of which has a different optimization goal. We have built a wireless network simulator, called ESIM, to assist in the design and analysis of these algorithms. When the average communication energy required transferring a task is less than the average energy required to execute a task, our experimental results indicate that markets based on our protocol and decision-making algorithms fairly and effectively allocate energy resources among different tasks in both cooperative and competitive scenarios.
Robert P. Dick, Niraj K. Jha
IEEE Trans. Mob. Comput.3
2004 Input space adaptive design: a high-level methodology for optimizing energy and performance
abstract
This paper presents a high-level design methodology, called input space adaptive design, and new design automation algorithms for optimizing energy consumption and performance. Our techniques can be applied to behaviors described in hardware description languages, predesigned register-transfer level (RTL) circuits, or in the context of traditional high-level design methodologies. An input space adaptive design exploits the well-known fact that the quality of circuits can be significantly optimized by employing algorithms and implementation architectures that adapt to input statistics. This paper shows that harnessing the principles of input space adaptive design into a structured high-level design methodology can lead to large improvements in performance and energy consumption. We illustrate the tradeoffs involved in such designs, and demonstrate the need for a systematic design methodology in order to realize the full potential for performance and energy improvements. We propose a methodology for input space adaptive design that consists of the following steps: identification of parts of the behavior that hold the highest potential for optimization, selection of input subspaces whose occurrence can lead to significant reductions in implementation complexity, and transformation of the behavior to realize performance and/or energy savings. Evaluations of performance, energy, and area characteristics of input space adaptive designs in the context of a commercial high-level design flow indicate that such designs can reduce energy consumption by up to 58.9% (average of 40.0%), and simultaneously improve performance by up to 57.5% (average of 41.8%) compared to well-optimized designs that do not employ such techniques. The energy-delay product is reduced by up to 77.9% (average of 64.8%). When the performance improvements are translated into additional energy savings through supply voltage reduction, input space adaptive designs consume up to 74.2% (average of 68.8%) less energy at the same performance. The average area overhead is only 7.4%.
Anand Raghunathan, Ganesh Lakshminarayana, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.4
2003 Graphical user interface energy characterization for handheld computers
abstract
A significant fraction of the software and resource usage of a modern handheld computer is devoted to its graphical user interface (GUI). Moreover, GUIs are direct users of the display and also determine how users interact with software. Given that displays consume a significant fraction of system energy, it is very important to optimize GUIs for energy consumption. This work presents the first GUI energy characterization methodology. Energy consumption is characterized for three popular GUI platforms (Windows, X Window system, and Qt) from the hardware, software, and application perspectives. Based on this characterization, insights are offered for improving GUI platforms, and designing GUIs in an energy-efficient and aware fashion. Such a characterization also provides a firm basis for further research on GUI energy optimization.
Lin Zhong 0001, Niraj K. Jha
CASES2
2003 Energy Estimation for Extensible Processors
Yunsi Fei, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
DATE4
2003 Simultaneous Dynamic Voltage Scaling of Processors and Communication Links in Real-Time Distributed Embedded Systems
abstract
Dynamic voltage scaling has been widely acknowledged as a powerful technique for trading off power consumption and delay for processors. Recently, variable-frequency (and variable-voltage) parallel and serial links have also been proposed, which can save link power consumption by exploiting variations in bandwidth requirement. In this paper, we address joint dynamic voltage scaling for variable-voltage processors and communication links in such systems. We propose a scheduling algorithm for real-time applications, with both data flow and control flow information captured. It performs efficient routing of communication events through multi-hops, as well as efficient slack allocation among heterogeneous processors and communication links to maximize energy savings, while meeting all real-time constraints.
Jiong Luo, Li-Shiuan Peh, Niraj K. Jha
DATE3
2003 Software Architectural Transformations: A New Approach to Low Energy Embedded Software
Tat Kee Tan, Anand Raghunathan, Niraj K. Jha
DATE3
2003 A comprehensive high-level synthesis system for control-flow intensive behaviors
abstract
In this paper, we describe a comprehensive high-level synthesis system for control-flow intensive as well as data-dominated behaviors. We propose a new control-data flow graph model to preserve the parallelism inherent in the application, as well as to facilitate high-level synthesis. Our algorithm, which is based on an iterative improvement strategy, performs clock selection, scheduling, module selection, resource allocation and assignment simultaneously to fully derive the benefits of design space exploration at the behavior level. The system can be used to optimize area, power or energy, by selecting the cost function accordingly. Experimental results show that for energy-optimized designs, energy is reduced by up to 79.4% (an average of 42.2%), with an average of 24.8% area overhead, compared to area-optimized designs. For power-optimized designs, power is reduced by up to 70.8% (an average of 56.7%), with an average of 25.2% area overhead, compared to area-optimized designs. No Vdd scaling is performed to obtain the above results.
Tat Kee Tan, Jiong Luo, Yunsi Fei, Keith S. Vallerio, Lin Zhong 0001, Anand Raghunathan, Niraj K. Jha
ACM Great Lakes Symposium on VLSI9
2003 Dynamic Voltage Scaling with Links for Power Optimization of Interconnection Networks
abstract
Originally developed to connect processors and memories in multicomputers, prior research and design of interconnection networks have focused largely on performance. As these networks get deployed in a wide range of new applications, where power is becoming a key design constraint, we need to seriously consider power efficiency in designing interconnection networks. As the demand for network bandwidth increases, communication links, already a significant consumer of power now, will take up an ever larger portion of total system power budget. In this paper we motivate the use of dynamic voltage scaling (DVS) for links, where the frequency and voltage of links are dynamically adjusted to minimize power consumption. We propose a history-based DVS policy that judiciously adjusts link frequencies and voltages based on past utilization. Our approach realizes up to 6.3/spl times/ power savings (4.6/spl times/ on average). This is accompanied by a moderate impact on performance (15.2% increase in average latency before network saturation and 2.5% reduction in throughput.) To the best of our knowledge, this is the first study that targets dynamic power optimization of interconnection networks.
Li-Shiuan Peh, Niraj K. Jha
HPCA3
2003 A High-level Interconnect Power Model for Design Space Exploration
Pallav Gupta, Lin Zhong 0001, Niraj K. Jha
ICCAD3
2003 Synthesis of Heterogeneous Distributed Architectures for Memory-Intensive Applications
Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
ICCAD4
2003 A Scalable Application-Specific Processor Synthesis Methodology
Fei Sun 0002, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
ICCAD4
2003 Combined Dynamic Voltage Scaling and Adaptive Body Biasing for Heterogeneous Distributed Real-time Embedded Systems
Le Yan 0003, Jiong Luo, Niraj K. Jha
ICCAD3
2003 Test Generation for Non-separable RTL Controller-datapath Circuits using a Satisfiability based Approach
abstract
We present a satisfiability-based algorithm for automatically generating test sequences that target gate-level stuck-at faults in a circuit by using its register-transfer level (RTL) description. Our methodology exploits a popular, unified RTL circuit representation, called assignment decision diagrams, for its analysis and justifies module-level precomputed test vectors on this representation. Test generation proceeds by abstracting the components in this unified representation using input/output propagation rules, so that any justification/propagation event can be captured as a Boolean implication. Consequently, we reduce RTL test generation to a satisfiability (SAT) instance that has a significantly lower complexity than the equivalent problem at the gate-level. Using the state-of-the-art SAT solver ZCHAFF, we show that our RTL test generator can outperform gate-level sequential automatic test pattern generation (ATPG) in terms of both fault coverage and test generation time (two-to-three orders of magnitude speed-up), in comparable test application times. Furthermore, we show that in a bi-level testing scenario, in which RTL ATPG is followed by gate-level sequential ATPG on the remaining faults, we improve the fault coverage even further, while maintaining a high speed-up in test generation time (nearly 29X) over pure gate-level sequential ATPG, at comparable test application times.
Loganathan Lingappan, Srivaths Ravi 0001, Niraj K. Jha
ICCD3
2003 PowerHerd: dynamic satisfaction of peak power constraints in interconnection networks
abstract
Power consumption is a critical issue in interconnection network design, driven by power- related design constraints, such as thermal and power delivery design. Usually, off-line worst-case power analysis is used in network design to guarantee safe on-line operation, which not only increases system cost but also constrains network performance. In this work, we present an on-line mechanism, called PowerHerd, which can dynamically regulate network power consumption, and guarantee that network peak power constraints are not exceeded. PowerHerd is a distributed approach -- within the interconnection network, each router dynamically maintains a local power budget, controls its local power dissipation, and exchanges spare power resources with its neighboring routers to optimize network performance.Experiments demonstrate that PowerHerd can effectively regulate network power consumption meeting peak power constraints with negligible network performance penalty. Armed with PowerHerd, network designers can focus on system performance and power optimization for the average case rather than the worst case, thus making it possible to employ a more powerful interconnection network in the system.
Li-Shiuan Peh, Niraj K. Jha
ICS3
2003 Analyzing the energy consumption of security protocols
abstract
Security is critical to a wide range of wireless data applications and services. While several security mechanisms and protocols have been developed in the context of the wired Internet, many new challenges arise due to the unique characteristics of battery powered embedded systems. In this work, we focus on an important constraint of such devices -- battery life -- and examine how it is impacted by the use of security protocols.We present a comprehensive analysis of the energy requirements of a wide range of cryptographic algorithms that are used as building blocks in security protocols. Furthermore, we study the energy consumption requirements of the most popular transport-layer security protocol SSL (Secure Sockets Layer). To our knowledge, this is the first comprehensive analysis of the energy requirements of SSL. For our studies, we have developed a measurement-based experimental testbed that consists of an iPAQ PDA connected to a wireless LAN and running Linux, a PC-based data acquisition system for real-time current measurement, the OpenSSL implementation of the SSL protocol, and parametrizable SSL client and server test programs. We investigate the impact of various parameters at the protocol level (such as cipher suites, authentication mechanisms, and transaction sizes, etc.) and the cryptographic algorithm level (cipher modes, strength) on overall energy consumption for secure data transactions.Based on our results, we discuss various opportunities for realizing energy-efficient implementations of security protocols. We believe such investigations to be an important first step towards addressing the challenges of energy efficient security for battery-constrained systems.
Nachiketh R. Potlapally, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
ISLPED4
2003 Analysis of power dissipation in embedded systems using real-time operating systems
abstract
The increasing complexity and software content of embedded systems has led to the frequent use of system software to help applications access hardware resources easily and efficiently. In this paper, we present a method for detailed analysis of real-time operating system (RTOS) power consumption. RTOSs form an important component of the system software layer. Despite the widespread use of, and significant role played by, RTOSs in mobile and low-power embedded systems, little is known about their power-consumption effects. This paper presents a method of producing a hierarchical energy-consumption profile for applications as they interact with an RTOS. As a proof-of-concept, we use our infrastructure to produce the power profiles for a commercial RTOS, /spl mu/C/OS-II, running several applications on an embedded system based on the Fujitsu SPARClite processor. These examples demonstrate that an RTOS can have a significant impact on power consumption. We discuss ways in which application software can be designed to use an RTOS in a power-efficient manner. We believe that this is a first step toward establishing a systematic approach to power optimization of embedded systems containing RTOSs.
Robert P. Dick, Ganesh Lakshminarayana, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2003 A simulation framework for energy-consumption analysis of OS-driven embedded applications
abstract
Energy consumption has become a major focus in the design of embedded systems (e.g., mobile computing and wireless communication devices). In particular, a shift of emphasis from hardware-oriented low-energy design techniques to energy-efficient embedded software design has occurred progressively in the past few years. To that end, various techniques have been developed for the design of energy-efficient embedded software. In operating system (OS)-driven embedded systems, the OS has a significant impact on the system's energy consumption directly (energy consumption associated with the execution of the OS functions and services), as well as indirectly (interaction of the OS with the application software). As a first step toward designing energy-efficient OS-based embedded systems, it is important to analyze the energy consumption of embedded software by taking the OS energy characteristics into account. To facilitate such studies, we present, in this work, an energy simulation framework that can be used to analyze the energy consumption characteristics of an embedded system featuring the embedded Linux OS running on the StrongARM processor. The framework allows software designers to study the energy consumption of the system software in relation to the application software, identify the energy hot spots, and perform design changes based on the knowledge of the OS energy consumption characteristics as well as application-OS interactions.
Tat Kee Tan, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2003 High-level macro-modeling and estimation techniques for switching activity and power consumption
abstract
We present efficient techniques for estimating switching activity and power consumption at the register-transfer level (RTL), using a combination of macro-modeling for datapath blocks, and control logic analysis techniques based on partial delay information. Previous work on estimating switching activity and power at the RTL has ignored the presence of glitches at various datapath and control signals. We demonstrate that glitches can form a significant component of the switching activity at signals in typical RTL circuits. In particular, for control-flow intensive designs, we show that the controller substantially affects the activity and power consumption in the datapath due to the presence of glitches at control signals. Since the final implementation of the controller is not available during high-level design iterations, we develop techniques that estimate glitching activity at control signals using control expressions and partial delay information. For datapath blocks that operate on word-level data, we construct piecewise linear models that capture the variation of output glitching activity and power consumption with various word-level parameters like mean, standard deviation, spatial and temporal correlations, and glitching activity at the block's inputs. For RTL blocks that operate on bit vectors that need not have an associated word-level value, we present accurate bit-level modeling techniques for glitching activity as well as power consumption. This allows us to perform accurate power estimation for control-flow intensive circuits, where most of the power consumed is dissipated in non-arithmetic components like multiplexers, registers, vector logic operators, etc. Experimental results on several RTL designs demonstrate the accuracy of the proposed estimation techniques. Our RTL power estimator produced estimates that were within 7% of those produced by an in-house power analysis tool on the final gate-level implementation, while being over 50/spl times/ faster than its gate-level counterpart.
Anand Raghunathan, Sujit Dey, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
2002 Low Power Distributed Embedded Systems: Dynamic Voltage Scaling and Synthesis
Jiong Luo, Niraj K. Jha
HiPC2
2002 High-level synthesis of distributed logic-memory architectures
abstract
With the increasing cost of global communication on-chip, high-performance designs for data-intensive applications require architectures that distribute hardware resources (computing logic, memories, interconnect, etc.) throughout a chip, while restricting computations and communications to geographic proximities. In this paper, we present a methodology for high-level synthesis (HLS) of distributed logic-memory architectures, i.e., architectures that have logic and memory distributed across several partitions in a chip. Conventional HLS tools are capable of extracting parallelism from a behavior for architectures that assume a monolithic controller/datapath communicating with a memory or memory hierarchy. This work provides techniques to extend the synthesis frontier to more general architectures that can extract both coarse- and fine-grained parallelism from data accesses and computations in a synergistic manner. Our methodology selects many possible ways of organizing data and computations, carefully examines the trade-offs (i.e., communication overheads, synchronization costs, area overheads) in choosing one solution over another, and utilizes conventional HLS techniques for intermediate steps.We have evaluated the proposed framework on several benchmarks by generating register-transfer level (RTL) implementations using an existing commercial HLS tool with and without our enhancements, and by subjecting the resulting RTL circuits to logic synthesis and layout. The results show that circuits designed as distributed logic-memory architectures using our framework achieve significant (upto, 5.31X average of 3.45X) performance improvements over well-optimized conventional designs with small area overheads (upto 19.3%, 15.1% on average).
Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
ICCAD4
2002 Synthesis of custom processors based on extensible platforms
abstract
Efficiency and flexibility are critical, but often conflicting, design goals in embedded system design. The recent emergence of extensible processors promises a favorable tradeoff between efficiency and flexibility, while keeping design turnaround times short. Current extensible processor design flows automate several tedious tasks, but typically require designers to manually select the parts of the program that are to be implemented as custom instructions.In this work, we describe an automatic methodology to select custom instructions to augment an extensible processor, in order to maximize its efficiency for a given application program. We demonstrate that the number of custom instruction candidates grows rapidly with program size, leading to a large design space, and that the quality (speedup) of custom instructions varies significantly across this space, motivating the need for the proposed flow. Our methodology features cost functions to guide the custom instruction selection process, as well as static and dynamic pruning techniques to eliminate inferior parts of the design space from consideration. Further, we employ a two-stage process, wherein a limited number of promising instruction candidates are first selected, and then evaluated in more detail through cycle-accurate instruction set simulation and synthesis of the corresponding hardware, to identify the custom instruction combinations that result in the highest program speedup or maximize speedup under a given area constraint.We have evaluated the proposed techniques using a state-of-the-art extensible processor platform, in the context of a commercial design flow. Experiments with several benchmark programs indicate that custom processors synthesized using automatic custom instruction selection can result in large improvements in performance (upto 5.4X, average of 3.4X), energy (upto 4.5X, average of 3.2X), and energy-delay product (upto 24.2X, average of 12.6X), while speeding up the design process significantly.
Fei Sun 0002, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha
ICCAD4
2002 Interconnect-aware high-level synthesis for low power
abstract
Interconnects (wires, buffers, clock distribution networks, multiplexers and busses) consume a significant fraction of total circuit power. In this work, we demonstrate the importance of optimizing on-chip interconnects for power during high-level synthesis. We present a methodology to integrate interconnect power optimization into high-level synthesis. Our binding algorithm not only reduces power consumption in functional units and registers in the resultant register-transfer level (RTL) architecture, but also optimizes interconnects for power. We take physical design information into account for this purpose. To estimate interconnect power consumption accurately for deep sub-micron (DSM) technologies, wire coupling capacitance is taken into consideration. We observed that there is significant spurious (i.e., unnecessary) switching activity in the interconnects and propose techniques to reduce it. Compared to interconnect-unaware power-optimized circuits, our experimental results show that interconnect power can be reduced by 53.1% on an average, while reducing overall power by an average of 26.8% with 0.5% area overhead. Compared to area-optimized circuits, the interconnect power reduction is 72.9% and overall power reduction is 56.0% with 44.4% area overhead.
Lin Zhong 0001, Niraj K. Jha
ICCAD2
2002 Embedded Operating System Energy Analysis and Macro-Modeling
abstract
A large and increasing number of modern embedded systems are subject to tight power/energy constraints. It has been demonstrated that the operating system (OS) can have a significant impact on the energy efficiency of the embedded system. Hence, analysis of the energy effects of the OS is of great importance. Conventional approaches to energy analysis of the OS (and embedded software, in general) require the application software to be completely developed and integrated with the system software, and that either measurement on a hardware prototype or detailed simulation of the entire system be performed. Since this process requires significant design effort, unfortunately, it is typically too late in the design cycle to perform high-level or architectural optimizations on the embedded software, restricting the scope of power savings. Our work recognizes the need to provide embedded software designers with feedback about the effect of different OS services on energy consumption early in the design cycle. As a first step in that direction, this paper presents a systematic methodology to perform energy analysis and macro-modeling of an embedded OS. Our energy macro-models provide software architects and developers with an intuitive model for the OS energy effects, since they directly associate energy consumption with OS services and primitives that are visible to the application software. Our methodology consists of (i) an analysis stage, where we identify a set of energy components, called energy characteristics, which are useful to the designer in making OS-related design trade-offs, and (ii) a subsequent macromodeling stage, where we collect data for the identified energy components and automatically derive macro-models for them. We validate our methodology by deriving energy macro-models for two state-of-the-art embedded OS's, /spl mu/C/OS and Linux OS.
Tat Kee Tan, Anand Raghunathan, Niraj K. Jha
ICCD3
2002 Register Binding Based Power Management for High-level Synthesis of Control-Flow Intensive Behaviors
abstract
A circuit or circuit component that does not contain any spurious switching activity, i.e., activity that is not required by its specified functionality, is called perfectly power managed (PPM). We present a general sufficient condition for register binding to ensure that a given set of functional units is PPM. This condition not only applies to data-flow intensive (DFI) behaviors but also to control-flow intensive (CFI) behaviors. It leads to a straightforward power-managed (PM) register binding algorithm. The proposed algorithm is independent of the functional unit binding and scheduling algorithms. Hence, it can be easily incorporated into existing high-level synthesis systems. For the benchmarks we experimented with, an average 45.9% power reduction was achieved by our method at the cost of 7.7% average area overhead, compared to power-optimized register-transfer level (RTL) circuits which did not use PM register binding.
Lin Zhong 0001, Jiong Luo, Yunsi Fei, Niraj K. Jha
ICCD4
2002 Test synthesis of systems-on-a-chip
abstract
Embedded systems are increasingly synthesized today as systems-on-a-chip (SOCs), wherein existing functional blocks (also called cores) are used to implement different functions in the system specification. Emphasis during system synthesis is usually on optimizing one or more objectives such as price, area, performance and power. Testability enhancement of the SOC solution so obtained follows as a postprocessing step to enable the application of precomputed test sequences to each embedded core and observe its responses. Unfortunately, performing test modifications after the SOC design has been optimized for target design constraints does not preserve the optimality of the solution obtained. In this paper, the authors present the first system-level synthesis for testability framework that integrates testability considerations into the synthesis process. Their work incorporates finite-state automata (FSA)-based symbolic testability analysis within the framework of an existing multiobjective optimization-based system synthesis tool to provide a viable solution. The experimental results show that the use of FSA-based testability analysis facilitates low test overheads and test application times without sacrificing the test coverage of the embedded cores. System synthesis through the integrated framework for a number of examples indicates that efficient SOC architectures, which tradeoff different architectural features such as integrated circuit price, power consumption, and area under real-time constraints, can now be easily generated with low testability costs.
Srivaths Ravi 0001, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 High-level test compaction techniques
abstract
Available register-transfer level (RTL) test generation techniques do not make a concerted effort to reduce the test application time associated with the derived tests. Chip tester memory limitations, increasing tester costs, etc., make it imperative that the issue of generating compact tests at the RTL be addressed and consolidated with the known gains of high-level testing. In this paper, the authors provide a comprehensive framework for generating compact tests for an RTL circuit. They develop a series of techniques that exploit the inherent parallelism available in symbolic test(s) derived for RTL module(s). These techniques enable them to schedule testing of multiple modules in parallel as well as perform test pipelining. In addition, the authors also present design for testability (DFT) techniques for lowering test application time. Using a maximum bipartite matching formulation, they choose a low-overhead set of test enhancements that can achieve compact tests. The authors' techniques can seamlessly plug into any generic high-level test framework. Their experimental results in the context of one such framework indicate that the proposed methodology achieves an average reduction in test application time of 54.2% for the example circuits.
Srivaths Ravi 0001, Ganesh Lakshminarayana, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2002 High-level energy macromodeling of embedded software
abstract
Presents an efficient and accurate high level software energy estimation methodology using the concept of characterization-based macromodeling. In characterization-based macromodeling, a function or subroutine is characterized using an accurate lower level energy model of the target processor to construct a macromodel that relates the energy consumed in the function under consideration to various parameters that can be easily observed or calculated from a high-level programming language description. The constructed macromodels eliminate the need for significantly slower instruction-level interpretation or hardware simulation that is required in conventional approaches to software energy estimation. Two different approaches to macromodeling for embedded software offer distinct efficiency-accuracy characteristics: 1) complexity-based macromodeling, where the variables that determine the algorithmic complexity of the function under consideration are used as macromodeling parameters and 2) profiling-based macromodeling, where internal profiling statistics for the functions are used as the parameters in the energy macromodels.
Tat Kee Tan, Anand Raghunathan, Ganesh Lakshminarayana, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2002 Leakage power analysis and reduction during behavioral synthesis
abstract
This paper presents a high-level leakage power analysis and reduction algorithm. The algorithm uses device-level models for leakage to precharacterize a given register-transfer level module library. This is used to estimate the power consumption of a circuit due to leakage. The algorithm can also identify and extract the frequently idle modules in the datapath, which may be targeted for low-leakage optimization. Leakage optimization is based on the use of dual threshold voltage (V/sub T/) technology. The algorithm prioritizes modules giving a high-level synthesis system an indication of where most gains for leakage reduction may be found. We tested our algorithm using a number of benchmarks from various sources. We ran a series of experiments by integrating our algorithm into a low-power high-level synthesis system. In addition to reducing the power consumption due to switching activity, our algorithm provides the high-level synthesis system with the ability to detect and reduce leakage power consumption, hence, further reducing total power consumption. This is shown over a number of technology generations. The trend in these generations indicates that leakage becomes the dominant component of power at smaller feature size and lower supply voltages. Results show that using a dual-V/sub T/ library during high-level synthesis can reduce leakage power by an average of 58% for the different technology generations. Total power can be reduced by an average of 15.0%-45.0% for 0.18-0.07 /spl mu/m technologies, respectively. The contribution of leakage power to overall power consumption ranges from 22.6% to 56.2%. Our approach reduced these values to 11.7%-26.9%.
Kamal S. Khouri, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
2001 Battery-Aware Static Scheduling for Distributed Real-Time Embedded Systems
abstract
This paper addresses battery-aware static scheduling in battery-powered distributed real-time embedded systems. As suggested by previous work, reducing the discharge current level and shaping its distribution are essential for extending the battery lifespan. We propose two battery-aware static scheduling schemes. The first one optimizes the discharge power profile in order to maximize the utilization of the battery capacity. The second one targets distributed systems composed of voltage-scalable processing elements (PEs). It performs variable-voltage scheduling via efficient slack time re-allocation, which helps reduce the average discharge power consumption as well as flatten the discharge power profile. Both schemes guarantee the hard real-time constraints and precedence relationships in the real-time distributed embedded system specification. Based on previous work, we develop a battery lifespan evaluation metric which is aware of the shape of the discharge power profile. Our experimental results show that the battery lifespan can be increased by up to 29% by optimizing the discharge power file alone. Our variable-voltage scheme increases the battery lifespan by up to 76% over the non-voltage-scalable scheme and by up to 56% over the variable-voltage scheme without slack-time re-allocation.
Jiong Luo, Niraj K. Jha
DAC2
2001 High-level Software Energy Macro-modeling
abstract
This paper presents an efficient and accurate high-level software energy estimation methodology using the concept of characterization-based macro-modeling. In characterization-based macro-modeling, a function or sub-routine is characterized using an accurate lower-level energy model of the target processor, to construct a macro-model that relates the energy consumed in the function under consideration to various parameters that can be easily observed or calculated from a high-level programming language description. The constructed macro-models eliminate the need for significantly slower instruction-level interpretation or hardware simulation that is required in conventional approaches to software energy estimation.
Tat Kee Tan, Anand Raghunathan, Ganesh Lakshminarayana, Niraj K. Jha
DAC4
2001 Input Space Adaptive Design: A High-level Methodology for Energy and Performance Optimization
abstract
This paper presents a high-level design methodology, called input space adaptive design, and new design automation algorithms for optimizing energy consumption and performance. An input space adaptive design exploits the well-known fact that the quality of hardware circuits and software programs can be significantly optimized by employing algorithms and implementation architectures that adapt to the input statistics. We propose a methodology for such designs which includes identifying parts of the behavior to be optimized, selecting appropriate input sub-spaces, transforming the behavior, and verifying the equivalence of the original and optimized designs. Experimental results indicate that such designs can reduce energy consumption by up to 70.6% (average of 55.4%), and simultaneously improve performance by up to 85.1% (average of 58.1%), leading to a reduction in the energy-delay product by up to 95.6% (average of 80.7%), compared to well-optimized designs that do not employ such techniques.
Anand Raghunathan, Ganesh Lakshminarayana, Niraj K. Jha
DAC4
2001 Low Power System Scheduling and Synthesis
abstract
Many scheduling techniques have been presented recently which exploit dynamic voltage scaling (DVS) and dynamic power management (DPM) for both uniprocessors, and distributed systems, as well as both real-time and non-real-time systems. While such techniques are power-aware and aim at extending battery lifetimes for portable systems, they need to be augmented to make them battery-aware as well. This paper surveys such power-aware and battery-aware scheduling algorithms. In addition system synthesis algorithms for real-time systems-on-a-chip (SOCs), distributed and wireless client-server embedded systems, etc., have begun optimizing power consumption in addition to system price. The paper also surveys such algorithms as well, and points out some open problems.
Niraj K. Jha
ICCAD1
2001 High-Level Power Modeling of CPLDs and FPGAs
abstract
We present a high-level power modeling technique to estimate the power consumption of reconfigurable devices such as complex programmable logic devices (CPLDs) and field programmable gate arrays (FPGAs). For simplicity of reference, we simply refer to these devices as FPGAs. First, we capture the relationship between FPGA power dissipation and I/O signal statistics. We then use an adaptive regression method to model the FPGA power consumption. Such a high-level model can be used in the inner loop of a system-level synthesis tool to estimate the power consumed by different FPGA resources for different potential system-level synthesis solutions. It can also be used to verify the power budgets during embedded system design. With our high-level power model, the FPGA power consumption can be obtained very quickly. Experimental results indicate that the average relative error is only 3.1 % compared to low-level FPGA power simulation methods.
Niraj K. Jha
ICCD2
2001 Fast test generation for circuits with RTL and gate-level views
abstract
In this paper, we propose a simple two-pass strategy that couples register-transfer level (RTL) test generation with gate-level sequential test generation through fault lists. We motivate this approach by showing that faults found hard-to-test by gate-level sequential test generators are often easily testable at the RTL. Likewise, modules found symbolically untestable at the RTL have many of their faults testable at the gate level. Therefore, a two-pass strategy, which runs a fast RTL test generator followed by a gate-level sequential test generator on the remaining untested faults, can leverage off the strengths of each test generator. No modifications are necessary to the source code of either test generator to make this approach work. This makes it particularly attractive to industrial test flows, where the available gate-level test generator may be from a commercial vendor. This is in contrast to many hierarchical test generation techniques where there is significant interdependence between test generation at the register transfer and gate levels. For several benchmark circuits, we experimentally studied the performance of the two-pass approach using a symbolic RTL test generator, TAO, and efficient gate-level test generators, HITEC and SEST. Experimental results show that the proposed two-pass approach achieves a maximum speedup of 103X over a single-pass gate-level sequential test generator. The average speedup was 38X. No design for testability modifications were assumed for the circuits.
Srivaths Ravi 0001, Niraj K. Jha
ITC2
2001 Clock selection for performance optimization of control-flowintensive behaviors
abstract
This paper presents a clock selection algorithm for control-flow intensive behaviors that are characterized by the presence of conditionals and deeply nested loops. Unlike previous papers, which are primarily geared toward data-dominated behaviors, this algorithm examines the effects of branch probabilities and their interaction with allocation constraints. Using examples, we demonstrate, how changing branch probabilities and resource allocation can dramatically affect the optimal clock period, and hence, the performance of the schedule, and show that the interaction of these two factors must also be taken into account when searching for an optimal clock period. We then introduce the clock selection algorithm, which employs a fast critical-path analysis engine that allows it to evaluate what effect different clock periods, branch probabilities, and resource allocations may ultimately have on the performance of the behavior. When evaluating the critical path, we exploit the fact that our target behaviors exhibit locality of execution. We tested our algorithm using a number of benchmarks from various sources. A series of experiments demonstrates that our algorithm is quickly capable of selecting a small set of performance-enhancing clock periods, among which the optimal clock period typically lies. Another experiment demonstrates that the algorithm can adapt to varying resource constraints.
Kamal S. Khouri, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2001 Fault-diagnosis-based technique for establishing RTL and gate-levelcorrespondences
abstract
In this paper, we address an important problem associated with hierarchical design flows (termed the mapping problem): identifying correspondences between a signal in a high-level specification and a net in its lower level implementation. Conventional techniques use shared names to associate a signal with a net whenever possible. However, given that a synthesis flow may not preserve names, such a solution is not universally applicable. This work provides a robust framework for establishing register-transfer level (RTL) signal to gate-level net correspondences for a given design. Our technique exploits the observation that circuit diagnosis provides a convenient means for locating faults in a gate-level network. Since our problem requires locating gate-level nets corresponding to RTL signals, we formulate the mapping problem as a query whose solution is provided by a circuit diagnosis engine. Our experimental work with industrial designs for many mapping cases shows that our solution to the mapping problem is 1) fast and 2) precise in identifying the gate-level equivalents (the number of nets returned by our mapping engine for a query is typically one or two even for designs with tens of thousands of VHDL lines).
Srivaths Ravi 0001, Indradeep Ghosh, Vamsi Boppana, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2001 Testing of core-based systems-on-a-chip
abstract
Available techniques for testing of core-based systems-on-a-chip (SOCs) do not provide a systematic means for synthesizing low-overhead test architectures and compact test solutions. In this paper, we provide a comprehensive framework that generates low-overhead compact test solutions for SOCs. First, we develop a common ground for addressing issues such as core test requirements, core access, and testing hardware additions. For this purpose, we introduce finite-state automata (FSA) for modeling tests, transparency modes, and testing hardware behavior. In many cases, the tests repeat a basic set of test actions for different test data that can again be modeled using FSA. While earlier work can derive a single symbolic test for a module in a register-transfer level (RTL) circuit as a finite-state automaton, this work extends the methodology to the system level and additionally contributes a satisfiability-based solution to the problem of applying a sequence of tests phased in time. This problem is known to be a bottleneck in testability analysis not only at the system level, but also at the RTL. Experimental results show that the system-level average area overhead for making SOCs testable with our method is only 4.5%, while achieving an average test application time reduction of 80% over recent approaches. At the same time, it provides 100% test coverage of the precomputed test sets/sequences of the embedded cores.
Srivaths Ravi 0001, Ganesh Lakshminarayana, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2001 TAO: regular expression-based register-transfer level testability analysis and optimization
abstract
In this paper, we present testability analysis and optimization (TAO), a novel methodology for register-transfer level (RTL) testability analysis and optimization of RTL controller/data path circuits. Unlike existing high-level testing techniques that cater restrictively to certain classes of circuits or design styles, TAO exploits the algebra of regular expressions to provide a unified framework for handling a wide variety of circuits including application-specific integrated circuits (ASICs), application-specific programmable processors (ASPPs), application-specific instruction processors (ASIPs), digital signal processors (DSPs), and microprocessors. We also augment TAO with a design-for-test (DFT) framework that can provide a low-cost testability solution by examining the tradeoffs in choosing from a diverse array of testability modifications like partial scan or test multiplexer insertion in different parts of the circuit. Test generation is symbolic and, hence, independent of bit width. Experimental results on benchmark circuits show that TAO is very efficient, in addition to being comprehensive. The fault coverage obtained is above 99% in all cases. The average area and delay overheads for incorporating testability into the benchmarks are only 3.2% and 1.0%, respectively. The test generation time is two-to-four orders of magnitude smaller than that associated with gate-level sequential test generators, while the test application times are comparable.
Srivaths Ravi 0001, Ganesh Lakshminarayana, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
2000 Power analysis of embedded operating systems
abstract
The increasing complexity and software content of embedded systems has led to the frequent use of system software that helps applications access underlying hardware resources easily and efficiently. In this paper, we analyze the power consumption of real-time operating systems (RTOSs), which form an important component of the system software layer. Despite the widespread use of, and significant role played by, RTOSs in mobile and low-power embedded systems, little is known about their power consumption characteristics. This work presents the power profiles for a commercial RTOS, μC/OS, running several applications on an embedded system based on the Fujitsu SPARClite processor. Our work demonstrates that the RTOS can consume a significant fraction of the system power and, in addition, impact the power consumed by other software components. We illustrate the ways in which application software can be designed to use the RTOS in a power-efficient manner. We believe that this work is a first step towards establishing a systematic approach to RTOS power modeling and optimization.
Robert P. Dick, Ganesh Lakshminarayana, Anand Raghunathan, Niraj K. Jha
DAC4
2000 Power-Conscious Joint Scheduling of Periodic Task Graphs and Aperiodic Tasks in Distributed Real-Time Embedded Systems
abstract
In this paper, we present a power-conscious algorithm for jointly scheduling multi-rate periodic task graphs and aperiodic tasks in distributed real-time embedded systems. While the periodic task graphs have hard deadlines, the aperiodic tasks can have either hard or soft deadlines. Periodic task graphs are first scheduled statically. Slots are created in this static schedule to accommodate hard aperiodic tasks. Soft aperiodic tasks are scheduled dynamically with an on-line scheduler. Flexibility is introduced into the static schedule and optimized to allow the on-line scheduler to make dynamic modifications to the static schedule. This helps minimize the response times of soft aperiodic tasks through both resource reclaiming and slack stealing. Of course, the validity of the static schedule is maintained. The on-line scheduler also employs dynamic voltage scaling and power management to obtain a power-efficient schedule. Experimental results show that the flexibility introduced into the static schedule helps improve the response times of soft aperiodic tasks by up to 43%. Dynamic voltage scaling and power management reduce power by up to 68%. The scheme in which the static schedule is allowed to be flexible achieves up to 3.2% more power saving compared to the scheme in which no flexibility is allowed, when both schemes are power-conscious. Our work gives an average architecture price saving of 30% over a previous approach for embedded system architectures synthesized with execution slots for hard aperiodic tasks present.
Jiong Luo, Niraj K. Jha
ICCAD2
2000 Leakage Power Analysis and Reduction during Behavioral Synthesis
abstract
This paper presents a high-level leakage power analysis and reduction algorithm. The algorithm uses device-level models for leakage to pre-characterize a given register-transfer level module library. This is used to estimate the power consumption of a circuit due to leakage. The algorithm can also identify and extract the frequently idle modules in the datapath, which may be targeted for low-leakage optimization. Leakage optimization is based on the use of dual threshold voltage (V/sub T/) technology. The algorithm prioritizes modules giving a high level synthesis (HLS) system an indication of where most gains for leakage reduction may be found. Results show that using a dual-V/sub T/ library during HLS can reduce leakage power by an average of 59% for the different technology generations. Total power can be reduced by an average of 18.8% to 45.4% for 0.18 /spl mu/m to 0.07 /spl mu/m technologies, respectively, compared to register-transfer level (RTL) circuits optimized for switching power only. The contribution of leakage power to overall power consumption of switching power optimized RTL circuits ranges from 23.5% to 54.1%. Our approach reduced these values to 11.4% to 25.9%.
Kamal S. Khouri, Niraj K. Jha
ICCD2
2000 A Technique for Identifying RTL and Gate-Level Correspondences
abstract
In this paper, we consider the mapping problem of identifying correspondences between a signal in a high-level specification and a net in a lower-level implementation for a given design. Conventional techniques use shared names to associate a signal with a net whenever possible. However, given that a synthesis flow may not preserve names, solutions to the above problem eventually take recourse to expensive alternatives such as formal verification. This paper provides a robust framework for identifying RTL signal to gate-level net correspondences for a given design. Our technique exploits the observation that circuit diagnosis provides a convenient means for locating faults in a gate-level network. Since our problem requires locating gate-level nets corresponding to RTL signals, we formulate the mapping problem as a query whose solution is provided by a circuit diagnosis engine. Our experimental work with industrial designs for many mapping cases shows that our solution to the mapping problem is (i) fast, and (ii) precise in identifying the gate-level equivalents (the number of nets returned by our mapping engine for a query is typically or 2 even for designs with tens of thousands of VHDL lines).
Srivaths Ravi 0001, Niraj K. Jha, Indradeep Ghosh, Vamsi Boppana
ICCD2
2000 : Reducing test application time in high-level test generation
abstract
Available register-transfer level (RTL) test generation techniques do not make a concerted effort to reduce the test application time associated with the derived tests. Chip tester memory limitations, increasing tester costs, etc., make it imperative that the issue of generating compact tests at the RTL be addressed and consolidated with the known gains of high-level testing. In this paper, we provide a comprehensive framework for generating compact tests for an RTL circuit. We develop a series of techniques that exploit the inherent parallelism available in symbolic test(s) derived for RTL module(s). These techniques enable us to schedule testing of multiple modules in parallel as well as perform test pipelining. In addition, we also present design for testability (DFT) techniques for lowering test application time. Using a maximum bipartite matching formulation, we choose a low-overhead set of test enhancements that can achieve compact tests. Our techniques can seamlessly plug into any generic high-level test framework. Our experimental results in the context of one such framework indicate that the proposed methodology achieves an average reduction in test application time of 61.1% for the example circuits.
Srivaths Ravi 0001, Ganesh Lakshminarayana, Niraj K. Jha
ITC3
2000 Behavioral Synthesis of Fault Secure Controller/Datapaths Based on Aliasing Probability Analysis
abstract
This paper addresses the problem of synthesizing fault-secure controller/data path circuits from behavioral specifications. These circuits are guaranteed to either produce the correct output, or to flag an error. We use an iterative improvement-based behavioral synthesis framework that performs functional unit selection, clock selection, scheduling, and resource sharing with the aim of minimizing the area of the synthesized circuit, while allowing multicycling, chaining, and functional unit pipelining. We present a dynamic comparison selection algorithm that can be used during behavioral synthesis to determine which intermediate results in the computation need to be secured in order to enable maximal resource sharing. Previous work on synthesizing fault-secure data paths has focused on ensuring that aliasing (a condition when the circuit produces an incorrect output and does not flag an error) cannot occur in any part of the design. We demonstrate that such an approach can lead to unnecessarily large overheads. In order to alleviate the overheads incurred for fault security, our behavioral synthesis framework uses ALiasing Probability analysiS (ALPS) in order to identify resource sharing configurations that reduce area while introducing a very low probability of aliasing (of the order of 10/sup -10/ for a bit-width of 32) in the resultant data path. Experimental results performed for several behavioral descriptions demonstrate that our techniques synthesize more compact circuits than techniques available in the literature, e.g., double modular redundancy or zero-aliasing techniques.
Ganesh Lakshminarayana, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Computers3
2000 A fast and low-cost testing technique for core-based system-chips
abstract
This paper proposes a new methodology for testing a core-based system chip, targeting the simultaneous reduction of test area overhead and test application time. At the core level, testability and transparency can be achieved by the core provider by reusing existing logic inside the core, providing different versions of the core having different area overheads and transparency latencies. The technique analyzes the topology of the system-chip to select the core versions that best meet the user's desired test area overhead and test application time objectives. Application of the method to example system-chips demonstrates the ability to design highly testable system-chips with minimized test area overhead, minimized test application time, or a desired tradeoff between the two. Significant reduction in area overhead and test application time compared to existing system chip testing techniques is also demonstrated.
Indradeep Ghosh, Sujit Dey, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2000 A BIST scheme for RTL circuits based on symbolic testabilityanalysis
abstract
This paper introduces a novel scheme for testing register-transfer level (RTL) controller/data paths using built-in self-test (BIST). The scheme uses the controller netlist and the data path of a circuit to extract a test control/data flow (TCDF) graph. This TCDF is used to derive a set of symbolic justification and propagation paths (known as the test environment) to test some of the operations and variables present in it. If it becomes difficult to generate such test environments with the derived TCDF's, a few test multiplexers are added at suitable points in the circuit to increase its controllability and observability. The test environment of an operation (variable) guarantees the existence of a path from the primary inputs of the circuit to the inputs of the module (register) to which the operation (variable) is mapped, and a path from the output of the module (register) to a primary output of the circuit. Since the search for a test environment is done symbolically, it is very fast and needs to be done only once for each module or register,in the circuit. This test environment can then be used to exercise a module or register in the circuit with pseudorandom pattern generators which are placed only at the primary inputs of the circuit. The test responses can he analyzed with signature analyzers which are only placed at the primary outputs of the circuit. Unlike many RTL BIST schemes, an increase in the data path bit-width does not adversely impact the complexity of our testability analysis scheme since the analysis is symbolic. Every module in the module library is made random-pattern testable, whenever possible, using gate-level testability insertion techniques. This is a one-time cost. Finally, a BIST controller is synthesized to provide the necessary control signals to form the different test environments during testing, and a BIST architecture is superimposed on the circuit. Experimental results on a number of industrial and university benchmarks show that high fault coverage (>99%) can be obtained with our scheme. The average area overhead of the scheme is 6.9% which is much lower than many existing logic-level BIST schemes. The average delay overhead is only 2.5%. The test application time to achieve the high fault coverage for the whole circuit is also quite low.
Indradeep Ghosh, Niraj K. Jha, Sudipta Bhawmik
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2000 Incorporating speculative execution into scheduling ofcontrol-flow-intensive designs
abstract
Speculative execution refers to the execution of parts of a computation before the execution of the conditional operations that decide whether they need to be executed. It has been shown to be a promising technique for eliminating performance bottlenecks imposed by control flow in hardware and software implementations alike. In this paper, we present techniques to incorporate speculative execution in a fine-grained manner into scheduling of control-flow-intensive behavioral descriptions. We demonstrate that failing to take into account information such as resource constraints and branch probabilities can lead to significantly suboptimal performance. We also demonstrate that it may be necessary to speculate simultaneously along multiple paths, subject to resource constraints, in order to minimize the delay overheads incurred when prediction errors occur. Experimental results on several benchmarks show that our speculative scheduling algorithm can result in significant (up to seven-fold) improvements in performance (measured in terms of the average number of clock cycles) as compared to scheduling without speculative execution. Also, the best and worst case execution times for the speculatively performed schedules are the same as or better than the corresponding values for the schedules obtained without speculative execution.
Ganesh Lakshminarayana, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2000 TAO-BIST: A framework for testability analysis and optimization forbuilt-in self-test of RTL circuits
abstract
In this paper, we present TAO-BIST, a framework for testing register-transfer level (RTL) controller-datapath circuits using built-in self-test (BIST). Conventional BIST techniques at the RTL generally introduce more testability hardware than is necessary thereby causing unnecessary area, delay and power overheads. They have typically been applied to only application-specific integrated circuits (ASICs). TAO-BIST adopts a three-phased approach to provide an efficient BIST framework at the RTL. In the first phase, we identify and add an initial set of test enhancements to the given circuit. In the second phase, we use regular-expression based high-level symbolic testability analysis of a BIST model of the circuit to completely encapsulate justification/propagation information for the modules under test. The regular expressions so obtained are then used to construct a Boolean function in the final phase for determining a test enhancement solution that meets delay constraints with minimal area overheads. Our method is applicable to a wide spectrum of circuits including ASICs, application-specific programmable processors (ASPPs), application-specific instruction processors (ASIPs), digital signal processors (DSPs), and microprocessors. Experimental results on a number of benchmark circuits show that high fault coverage (>99%) can be obtained with our scheme. The average area and delay overheads due to TAO-BIST are only 6.0% and 1.5%, respectively, for a bit-width of 16. These overheads decrease further with an increase in bit-width. The test application time to achieve the high fault coverage for the whole controller-datapath circuit is also quite low.
Srivaths Ravi 0001, Ganesh Lakshminarayana, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1999 Common-Case Computation: A High-Level Technique for Power and Performance Optimization
abstract
This paper presents a design methodology, called common-case computation (CCC), and new design automation algorithms for optimizing power consumption or performance. The proposed techniques are applicable in conjunction with any high-level design methodology where a structural register-transfer level (RTL) description and its corresponding scheduled behavioral (cycle-accurate functional RTL) description are available. It is a well-known fact that in behavioral descriptions of hardware (also in software), a small set of computations (CCCs) often accounts for most of the computational complexity. However, in hardware implementations (structural RTL or lower level), CCCs and the remaining computations a typically treated alike. This paper shows that identifying and exploiting CCCs during the design process can lead to implementations that are much more efficient in terms of power consumption or performance. We propose a CCC-based high-level design methodology with the following steps: extraction of common-case behaviors and execution conditions from the scheduled description, simplification of the common-case behaviors in a stand-alone manner, synthesis of common-case detection and execution circuits from the common-case behaviors, and composing the original design with the common-case circuits, resulting in a CCC-optimized design. We demonstrate that CCC-optimized designs reduce power consumption by up to 91.5%, or improve performance by up to 76.6% compared to designs derived without special regard for CCCs.
Ganesh Lakshminarayana, Anand Raghunathan, Kamal S. Khouri, Niraj K. Jha, Sujit Dey
DAC4
1999 MOCSYN: Multiobjective Core-Based Single-Chip System Synthesis
abstract
In this paper we present a system synthesis algorithm, called MOCSYN, which partitions and schedules embedded system specifications to intellectual property cores in an integrated circuit. Given a system specification consisting of multiple periodic task graphs as well as a database of core and integrated circuit characteristics, MOCSYN synthesizes real-time heterogeneous single-chip hardware software architectures using an adaptive multiobjective genetic algorithm that is designed to escape local minima. The use of multiobjective optimization allows a single system synthesis run to produce multiple designs which trade off different architectural features. Integrated circuit price, power consumption, and area are optimized under hard real-time constraints. MOCSYN differs from previous work by considering problems unique to single-chip systems. It solves the problem of providing clock signals to cores composing a system-on-a-chip. It produces a bus structure which balances ease of layout with the reduction of bus contention. In addition, it carries out floorplan block placement within its inner loop allowing accurate estimation of global communication delays and power consumption.
Robert P. Dick, Niraj K. Jha
DATE2
1999 Memory binding for performance optimization of control-flow intensive behaviors
abstract
The paper presents a memory binding algorithm for behaviors that are characterized by the presence of conditionals and deeply-nested loops that access memory extensively through arrays. Unlike previous works, this algorithm examines the effects of branch probabilities and allocation constraints. First, we demonstrate through examples, the importance of incorporating branch probabilities and allocation constraint information when searching for a performance-efficient memory binding. We also show the interdependence of these two factors and how varying one without considering the other may greatly affect the performance of the behavior. Second, we introduce a memory binding algorithm that has the ability to examine numerous bindings by employing an efficient performance estimation procedure. The estimation procedure exploits locality of execution, which is an inherent characteristic of target behaviors. This enables the performance estimation technique to look at the global impact of the different bindings, given the allocation constraints. We tested our algorithm using a number of benchmarks from the parallel computing domain. A series of experiments demonstrates the algorithm's ability to produce bindings that optimize performance, meet memory allocation constraints, and adapt to different resource constraints and branch probabilities. Results show that the algorithm requires 37% fewer memories with a performance loss of only 0.3% when compared to a parallel memory architecture. When compared to the best of a series of random memory bindings, the algorithm improves schedule performance by 21%.
Kamal S. Khouri, Ganesh Lakshminarayana, Niraj K. Jha
ICCAD3
1999 A framework for testing core-based systems-on-a-chip
abstract
Available techniques for testing core-based systems-on-a-chip (SOCs) do not provide a systematic means for synthesising low-overhead test architectures and compact test solutions. In this paper, we provide a comprehensive framework that generates low-overhead compact test solutions for SOCs. First, we develop a common ground for addressing issues such as core test requirements, core access and test hardware additions. For this purpose, we introduce finite-state automata for modeling tests, transparency modes and test hardware behavior. In many cases, the tests repeat a basic set of test actions for different test data which can again be modeled using finite-state automata. While earlier work can derive a single symbolic test for a module in a register-transfer level (RTL) circuit as a finite-state automaton, this work extends the methodology to the system level, and, additionally contributes a satisfiability-based solution to the problem of applying a sequence of tests phased in time. This problem is known to be a bottleneck in testability analysis not only at the system level, but also at the RTL. Experimental results show that the system-level average area overhead for making SOCs testable with our method is only 4.4%, while achieving an average test application time reduction of 78.5% over recent approaches. At the same time, it provides 100% test coverage of the precomputed test sets/sequences of the embedded cores.
Srivaths Ravi 0001, Ganesh Lakshminarayana, Niraj K. Jha
ICCAD3
1999 TAO-BIST: A Framework for Testability Analysis and Optimizationb of RTL Circuits for BIST
abstract
In this paper, we present TAO-BIST, a framework for testing register-transfer level (RTL) controller-datapath circuits using built-in self-test (BIST). Conventional BIST techniques at the RTL generally introduce more testability hardware than is necessary, thereby causing unnecessary area, delay and power overheads. They have typically been applied to only application-specific integrated circuits (ASICs). TAO-BIST adopts a three-phased approach to provide an efficient BIST framework at the RTL. In the first phase, we identify and add an initial set of test enhancements to the given circuit. In the second phase, we use regular-expression based high-level symbolic testability analysis of a BIST model of the circuit to completely encapsulate justification/propagation information for the modules under test. The regular expressions so obtained are then used to construct a Boolean function in the final phase for determining a test enhancement solution that meets delay constraints with minimal area overheads. Our method is applicable to a wide spectrum of circuits including ASICs, application-specific programmable processors (ASPPs), application-specific instruction processors (ASIPs), digital signal processors (DSPs) and microprocessors. Experimental results on a number of benchmark circuits show that high fault coverage (>99%) can be obtained with our scheme. The average area and delay overheads due to TAO-BIST are only 6.0%, and 1.5%, respectively. The test application time to achieve the high fault coverage for the whole controller-datapath circuit is also quite low.
Srivaths Ravi 0001, Ganesh Lakshminarayana, Niraj K. Jha
VTS3
1999 COFTA: Hardware-Software Co-Synthesis of Heterogeneous Distributed Embedded Systems
abstract
Embedded systems employed in critical applications demand high reliability and availability in addition to high performance. Hardware-software co-synthesis of an embedded system is the process of partitioning, mapping, and scheduling its specification into hardware and software modules to meet performance, cost, reliability, and availability goals. In this paper, we address the problem of hardware-software co-synthesis of fault-tolerant real-time heterogeneous distributed embedded systems. Fault detection capability is imparted to the embedded system by adding assertion and duplicate-and-compare tasks to the task graph specification prior to co-synthesis. The dependability (reliability and availability) of the architecture is evaluated during co-synthesis. Our algorithm, called COFTA (Co-synthesis Of Fault-Tolerant Architectures), allows the user to specify multiple types of assertions for each task. It uses the assertion or combination of assertions which achieves the required fault coverage without incurring too much overhead. We propose new methods to: 1) Perform fault tolerance based task clustering, which determines the best placement of assertion and duplicate-and-compare tasks, 2) Derive the best error recovery topology using a small number of extra processing elements, 3) Exploit multidimensional assertions, and 4) Share assertions to reduce the fault tolerance overhead. Our algorithm can tackle multirate systems commonly found in multimedia applications. Application of the proposed algorithm to a large number of real-life telecom transport system examples (the largest example consisting of 2,172 tasks) shows its efficacy. For fault secure architectures, which just have fault detection capabilities, COFTA is able to achieve up to 48.8 percent and 25.6 percent savings in embedded system cost over architectures employing duplication and task-based fault tolerance techniques, respectively. The average cost overhead of COFTA fault-secure architectures over simplex architectures is only 7.3 percent. In case of fault-tolerant architectures, which cannot only detect but also tolerate faults, COFTA is able to achieve up to 63.1 percent and 23.8 percent savings in embedded system cost over architectures employing triple-modular redundancy, and task-based fault tolerance techniques, respectively. The average cost overhead of COFTA fault-tolerant architectures over simplex architectures is only 55.4 percent.
Bharat P. Dave, Niraj K. Jha
IEEE Trans. Computers2
1999 Controller-based power management for control-flow intensive designs
abstract
This paper presents a low-overhead controller-based power management technique that redesigns control logic to reconfigure the existing data path components under idle conditions so as to minimize unnecessary activity. Controller-based power management exploits the fact that though the control signals in a register-transfer level implementation are fully specified, they can be respecified under certain states/conditions when the data path components that they control need not be active. We demonstrate that controller-based power management is often better-suited to control-flow intensive designs than comparable conventional power management techniques such as operand isolation. We present an algorithm to perform power management through controller redesign that consists of constructing an activity graph for each data path component, identifying conditions under which the component need not be active, and relabeling the activity graph resulting in redesign of the corresponding control expressions. We provide a comprehensive analysis of the potential side effects of controller-based power management on circuit delay, glitching activity at control and data path signals, and formation of false combinational cycles. Our algorithm avoids the above negative effects of controller-based power management to maximize power savings and minimize overheads. We present experimental results which demonstrate that (1) controller-based power management results in large power savings at minimal overheads for control-flow intensive designs, which pose several challenges to conventional power management techniques and (2) it is important to consider the various potential negative effects while performing controller-based power management in order to obtain maximal power savings.
Sujit Dey, Anand Raghunathan, Niraj K. Jha, Kazutoshi Wakabayashi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1999 Corrections to "mogac: a multiobjective genetic algorithm for hardware-software cosynthesis of distributed embedded systems"
abstract
In the above-named article [ibid., vol. 16, pp. 920-935, Oct. 1998] there is an error in the experimental results reported for the Hou 3&4 (clustered) benchmark, due to a typographical error in its task collection input file. The last line of the The Hou 3&4 (clustered) row of Table IV on page 932 of the original article should be replaced with Table I of this short paper. For this example, average price increases when the effort level is changed from three to four. In general, solution quality increases, i.e., the price decreases, with increasing optimizer effort. This trend is not expressed in this case as a result of the limited number of samples (30). When run with 500 samples, price strictly decreases with increasing effort. The Hou 3&4 (clustered) row of Table VI on page 933 of the above paper1 should be replaced with Table II of this short paper.
Robert P. Dick, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1999 A low overhead design for testability and test generation technique for core-based systems-on-a-chip
abstract
In a fundamental paradigm shift in system design, entire systems are being built on a single chip, using multiple embedded cores. Though the newest system design methodology has several advantages in terms of time-to-market and system cost, testing such core-based systems is difficult, mainly due to the problem of justifying test sequences at the inputs of a core embedded deep in the circuit and propagating test responses from the core outputs. In this paper, we first present a design for testability technique for testing such core-based systems. In this scheme, untestable cores are first made testable using hierarchical testability analysis techniques. If necessary, additional testability hardware is added to the cores to make them transparent so that they can propagate test data without information loss. This testability and transparency technique is currently applicable to cores of the following types: application-specific integrated circuits, application-specific programmable processors, and application-specific instruction processors. Other core types can be made testable and transparent using traditional techniques. The testable and transparent cores can then he integrated together with some system-level testability hardware to ensure justification of precomputed test sequences of each core from system primary inputs to the core inputs and propagation of test responses from core outputs to system primary outputs. Justification and propagation of test sequences are done at the system level by extending and suitably modifying the symbolic hierarchical testability analysis method that has been successfully applied to register-transfer level circuits. Since the testability analysis method is symbolic, the system test generation method is independent of the bit-width of the cores. The system-level test set is obtained as a byproduct of the testability analysis and insertion method without further search. The test methodology was applied to six example systems. Besides the proposed test method, the two methods that are currently used in the industry were also evaluated: (1) FScan-BScan, where each core is full-scanned, and system test is performed using boundary scan and (2) FScan-TBus, where each core is full-scanned, and system test is performed using a test bus. The experiments show that the proposed scheme has significantly lower area overhead, delay overhead, and test application time compared to FScan-BScan and FScan-TBus, without any compromise in the system fault coverage.
Indradeep Ghosh, Niraj K. Jha, Sujit Dey
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1999 Hierarchical test generation and design for testability methods for ASPPs and ASIPs
abstract
In this paper, we present design for testability (DFT) and hierarchical test generation techniques for facilitating the testing of application-specific programmable processors (ASPPs) and application-specific instruction processors (ASIPs). The method utilizes the register-transfer level (RTL) circuit description of an ASPP or ASIP to come up with a set of test microcode patterns which can be written into the instruction read-only memory (ROM) of the processor. These lines of microcode dictate a new control/data flow in the circuit and can be used to test modules which are not easily testable. The new control/data flow is used to justify precomputed test sets of a module from the system primary inputs to the module inputs and propagate output responses from the module output to the system primary outputs. The testability analysis, which is based on the relevant control/data flow extracted from the RTL circuit, is symbolic. Thus, it is independent of the bit-width of the data path and is extremely fast. The test microcode patterns are a by-product of this analysis. If the derived test microcode cannot test all untested modules in the circuit, then test multiplexers are added (usually to the off-critical paths of the data path) to test these modules. This is done to guarantee the testability of all modules in the circuit. If the control microcode memory of the processor is erasable, then the test microcode lines can be erased once the testing of the chip is over. In that case, the DFT scheme has very little overhead (typically less than 1%). Otherwise, the test microcode lines remain as an overhead in the control memory. The method requires the addition of only one external test pin. Application of this technique to several examples has resulted in a very high fault coverage (above 99.6%) for all of them. The test generation time is about three orders of magnitude smaller compared to an efficient gate-level sequential test generator. The average area overhead (without assuming an erasable ROM) is 3.1% while the delay overheads are negligible. This method does not require any scan in the controller or data path. It is also amenable to at-speed testing.
Indradeep Ghosh, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1999 High-level synthesis of low-power control-flow intensive circuits
abstract
In this paper, we present a comprehensive high-level synthesis system that is geared toward reducing power consumption in control-flow intensive as well as data-dominated circuits. An iterative improvement framework allows the system to search the design space by examining the interaction between the different high-level synthesis tasks. In addition to incorporating traditional high-level synthesis tasks such as scheduling, module selection and resource sharing, we introduce a new optimization that performs power-conscious structuring of multiplexer networks, which are predominant in control-flow intensive circuits. The scheduler employed is capable of loop optimizations within and across loop boundaries. We also introduce a fast power estimation technique, based on switching activity matrices, to drive the synthesis process. Experimental results for a number of control-flow intensive and data-dominated benchmarks demonstrate power reduction of up to 62% (58%) when compared to V/sub dd/-scaled area-optimized (delay-optimized) designs. The area overheads over area-optimized designs are less than 39%, whereas the area savings over delay-optimized designs are up to 40%.
Kamal S. Khouri, Ganesh Lakshminarayana, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1999 High-level synthesis of power-optimized and area-optimized circuits from hierarchical data-flow intensive behaviors
abstract
We present a technique for synthesizing power- as well as area-optimized circuits from hierarchical data flow graphs under throughput constraints. We allow for the use of complex register-transfer level (RTL) modules, such as fast Fourier transforms (FFT's) and filters, as building blocks for the RTL circuit, in addition to simple RTL modules such as adders and multipliers. Unlike past techniques in the area, we also customize the complex RTL modules to match the environment in which they find themselves. We present a fast and efficient algorithm for mapping multiple behaviors onto the same RTL module during the course of synthesis, thus allowing our synthesis system to explore previously unexplored regions of the design space. These techniques are at the core of an iterative improvement based approach which can accept temporary degradation in solution quality in its quest for a globally optimal solution. The moves in our iterative improvement procedure explore optimizations along different dimensions such as functional unit selection, resource allocation, resource sharing, resource splitting, and selection and resynthesis of complex RTL modules. These interrelated optimizations are dynamically traded off with each other during the course of synthesis, thus exploiting the benefits that arise from their interaction. The synthesis framework also tackles other related high-level synthesis tasks such as scheduling, clock selection, and V/sub dd/ selection. Experimental results demonstrate that our algorithm produces circuits whose area and power consumption are comparable to or better than those produced using flattened synthesis, within much shorter CPU times.
Ganesh Lakshminarayana, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1999 FACT: a framework for applying throughput and power optimizing transformations to control-flow-intensive behavioral descriptions
abstract
In this paper, we present an algorithm for the application of a general class of transformations to control-flow intensive behavioral descriptions. Our algorithm is based on the observation that incorporation of scheduling information can help guide the selection and application of candidate transformations, and significantly enhance the quality of the synthesized solution. The efficacy of the selected throughput and power optimizing transformations is enhanced by the ability of our algorithm to transcend basic blocks in the behavioral description. This ability is imparted to our algorithm by a general technique we have devised. Our system currently supports associativity, commutativity, distributility, constant propagation, code motion, and loop unrolling. It is integrated with a scheduler which performs implicit loop unrolling and functional pipelining, and has the ability to parallelize the execution of independent iterative constructs, whose bodies can share resources. Other transformations can easily be incorporated within the framework. We demonstrate the efficacy of our algorithm by applying it to several commonly available benchmarks. Upon synthesis, behaviors transformed by the application of our algorithm showed, on an average, a 2.5-fold improvement in throughput over an existing transformation algorithm, and a 57.6% improvement in power over designs produced without the benefit of our algorithm.
Ganesh Lakshminarayana, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1999 Wavesched: a novel scheduling technique for control-flow intensive designs
abstract
In this paper, we present a novel scheduling algorithm targeted toward minimizing the average execution time of control-flow intensive behavioral descriptions. Our algorithm uses a control/data flow graph model, which preserves the parallelism inherent in the application. It explores previously unexplored regions of the solution space by its ability to overlap the schedules of independent iterative constructs, whose bodies share resources. It also incorporates well known optimization techniques like loop unrolling in a natural fashion. This is made possible by a general loop-handling technique, which we have devised. Application of the algorithm to several common benchmarks demonstrates up to 4.8-fold improvement in expected schedule length over existing scheduling algorithms, without paying a price in terms of the best and worst case schedule lengths required to execute the behavioral description (in fact, frequently, the best/worst case schedule lengths are also better for our algorithm).
Ganesh Lakshminarayana, Kamal S. Khouri, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1999 Register transfer level power optimization with emphasis on glitch analysis and reduction
abstract
We present design-for-low-power techniques for register-transfer level (RTL) controller/data path circuits. We analyze the generation and propagation of glitches in both the control and data path parts of the circuit. In data-flow intensive designs, glitching power is primarily due to the chaining of arithmetic functional units. In control-flow intensive designs, on the other hand, multiplexer networks and registers dominate the total circuit power consumption, and the control logic can generate a significant amount of glitches at its outputs, which in turn propagate through the data path to account for a large portion of the glitching power in the entire circuit. Our analysis also highlights the relationship between the propagation of glitches from control signals and the bit-level correlation between data signals. Based on the analysis, we develop techniques that attempt to reduce glitching power consumption by minimizing propagation of glitches in the RTL circuit. Our techniques include restructuring multiplexer networks (to enhance data correlations and eliminate glitchy control signals), clocking control signals, and inserting selective rising/falling delays, in order to kill the propagation of glitches from control as well as data signals. In addition, we present a procedure to automatically perform the well-known power-reduction technique of clock gating through an efficient structural analysis of the RTL circuit, while avoiding the introduction of glitches on the clock signals. Application of the proposed power optimization techniques to several RTL circuits shows significant power savings, with negligible area and delay overheads.
Anand Raghunathan, Sujit Dey, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1999 Safety and Reliability Driven Task Allocation in Distributed Systems
abstract
Distributed computer systems are increasingly being employed for critical applications, such as aircraft control, industrial process control, and banking systems. Maximizing performance has been the conventional objective in the allocation of tasks for such systems. Inherently, distributed systems are more complex than centralized systems. The added complexity could increase the potential for system failures. Some work has been done in the past in allocating tasks to distributed systems, considering reliability as the objective function to be maximized. Reliability is defined to be the probability that none of the system components falls while processing. This, however, does not give any guarantees as to the behavior of the system when a failure occurs. A failure, not detected immediately, could lead to a catastrophe. Such systems are unsafe. In this paper, we describe a method to determine an allocation that introduces safety into a heterogeneous distributed system and at the same time attempts to maximize its reliability. First, we devise a new heuristic, based on the concept of clustering, to allocate tasks for maximizing reliability. We show that for task graphs with precedence constraints, our heuristic performs better than previously proposed heuristics. Next, by applying the concept of task-based fault tolerance, which we have previously proposed, we add extra assertion tasks to the system to make it safe. We present a new heuristic that does this in such a way that the decrease in reliability for the added safety is minimized. For the purpose of allocating the extra tasks, this heuristic performs as well as previously known methods and runs an order of magnitude faster. We present a number of simulation results to prove the efficacy of our scheme.
Santhanam Srinivasan, Niraj K. Jha
IEEE Trans. Parallel Distributed Syst.2
1999 COSYN: Hardware-software co-synthesis of heterogeneous distributed embedded systems
abstract
Hardware-software co-synthesis starts with an embedded-system specification and results in an architecture consisting of hardware and software modules to meet performance, power, and cost goals. Embedded systems are generally specified in terms of a set of acyclic task graphs. In this paper, we present a co-synthesis algorithm COSYN, which starts with periodic task graphs with real-time constraints and produces a low-cost heterogeneous distributed embedded-system architecture meeting these constraints. It supports both concurrent and sequential modes of communication and computation. It employs a combination of preemptive and nonpreemptive static scheduling. It allows task graphs in which different tasks have different deadlines. It introduces the concept of an association array to tackle the problem of multirate systems. It uses a new task-clustering technique, which takes the changing nature of the critical path in the task graph into account. It supports pipelining of task graphs and a mix of various technologies to meet embedded-system constraints and minimize power dissipation. In general, embedded-system tasks are reused across multiple functions. COSYN uses the concept of architectural hints and reuse to exploit this fact. Finally, if desired, it also optimizes the architecture for power consumption. COSYN produces optimal results for the examples from the literature while providing several orders of magnitude advantage in central processing unit time over an existing optimal algorithm. The efficacy of COSYN and its low-power extension COSYN-LP is also established through their application to very large task graphs (with over 1000 tasks).
Bharat P. Dave, Ganesh Lakshminarayana, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.3
1999 Power management in high-level synthesis
abstract
In this paper, we present a power-management methodology targeted toward high-level synthesis of data-dominated behavioral descriptions. It is founded on the observation that variable assignment can significantly affect power-management opportunities in the synthesized architecture, i.e., variable assignment determines whether or not spurious operations get executed by functional units in the architecture. We introduce perfectly power managed architectures, whose functional units do not execute any spurious operations. We present a variable assignment technique which, when used in high-level synthesis, produces architectures which are perfectly power-managed. Unlike many previously proposed power-management techniques, our method does not add latches or any other circuitry in front of functional units or registers and is, therefore, free of the attendant performance penalty. Experimental results indicate savings of up to 52.5% (average 23.0%) in power consumption over already power-optimized architectures. The area overheads due to our technique are also low and averaged 2.5% for our examples.
Ganesh Lakshminarayana, Anand Raghunathan, Niraj K. Jha, Sujit Dey
IEEE Trans. Very Large Scale Integr. Syst.3
1998 A Fast and Low Cost Testing Technique for Core-Based System-on-Chip
abstract
This paper proposes a new methodology for testing a core-based system-on-chip (SOC), targeting the simultaneous reduction of test area overhead and test application time. Testing of embedded cores is achieved using the transparency properties of surrounding cores. At the core level, testability and transparency can be achieved by reusing existing logic inside the core, and providing different versions of the core having different area overheads and transparency latencies. At the chip level, the technique analyzes the topology of the SOC to select the core versions that best meet the user's desired test area overhead and test application time objectives. Application of the method to example SOCs demonstrates the ability to design highly testable SOCs with minimized test area overhead, minimized test application time, or a desired trade-off between the two. Significant reduction in area overhead and test application time compared to an existing SOC testing technique is also demonstrated.
Indradeep Ghosh, Sujit Dey, Niraj K. Jha
DAC3
1998 A BIST Scheme for RTL Controller-Data Paths Based on Symbolic Testability Analysis
abstract
This paper introduces a novel scheme for testing register-transfer level controller/data paths using built-in self-test (BIST). The scheme uses the controller netlist and the data path of a circuit to extract a test control/data flow (TCDF) which consists of operations mapped to modules in the circuit and variables mapped to registers. This TCDF is used to derive a set of symbolic justification and propagation paths (known as test environment) to test some of the operations and variables present in it. If it becomes difficult to generate such test environments with the derived TCDFs, a few test multiplexers are added at suitable points in the circuit to increase its controllability and observability. This test environment can then be used to exercise a module or register in the circuit with pseudorandom pattern generators which are placed only at the primary inputs of the circuit. The test responses can be analyzed with signature analyzers which are only placed at the primary outputs of the circuit. Every module in the module library is made random-pattern testable, whenever possible, using gate-level testability insertion techniques. Finally, a BIST controller is synthesized to provide the necessary control signals to form the different test environments during testing, and a BIST architecture is superimposed on the circuit. Experimental results on a number of industrial and university benchmarks show that high fault coverage (>99%) can be obtained with our scheme in a small number of test cycles at an average area (delay) overhead of only 6.4% (2.5%).
Indradeep Ghosh, Niraj K. Jha, Sudipta Bhawmik
DAC2
1998 FACT: A Framework for the Application of Throughput and Power Optimizing Transformations to Control-Flow Intensive Behavioral Descriptions
abstract
In this paper, we present an algorithm for the application of a general class of transformations to control-flow intensive behavioral descriptions. Our algorithm is based on the observation that incorporation of scheduling information can help guide the selection and application of candidate transformations, and significantly enhance the quality of the synthesized solution. The efficacy of the selected throughput and power optimizing transformations is enhanced by the ability of our algorithm to transcend basic blocks in the behavioral description. This ability is imparted to our algorithm by a general technique we have devised. Our system currently supports associativity, commutativity, distributivity, constant propagation, code motion, and loop unrolling. It is integrated with a scheduler which performs implicit loop unrolling and functional pipelining, and has the ability to parallelize the execution of independent iterative constructs whose bodies can share resources. Other transformations can easily be incorporated within the framework. We demonstrate the efficacy of our algorithm by applying it to several commonly available benchmarks. Upon synthesis, behaviors transformed by the application of our algorithm showed up to 6-fold improvement in throughput over an existing transformation algorithm, and up to 4.5-fold improvement in power over designs produced without the benefit of our algorithm.
Ganesh Lakshminarayana, Niraj K. Jha
DAC2
1998 Synthesis of Power-Optimized and Area-Optimized Circuits from Hierarchical Behavioral Descriptions
abstract
We present a technique for synthesizing power- as well as area-optimized circuits from hierarchical data flow graphs under throughput constraints. We allow for the use of complex RTL modules, such as FFTs and filters, as building blocks for the RTL circuit, in addition to simple RTL modules such as adders and multipliers. Unlike past techniques in the area, we also customize the complex RTL modules to match the environment in which they find themselves. We present a fast and efficient algorithm for mapping multiple behaviors onto the same RTL module during the course of synthesis, thus allowing our synthesis system to explore previously unexplored regions of the design space. These techniques are at the core of an iterative improvement based approach which can accept temporary degradation in solution quality in its quest for a globally optimal solution. The moves in our iterative improvement procedure explore optimizations along different dimensions such as functional unit selection, resource allocation, resource sharing, resource splitting, and selection and resynthesis of complex RTL modules. These inter-related optimizations are dynamically traded off with each other during the course of synthesis, thus exploiting the benefits that arise from their interaction. The synthesis framework also tackles other related high-level synthesis tasks such as scheduling, clock selection, and Vdd selection. Experimental results demonstrate that our algorithm produces circuits whose area and power consumption are comparable to or better than those produced using flattened synthesis, within much shorter CPU times. The efficacy of our algorithm in the power-optimization mode is illustrated by the fact that it produces circuits that consume upto 6.7 times less power than area-optimized circuits working at 5 Volts at area overheads not exceeding 50%.
Ganesh Lakshminarayana, Niraj K. Jha
DAC2
1998 Incorporating Speculative Execution into Scheduling of Control-Flow Intensive Behavioral Descriptions
abstract
Speculative execution refers to the execution of parts of a computation before the execution of the conditional operations that decide whether it needs to be executed. It has been shown to be a promising technique for eliminating performance bottlenecks imposed by control flow in hardware and software implementations alike. In this paper, we present techniques to incorporate speculative execution in a fine-grained manner into scheduling of control-flow intensive behavioral descriptions. We demonstrate that failing to take into account information such as resource constraints and branch probabilities can lead to significantly sub-optimal performance. We also demonstrate that it may be necessary to speculate simultaneously along multiple paths, subject to resource constraints, in order to minimize the delay overheads incurred when prediction errors occur. Experimental results on several benchmarks show that our speculative scheduling algorithm can result in significant (upto seven-fold) improvements in performance (measured in terms of the average number of clock cycles) as compared to scheduling without speculative execution. Also, the best and worst case execution times for the speculatively performed schedules are the same as or better than the corresponding values for the schedules obtained without speculative execution.
Ganesh Lakshminarayana, Anand Raghunathan, Niraj K. Jha
DAC3
1998 CASPER: Concurrent Hardware-Software Co-Synthesis of Hard Real-Time Aperiodic and Periodic Specifications of Embedded System Architectures
abstract
Hardware-software co-synthesis of an embedded system requires mapping of its specifications into hardware and software modules such that its real-time and other constraints are met. Embedded system specifications are generally represented by acyclic task graphs. Many embedded system applications are characterized by aperiodic as well as periodic task graphs. Aperiodic task graphs can arrive for execution at any time and their resource requirements vary depending on how their constituent tasks and edges are allocated. Traditional approaches based on a fixed architecture coupled with slack stealing and/or on-line determination of how to serve aperiodic task graphs are not suitable for embedded systems with hard real-time constraints, since they cannot guarantee that such constraints would always be met. In this paper, we address the problem of concurrent co-synthesis of aperiodic and periodic specifications of embedded systems. We estimate the resource requirements of aperiodic task graphs and allocate execution slots on processing elements and communication links for executing them. Our approach guarantees that the deadlines of both aperiodic and periodic task graphs are always met. We have observed that simultaneous consideration of aperiodic task graphs while performing co-synthesis of periodic task graphs is vital for achieving superior results compared to the traditional slack stealing and dynamic scheduling approaches. To the best of our knowledge, this is the first co-synthesis algorithm which provides simultaneous support of periodic and aperiodic task graphs with hard real-time constraints. Application of the proposed algorithm to several examples from real-life telecom transport systems shows that up to 28% and 34% system cost savings are possible over co-synthesis algorithms which employ slack stealing and rate-monotonic scheduling, respectively.
Bharat P. Dave, Niraj K. Jha
DATE2
1998 IMPACT: A High-Level Synthesis System for Low Power Control-Flow Intensive Circuits
abstract
In this paper, we present a comprehensive high-level synthesis system that is geared towards reducing power consumption in control-flow intensive circuits. An iterative improvement algorithm is at the heart of the system. The algorithm searches the design space by handling scheduling, module selection, resource sharing and multiplexer network restructuring simultaneously. The scheduler performs concurrent loop optimization and implicit loop unrolling. It minimizes the expected number of cycles of the schedule without compromising on the minimum and maximum schedule lengths. A fast simulation technique based on trace manipulation aids power estimation in driving synthesis in the right direction. Experimental results demonstrate power reduction of up to 85% with minimal overhead in area over area-optimized designs operating at 5 V.
Kamal S. Khouri, Ganesh Lakshminarayana, Niraj K. Jha
DATE3
1998 CORDS: hardware-software co-synthesis of reconfigurable real-time distributed embedded systems
abstract
Article CORDS: hardware-software co-synthesis of reconfigurable real-time distributed embedded systems Share on Authors: Robert P. Dick Department of Electrical Engineering, Princeton University, 2 Princeton, New Jersey Department of Electrical Engineering, Princeton University, 2 Princeton, New JerseyView Profile , Niraj K. Jha Department of Electrical Engineering, Princeton University, 2 Princeton, New Jersey Department of Electrical Engineering, Princeton University, 2 Princeton, New JerseyView Profile Authors Info & Claims ICCAD '98: Proceedings of the 1998 IEEE/ACM international conference on Computer-aided designNovember 1998 Pages 62–67https://doi.org/10.1145/288548.288561Online:01 November 1998Publication History 66citation479DownloadsMetricsTotal Citations66Total Downloads479Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Robert P. Dick, Niraj K. Jha
ICCAD2
1998 Transforming control-flow intensive designs to facilitate power management
abstract
We present techniques to transform scheduled descriptions of control-flow intensive designs to facilitate power management. We investigate the factors that inhibit the application of power management in synthesized RTL implementations. Based on these insights, we present transformation techniques based on the concepts of variable protection, variable re-naming and re-assignment, and limited controller state memory insertion that result in inherently powermanaged architectures. Our transformation techniques can be easily used in conjunction with any existing resource sharing algorithm or in the framework of existing high-level synthesis tools. Experimental results on control-flow intensive designs indicated reductions of upto 76.6% (35.6% on average) in power consumption at area overheads not exceeding 10.1% (1.1% on average) over already poweroptimized designs.
Ganesh Lakshminarayana, Anand Raghunathan, Niraj K. Jha, Sujit Dey
ICCAD3
1998 Removal of memory access bottlenecks for scheduling control-flow intensive behavioral descriptions
abstract
Application domains like signal and image processing, mul-timedia and networking protocols involve processing of huge amounts of data stored in memory modules. The behavioral de-scriptions of these applications may contain a large number of array references for data accesses. Dependencies between array accesses cause bottlenecks in the derivation of high-performance schedules. In this paper, we introduce a scheduling-integrated technique to identify and remove these bottlenecks. We first demonstrate that there is a significant loss in the quality of a schedule if these bottlenecks are not taken into account by the scheduler. We then propose a technique to overcome these bottle-necks by introducing new operations in the schedule called veri-fication operations. Experimental results on several benchmarks show that a scheduler powered by our technique demonstrates a two-fold improvement in performance (measured in terms of the average number of clock cycles) over a recently-introduced sched-uler for control-flow intensive behavioral descriptions, called Wavesched. Wavesched itself has a two-fold performance advan-tage over traditional methods such as path-based scheduling and loop-directed scheduling. Also, the best- and worst-case execution times for the enhanced schedules obtained by our method are usu-ally equal to or much less than the corresponding values for the execution times obtained by previous schedulers. 1
Srivaths Ravi 0001, Ganesh Lakshminarayana, Niraj K. Jha
ICCAD3
1998 Fast high-level power estimation for control-flow intensive design
abstract
In this paper, we present a power estimation technique for control-flow intensive designs that is tailored towards driving iterative high-level synthesis systems, where hundreds of architectural trade-offs are explored and compared. Our method is fast and relatively accurate. The algorithm utilizes the behavioral information to extract branch probabilities, and uses these in conjunction with switching activity and circuit capacitance information, to estimate the power consumption of a given architecture.
Kamal S. Khouri, Ganesh Lakshminarayana, Niraj K. Jha
ISLPED3
1998 TAO: regular expression based high-level testability analysis and optimization
abstract
In this paper, we present TAO, a novel methodology for high-level testability analysis and optimization of register-transfer level controller/data path circuits. Unlike existing high-level testing techniques that cater restrictively to certain classes of circuits or design styles, TAO exploits the algebra of regular expressions to provide a unified framework for handling a wide variety of circuits including application-specific integrated circuits, application-specific programmable processors, application-specific instruction processors, digital signal processors and microprocessors. We also augment TAO with a design-for-test framework that can provide a low-cost testability solution by examining the trade-offs in choosing from a diverse array of testability modifications like partial scan or test multiplexer insertion in different parts of the circuit. Test generation is symbolic and, hence, independent of bit-width. Experimental results on benchmark circuits show that TAO is very efficient, in addition to being comprehensive. The fault coverage obtained is above 99% in all cases. The average area and delay overheads for incorporating testability into the benchmarks are only 3.3% and 1.1%, respectively. The test application time is comparable to that associated with gate-level sequential test generators.
Srivaths Ravi 0001, Ganesh Lakshminarayana, Niraj K. Jha
ITC3
1998 Guest Editorial
Niraj K. Jha
J. Electron. Test.1
1998 High-level test synthesis: a survey
Indradeep Ghosh, Niraj K. Jha
Integr.2
1998 COHRA: hardware-software cosynthesis of hierarchical heterogeneous distributed embedded systems
abstract
Hardware-software cosynthesis of an embedded system architecture entails partitioning of its specification into hardware and software modules such that its real-time and other constraints are met. Embedded systems are generally specified in terms of a set of acyclic task graphs. For medium- to large-scale embedded systems, the task graphs are usually hierarchical in nature. The embedded system architecture, which is the output of the cosynthesis system, may itself be nonhierarchical or hierarchical. Traditional nonhierarchical architectures create communication and processing bottlenecks and are impractical for large embedded systems. Such systems require a large number of processing elements and communication links connected in a hierarchical manner, thus forming a hierarchical distributed architecture, to meet performance and cost objectives. In this paper, we address the problem of hardware-software cosynthesis of hierarchical heterogeneous distributed embedded system architectures from hierarchical or nonhierarchical task graphs. Our cosynthesis algorithm has the following features: 1) it supports periodic task graphs with real-time constraints, 2) it supports pipelining of task graphs, 3) it supports a heterogeneous set of processing elements and communication links, 4) it allows both sequential and concurrent modes of communication and computation, 5) it employs a combination of preemptive and nonpreemptive static scheduling, 6) it employs a new task-clustering technique suitable for hierarchical task graphs, and 7) it uses the concept of association arrays to tackle the problem of multirate tasks encountered in multimedia systems. We show how our cosynthesis algorithm can be easily extended to consider fault tolerance or low-power objectives or both. Although hierarchical architectures have been proposed before, to the best of our knowledge, this is the first time the notion of hierarchical task graphs and hierarchical architectures has been supported in a cosynthesis algorithm.
Bharat P. Dave, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1998 MOGAC: a multiobjective genetic algorithm for hardware-software cosynthesis of distributed embedded systems
abstract
In this paper, we present a hardware-software cosynthesis system, called MOGAC, that partitions and schedules embedded system specifications consisting of multiple periodic task graphs. MOGAC synthesizes real-time heterogeneous distributed architectures using an adaptive multiobjective genetic algorithm that can escape local minima. Price and power consumption are optimized while hard real-time constraints are met. MOGAC places no limit on the number of hardware or software processing elements in the architectures it synthesizes. Our general model for bus and point-to-point communication links allows a number of link types to be used in an architecture. Application-specific integrated circuits consisting of multiple processing elements are modeled. Heuristics are used to tackle multirate systems, as well as systems containing task graphs whose hyperperiods are large relative to their periods. The application of a multiobjective optimization strategy allows a single cosynthesis run to produce multiple designs that trade off different architectural features. Experimental results indicate that MOGAC has advantages over previous work in terms of solution quality and running time.
Robert P. Dick, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1998 A design-for-testability technique for register-transfer level circuits using control/data flow extraction
abstract
In this paper, we present a technique for extracting functional (control/data flow) information from register-transfer level controller/data path circuits, and illustrate its use in design for hierarchical testability of these circuits. This scheme does not require any additional behavioral information. It identifies a suitable control and data flow from the register-transfer level circuit, and uses it to test each embedded element in the circuit by symbolically justifying its precomputed test set from the system primary inputs to the element inputs and symbolically propagating the output response to the system primary outputs. When symbolic justification and propagation become difficult, it inserts test multiplexers at suitable points to increase the symbolic controllability and observability of the circuit. These test multiplexers are mostly restricted to off-critical paths. Testability analysis and insertion are completely based on the register-transfer level circuit and the functional information automatically extracted from it, and are independent of the data path bit width owing to their symbolic nature. Furthermore, the data path test set is obtained as a byproduct of this analysis without any further search. Unlike many other design-for-testability techniques, this scheme makes the combined controller-data path very highly testable. It is general enough to handle control-flow-intensive register-transfer level circuits like protocol handlers as well as data-flow intensive circuits like digital filters. It results in low area/delay/power overheads, high fault coverage, and very low test generation times (because it is symbolic and independent of bit width). Also, a large part of our system-level test sets can be applied at speed. Experimental results on many benchmarks show the average area, delay, and power overheads for testability to be 3.1, 1.0, and 4.2%, respectively. Over 99% fault coverage is obtained in most cases with two-four orders of magnitude test generation time advantage over an efficient gate-level sequential test pattern generator and one-three orders of magnitude advantage over an efficient gate-level combinational test pattern generator (that assumes full scan). In addition, the test application times obtained for our method are comparable with those of gate-level sequential test pattern generators, and up to two orders of magnitude smaller than designs using full scan.
Indradeep Ghosh, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1998 Integration of hierarchical test generation with behavioral synthesis of controller and data path circuits
abstract
This paper describes the first behavioral synthesis system that incorporates a hierarchical test generation approach to synthesize area-efficient and highly testable controller/data path circuits. Functional information of circuit modules is used during the synthesis process to facilitate complete and easy testability of the data path. The controller behavior is taken into account while targeting data path testability. No direct controllability of the controller outputs through scan or otherwise is assumed. The test set for the combined controller/data path is generated during synthesis in a very short time. Near 100% testability of combined controller and data path is achieved. The synthesis system easily handles large bit-width data path circuits with sequential loops and conditional branches in their behavioral specification, and scheduling constructs like multicycling, chaining and structural pipelining. An improvement of about three to four orders of magnitude was usually obtained in the test generation time for the synthesized benchmarks as compared to an efficient gate-level sequential test generator. The testability overheads are almost zero. Furthermore, in many cases at-speed testing is also possible.
Sandeep Bhatia, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.2
1997 COSYN: Hardware-Software Co-Synthesis of Embedded Systems
abstract
Hardware-software co-synthesis is the process ofpartitioning an embedded system specification into hardware andsoftware modules to meet performance, power and cost goals. Inthis paper, we present a co-synthesis algorithm which starts withperiodic task graphs with real-time constraints and produces a low-costheterogeneous distributed embedded system architecturemeeting the constraints. The algorithm has the following features:1) it allows the use of multiple types of processing elements (PEs)and inter-PE communication links, where the links can take variousforms (point-to-point, bus, local area network (LAN), etc.), 2) itsupports both concurrent and sequential modes of communicationand computation, 3) it allows both preemptive and non-preemptivescheduling, 4) it employs the concept of an association array totackle the problem of multi-rate systems (which are commonlyfound in multimedia applications), 5) it uses a scheduler based ondynamic deadline-based priority levels for accurate performanceestimation of a co-synthesis solution, 6) it uses a new taskclustering technique which takes the dynamic nature of the criticalpath, and the existence of multiple critical paths in the task graphinto account, and 7) if desired, it also optimizes the architecture forpower consumption (we are not aware of any other co-synthesisalgorithm that optimizes power). Application of the proposedalgorithm to examples from the literature and real-life telecomtransport systems shows its efficacy.
Bharat P. Dave, Ganesh Lakshminarayana, Niraj K. Jha
DAC3
1997 Hierarchical Test Generation and Design for Testability of ASPPs and ASIPs
abstract
Abstract — In this paper, we present design for testability (DFT) and hierarchical test generation techniques for facilitating the testing of application-specific programmable processors (ASPP’s) and application-specific instruction processors (ASIP’s). The method utilizes the register-transfer level (RTL) circuit description of an ASPP or ASIP to come up with a set of test microcode patterns which can be written into the instruction read-only memory (ROM) of the processor. These lines of microcode dictate a new control/data flow in the circuit and can be used to test modules which are not easily testable. The new control/data flow is used to justify precomputed test sets of a module from the system primary inputs to the module inputs and propagate output responses from the module output to the system primary outputs. The testability analysis, which is based on the relevant control/data flow extracted from the RTL circuit,
Indradeep Ghosh, Anand Raghunathan, Niraj K. Jha
DAC3
1997 Power Management Techniques for Control-Flow Intensive Designs
abstract
This paper presents a low-overhead controller-based powermanagement technique that re-specifies control signals to reconfigureexisting multiplexer networks and functional units to minimizeunnecessary activity. We demonstrate that conventional powermanagement techniques may often not be suited to control-flowintensive designs, and provide a comprehensive analysis of thepotential negative effects of power management on circuit delay,glitching activity at control and data path signals, and formationof false combinational cycles. We present techniques to performpower management through controller re-specification while avoidingthe above negative effects, and demonstrate the efficiency ofthe techniques through experiments.
Anand Raghunathan, Sujit Dey, Niraj K. Jha, Kazutoshi Wakabayashi
DAC3
1997 MOGAC: a multiobjective genetic algorithm for the co-synthesis of hardware-software embedded systems
abstract
We present a hardware-software co-synthesis system, called MOGAC, that partitions and schedules embedded system specifications consisting of multiple periodic task graphs. MOGAC synthesizes real-time heterogeneous distributed architectures using an adaptive multiobjective genetic algorithm that can escape local minima. Price and power consumption are optimized while hard real-time constraints are met. MOGAC places no limit on the number of hardware or software processing elements in the architectures it synthesizes. Our general model for bus and point-to-point communication links allows a number of link types to be used in an architecture. Application-specific integrated circuits consisting of multiple processing elements are modeled. Heuristics are used to tackle multi-rate systems, as well as systems containing task graphs whose hyperperiods are large relative to their periods. The application of a multiobjective optimization strategy allows a single co-synthesis run to produce multiple designs which trade off different architectural features. Experimental results indicate that MOGAC has advantages over previous work in terms of solution quality and running time.
Robert P. Dick, Niraj K. Jha
ICCAD2
1997 Wavesched: a novel scheduling technique for control-flow intensive behavioral descriptions
abstract
Presents a novel scheduling algorithm targeted towards minimizing the average execution time of control-flow intensive behavioral descriptions. Our algorithm uses a control-data flow graph (CDFG) model, which preserves the parallelism inherent in the application. It explores previously unexplored regions of the solution space through its ability to overlap the schedules of independent iterative constructs whose bodies share resources. It also incorporates well-known optimization techniques like loop unrolling in a natural fashion. This is made possible by a general loop-handling technique which we have devised. Application of the algorithm to several common benchmarks demonstrates up to 4.8-fold improvement in expected schedule length over existing scheduling algorithms, without paying a price in terms of the best- and worst-case schedule lengths required to execute the behavioral description (in fact, frequently, the best/worst-case schedule lengths are also better for our algorithm).
Ganesh Lakshminarayana, Kamal S. Khouri, Niraj K. Jha
ICCAD3
1997 A Low-Overhead Design for Testability and Test Generation Technique for Core-Based Systems
abstract
In a fundamental paradigm shift in system design, entire systems are being built on a single chip, using multiple embedded cores. Though the newest system design methodology has several advantages in terms of time-to-market and system cost, testing such core-based systems is difficult due to the problem of justifying test sequences at the inputs of a core embedded deep in the system, and propagating test responses from the core outputs. In this paper, we present a design for testability and symbolic test generation technique for testing such core-based systems on a chip. The proposed method consists of two parts: (i) core-level DFT to make each core testable and transparent, the latter needed to propagate test data through the cores, and (ii) system-level DFT and test generation to ensure the justification and propagation of the precomputed test sequences and test responses of the core. Since the hierarchical testability analysis technique used to tackle the above problem is symbolic, the system test generation method is independent of the bit-width of the cores. The system-level test set is obtained as a by-product of the testability analysis and insertion method without further search. Besides the proposed test method, the two methods that are currently used in the industry were also evaluated on two example systems: (i) FScan-BScan, where each core is full-scanned, and system test is performed using boundary scan, and (ii) FScan-TBus, where each core is full-scanned, and system test is performed using a test bus. The experiments show that the proposed scheme has significantly lower area overhead, delay overhead, and test application time compared to FScan-BScan and FScan-TBus, without any compromise in the system fault coverage.
Indradeep Ghosh, Niraj K. Jha, Sujit Dey
ITC2
1997 Design for hierarchical testability of RTL circuits obtained by behavioral synthesis
abstract
In recent years, there has been growing interest in behavioral (high-level) synthesis for testability. This is due to the fact that testability features, such as scan or the built-in self-test, may incur large overheads if introduced during logic synthesis in the later phase of the design cycle. Related previous work attempted to generate system-level test sets using hierarchical testability during behavioral synthesis. There, the test generation scheme is independent of bit width and is, therefore, capable of handling complex controller/data path circuits with large data path bit widths (e.g., 32), which has posed a serious challenge to logic-level sequential test generators. However, this previous work is not applicable when another high-level synthesis system is used. In this paper, we present techniques that add minimal test hardware to a given register-transfer level (RTL) circuit obtained by behavioral synthesis in order to ensure that the embedded elements in the circuit are hierarchically testable. An important byproduct of our design for testability (DFT) procedure is a system-level test set that delivers precomputed test sets to each element in the RTL circuit. This eliminates the need to apply gate-level sequential test generation to the combined controller/data path. We performed extensive experiments with several complex controller/data path circuits synthesized by three different high-level synthesis systems which do not target testability. The key advantages of our method, illustrated by these experiments, include: 1) the area, delay, and power overheads incurred for testability are very low (the average area, delay, and power overheads for a large number of benchmarks are 3.5, 0.5, and 3.4%, respectively), 2) both the DFT hardware addition and test generation algorithms are independent of the data path bit width.
Indradeep Ghosh, Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1997 Synthesis of asynchronous circuits for stuck-at and robust path delay fault testability
abstract
In this paper, we present methods for synthesizing multilevel asynchronous circuits to be both hazard free and completely testable. Making an asynchronous two-level circuit hazard free usually requires the introduction of either redundant or nonprime cubes or both. This adversely affects the circuit's testability. However, using extra inputs, which is seldom necessary, and a synthesis-for-testability method, we convert the two-level circuit into a multilevel circuit that is completely testable. To avoid the addition of extra inputs as much as possible, we introduce new exact minimization algorithms for hazard-free two-level logic where we first minimize the number of redundant cubes and then minimize the number of nonprime cubes. We target both the stuck-at and robust path delay fault models using similar methods. However, the area overhead for the latter may be slightly higher than for the former.
Steven M. Nowick, Niraj K. Jha, Fu-Chiung Cheng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1997 SCALP: an iterative-improvement-based low-power data path synthesis system
abstract
In this paper, we present SCALP, a comprehensive low-power data path synthesis system that performs the various high-level synthesis tasks (transformations, scheduling, clock selection, module selection, and hardware allocation and assignment) with an aim of reducing the power consumption in the synthesized data path. Focusing on only one or a small subset of the high-level synthesis tasks makes it difficult to realize the full potential for power savings at the algorithm and architecture levels. Our synthesis algorithms, which are based on an iterative improvement strategy with efficient pruning techniques, are capable of performing the various high-level synthesis tasks (and considering their interactions) in an efficient manner. Supply voltage and clock period pruning strategies are used for quickly eliminating inferior design points during the search for the minimum power solution. Estimating switched capacitance accurately at intermediate stages during high-level synthesis can be challenging since the exact structure of the circuit, which affects both physical capacitance and switching activity, may not be available, and due to the high computational complexity of running register-transfer level power analysis tools several times during high-level synthesis. SCALP overcomes the above problems by maintaining a complete image of the structural register-transfer level (RTL) circuit (this is possible since we have a complete solution at any point during iterative improvement), and employing a very fast switched capacitance estimation technique that Is based on the concept of switched capacitance matrices. Our system can handle diverse module libraries and utilize complex scheduling constructs such as multicycling, chaining, and structural pipelining. Retiming and functional pipelining are used in our system to meet tight performance constraints, and to enable the ensuing synthesis steps to better explore the implementation space. Results on several real-life examples are presented to demonstrate the effectiveness of the algorithm. Power estimates obtained using switch-level simulation after layout indicate that up to an order-of-magnitude of power savings can be obtained using our synthesis system.
Anand Raghunathan, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1997 Analysis and Randomized Design of Algorithm-Based Fault Tolerant Multiprocessor Systems Under an Extended Model
abstract
Reliability of compute-intensive applications can be improved by introducing fault tolerance into the system. Algorithm based fault tolerance (ABFT) is a low-cost scheme which provides the required fault tolerance to the system through system level encoding. In this paper, we propose randomized construction techniques, under an extended model, for the design of ABFT systems with the required fault tolerance capability. The model considers failures in the processors performing the checking operations.
Shalini Yajnik, Niraj K. Jha
IEEE Trans. Parallel Distributed Syst.2
1997 Graceful Degradation in Algorithm-Based Fault Tolerant Multiprocessor Systems
abstract
Algorithm-based fault tolerance (ABFT) is a technique which improves the reliability of a multiprocessor system by providing concurrent error detection and fault location capability to it. It encodes data at the system level and modifies the algorithm to operate on the encoded data in order to expose both transient and permanent faults in any processor. Work done till now in this area takes care of only the fault detection and location part of the problem. However, if spare processors are not available, then after a faulty processor has been located, the work initially assigned to it has to be mapped to some nonfaulty processors in the system in such a way that the fault tolerance capability of the system is still maintained with as small a degradation in performance as possible. In this paper, we propose an integrated deterministic solution to the above problem which combines concurrent error detection and fault location with graceful degradation. There exists no previous deterministic ABFT method for the design of general t-fault locating systems, even for the case of t=1. We propose a general method for designing one-fault locating/s-fault detecting systems. We use an extended model for representing ABFT systems. This model considers the processors computing the checks to be a part of the ABFT system, so that faults in the check computing processors can also be detected and located using a simple diagnosis algorithm, and the checks can be mapped to other nonfaulty processors in the system.
Shalini Yajnik, Niraj K. Jha
IEEE Trans. Parallel Distributed Syst.2
1996 Glitch Analysis and Reduction in Register Transfer Level
abstract
We present design-for-low-power techniques based on glitch reduction for register-transfer level circuits.We analyze the generation and propagation of glitches in both the control and data path parts of the circuit.Based on the analysis, we develop techniques that attempt to reduce glitching power consumption by minimizing generation and propagation of glitches in the RTL circuit.Our techniques include restructuring multiplexer networks (to enhance data correlations, eliminate glitchy control signals, and reduce glitches on data signals), clocking control signals, and inserting selective rising/falling delays.Our techniques are suited to control-flow intensive designs, where glitches generated at control signals have a significant impact on the circuit's power consumption, and multiplexers and registers often account for a major portion of the total power.Application of the proposed techniques to several examples shows significant power savings, with negligible area and delay overheads.
Anand Raghunathan, Sujit Dey, Niraj K. Jha
DAC3
1996 A design for testability technique for RTL circuits using control/data flow extraction
abstract
We present a technique for extracting functional (control/dataflow) information from register transfer level (RTL) controller/data path circuits and illustrate its use in design for hierarchical testability of these circuits. This testing procedure and design for testability (DFT) technique is general enough to handle RTL control flow intensive circuits like protocol handlers as well as data flow intensive circuits like digital filters. It makes the combined controller-data path highly testable and does not require any external behavioral information. This scheme has the advantages of low area/delay/power overheads (average of 3.2%, 0.9% and 4.1%, respectively, for benchmarks), high fault coverage (over 99% for most cases), very low test generation times (because it is independent of bit-width), and the advantage of at-speed testing. Experiments show a 2-to-4 (1-to-3) orders of magnitude test generation time advantage over an efficient gate-level sequential test generator (combinational test generator that assumes full scan).
Indradeep Ghosh, Anand Raghunathan, Niraj K. Jha
ICCAD3
1996 Register-transfer level estimation techniques for switching activity and power consumption
abstract
We present techniques for estimating switching activity and power consumption in register-transfer level (RTL) circuits. Previous work on this topic has ignored the presence of glitching activity at various data path and control signals, which can lead to significant underestimation of switching activity. For data path blocks that operate on word-level data, we construct piecewise linear models that capture the variation of output glitching activity and power consumption with various word-level parameters like mean, standard deviation, spatial and temporal correlations, and glitching activity at the block's inputs. For RTL blocks that operate on data that need not have an associated word-level value, we present accurate bit-level modeling techniques for glitching activity as well as power consumption. This allows us to perform accurate power estimation for control-flow intensive circuits, where most of the power consumed is dissipated in non-arithmetic components like multiplexers, registers, vector logic operators, etc. Since the final implementation of the controller is not available during high-level design iterations, we develop techniques that estimate glitching activity at control signals using control expressions and partial delay information. Experiments on example RTL designs resulted in power estimates that were within 7% of those produced by an inhouse power analysis tool on the final gate-level implementation.
Anand Raghunathan, Sujit Dey, Niraj K. Jha
ICCAD3
1996 Controller re-specification to minimize switching activity in controller/data path circuits
abstract
This paper proposes a controller-based technique for minimizing switching activity in controller/data path circuits. Though the control signals in a register transfer level (RTL) implementation are fully specified, they can be respecified under certain states/conditions when the data path components that they control need not be active. Unlike techniques that insert extra circuitry like transparent latches, controller re-specification is a low-overhead technique that merely reconfigures existing multiplexer networks and functional units to minimize activity in the data path. Hence, it is well suited to control-flow intensive designs, where power consumption in multiplexer networks forms a major component of the total power consumption. Our controller re-specification algorithm consists of constructing an activity graph for each data path component, identifying conditions under which the component need not be active, and re-labeling the activity graph resulting in re-specification of the corresponding control expressions. Application of the proposed technique to several RTL circuits demonstrated the ability to reduce the total (controller+data path) power consumption by up to 51.8% compared to the initial area-optimized implementations, with nominal area and delay overheads.
Anand Raghunathan, Sujit Dey, Niraj K. Jha, Kazutoshi Wakabayashi
ISLPED3
1996 Hardware-Software Co-Design for Test: It's the Last Straw!
J. El-Ziq, Najmi T. Jarwala, Niraj K. Jha, Peter Marwedel, Christos A. Papachristou, Janusz Rajski, John W. Sheppard
VTS3
1996 Synthesis for parallel scan: applications to partial scan and robust path-delay fault testability
abstract
In this paper, we present methods for synthesis of sequential circuits for easy testability based on a scheme previously referred to as synthesis for parallel scan. A heuristic is used to achieve maximal merging of the testability logic with the normal combinational logic of the sequential circuit in order to minimize the area overhead. The latches used in the proposed scheme are normal nonscan latches. The synthesis for parallel scan scheme is applied to two different fault models. To facilitate stuck-at-fault testability of sequential circuits, the synthesis for parallel scan method is augmented with a novel structural analysis technique for selection of latches for partial scan with emphasis on minimization of delay overhead. Experimental results on ISCAS '89 benchmarks resynthesized by our method indicate that one can, in general, achieve the same level of testability with less area and delay overheads, as compared to the conventional structural analysis method for partial scan. A method for synthesizing fully hazard-free robust path-delay fault testable sequential circuits using the concept of synthesis for parallel scan is also presented. When the number of primary inputs of the sequential circuit is at least as large as the number of latches, only normal nonscan latches are required. Experimental results on MCNC and ISCAS '89 benchmarks show the area overheads to be very reasonable. No previous method could ensure complete robust testability of all path-delay faults in a sequential circuit without assuming a completely enhanced scan, even under the above condition. However, when this condition is not met, i.e., the number of primary inputs is less than the number of latches, we introduce the concept of synthesis for maximally parallel enhanced scan. This still provides complete robust testability, however, at the expense of a higher area overhead.
Sandeep Bhatia, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1995 An iterative improvement algorithm for low power data path synthesis
abstract
We address the problem of minimizing power consumption in behavioral synthesis of data-dominated circuits. The complex nature of power as a cost function implies that the effects of several behavioral synthesis tasks like module selection, clock selection, scheduling and resource sharing on supply voltage and switched capacitance need to be considered simultaneously to fully derive the benefits of design space exploration at the behavior level. We present an efficient algorithm for performing scheduling, clock selection, module selection, and resource allocation and assignment simultaneously with an aim of reducing the power consumption in the synthesized data path. The algorithm, which is based on an iterative improvement strategy, is capable of escaping local minima in its search for a low power solution. The algorithm considers diverse module libraries and complex scheduling constructs such as multicycling chaining, and structural pipelining. We describe supply voltage and clock pruning strategies that significantly improve the efficiency of our algorithm by cutting down on the computational effort involved in exploring candidate supply voltages and clock periods that are unlikely to lead to the best solution. Experimental results are reported to demonstrate the effectiveness of the algorithm. Our techniques can be combined with other known methods of behavioral power optimization like data path replication and transformations, to result in a complete data path synthesis system for low power applications.
Anand Raghunathan, Niraj K. Jha
ICCAD2
1995 Design for hierarchical testability of RTL circuits obtained by behavioral synthesis
abstract
Most behavioral synthesis and design for testability techniques target subsequent gate-level sequential test generation, which is frequently incapable of handling complex controller/data path circuits with large data path bit-widths. Hierarchical testing attempts to counter the complexity of test generation by exploiting information from multiple levels of the design hierarchy. We present techniques that add minimal test hardware to the given register-transfer level (RTL) design obtained through behavioral synthesis in order to ensure that all the embedded modules in the circuit are hierarchically testable. An important by-product of our DFT procedure is a system-level test set that is guaranteed to deliver pre-computed module test sets to each module in the RTL circuit. This eliminates the need to apply gate-level sequential test generation to the controller/data path. We performed extensive experiments with several complex data path/controller circuits synthesized by two different high level synthesis systems which do not target testability.
Indradeep Ghosh, Anand Raghunathan, Niraj K. Jha
ICCD3
1995 Task Allocation for Safety and Reliability in Distributed Systems
Santhanam Srinivasan, Niraj K. Jha
ICPP (2)2
1995 An ILP Formulation for Low Power Based on Minimizing Switched Capacitance During Data Path Allocation
abstract
Power consumption has been established as an important factor in the synthesis of VLSI circuits. However, very little work has been done at the behavior level for power estimation and low power synthesis. We demonstrate how hardware sharing affects the capacitance and transition activity, and hence the power dissipation in the synthesized data paths, and that the power-optimal allocation varies with input characteristics. We present a simulation-based method that measures both intra- and inter-iteration effects of hardware sharing on switched capacitance. This method enables us to accurately capture sufficient information to evaluate allocation options from a switched capacitance point of view. Using the information thus obtained, we formulate allocation as an Integer Linear Programming (ILP) problem with the total switched capacitance in the data path as the objective function. Thus, the decisions of how much hardware sharing should be performed and what allocation is best from the power point of view, are performed optimally for the given model. Experimental results are presented including accurate power measurements made at the switch level.
Anand Raghunathan, Niraj K. Jha
ISCAS2
1994 Behavioral Synthesis for Hierarchical Testability of Controller/Data Path Circuits with Conditional Branches
abstract
Most existing behavioral synthesis systems put emphasis on optimizing area and performance. Only recently has some research been done to consider testability during behavioral synthesis. In our previous work, we integrated hierarchical testability with behavioral synthesis of simple digital data path circuits to synthesize highly testable circuits (S. Bhatia and N.K. Jha, 1994). In the current work, we consider the testability of complete controller and data path during behavioral synthesis. The methods presented can easily handle large and complex behavioral descriptions with loops, conditionals, and allow different scheduling constructs, such as pipelining, multicycling, and chaining. The test set for the combined controller/data path is generated during synthesis in a very short time and near 100% fault coverage is obtained for almost all the synthesized circuits at practically zero overheads.>
Sandeep Bhatia, Niraj K. Jha
ICCD2
1994 Behavioral Synthesis for low Power
abstract
We present a behavioral synthesis method targeting low power consumption for data-dominated CMOS circuits. A study of how the high-level synthesis process affects power consumption is presented, based on, which we have developed the first allocation method for low power. We also present a method of optimizing the controller to reduce data path power dissipation. We consider loops, conditional branches, and scheduling constructs such as multicycling, chaining and structural pipelining. The techniques were implemented within the framework of an existing behavioral synthesis system. Experiments performed on various examples and benchmarks show that low power circuits can be synthesized by our method with very low or zero overheads.>
Anand Raghunathan, Niraj K. Jha
ICCD2
1994 Graceful Degradation in Algorithm-Based Fault Tolerant Multiprocessor Systems
abstract
Algorithm-based fault tolerance (ABFT) is a technique for improving the reliability of a multiprocessor system by providing concurrent error detection and fault location capability to it. In this paper, we propose the first integrated solution to the problem of fault detection, location and graceful degradation in ABFT systems. Unlike most previous methods, we use an extended model for representing ABFT systems, which allows faults to occur in check computing processors.>
Shalini Yajnik, Niraj K. Jha
ISCAS2
1994 Synthesis of Fault Tolerant Architectures for Molecular Dynamics
abstract
Molecular dynamics (MD) simulations play an important role in the study of conformations and dynamics of proteins and nucleic acids. This paper suggests ways of mapping the molecular dynamics calculations to various types of parallel architectures, e.g. mesh array, linear array, such that the inherent redundancy involved in the simulation algorithm can be used to make the computation fault tolerant without much overhead.>
Shalini Yajnik, Niraj K. Jha
ISCAS2
1994 A Totally Self-Checking Checker for a Parallel Unordered Coding Scheme
abstract
Bose has developed a parallel unordered coding scheme using only r checkbits for 2/sup r/ information bits. This code can detect all unidirectional errors and requires simple parallel encoding/decoding. The information symbols can be separated from the check symbols. However, the information symbols containing all zeros and all ones need to be transformed to two other information symbols. This allows one to reduce the number of checkbits over Berger code by 1. Since information symbols containing a power-of-two number of bits are quite common, this coding scheme should become quite popular. The authors describe a modular, economical, and easily testable totally self-checking (TSC) checker design for the above code. The TSC concept is well known for providing concurrent error detection of transient as well as permanent faults. The design is self-testing with at most only 2r+16 codeword tests. This means that if k is the number of information bits, the size of the codeword test set is only O(log/sub 2/ k). This is the first known TSC checker design for this code.>
Steven W. Burns, Niraj K. Jha
IEEE Trans. Computers2
1994 Algorithm-Based Fault Tolerance for FFT Networks
abstract
Algorithm-based fault tolerance (ABFT) is a low-overhead system-level fault tolerance technique. Many ABFT schemes have been proposed in the past for fast Fourier transform (FFT) networks. In this paper, a new ABFT scheme for FFT networks is proposed. We show that the new approach maintains the high throughput of previous schemes, yet needs lower hardware overhead and achieves higher fault converge than previous schemes by J.Y. Jou et al. (1988) and D.I. Tao et al. (1990).>
Sying-Jyan Wang, Niraj K. Jha
IEEE Trans. Computers2
1994 Partitioned Encoding Schemes for Algorithm-Based Fault Tolerance in Massively Parallel Systems
abstract
Considers the applicability of algorithm based fault tolerance (ABET) to massively parallel scientific computation. Existing ABET schemes can provide effective fault tolerance at a low cost For computation on matrices of moderate size; however, the methods do not scale well to floating-point operations on large systems. This short note proposes the use of a partitioned linear encoding scheme to provide scalability. Matrix algorithms employing this scheme are presented and compared to current ABET schemes. It is shown that the partitioned scheme provides scalable linear codes with improved numerical properties with only a small increase in hardware and time overhead.>
Jennifer Rexford, Niraj K. Jha
IEEE Trans. Parallel Distributed Syst.2
1994 Design of Algorithm-Based Fault-Tolerant Multiprocessor Systems for Concurrent Error Detection and Fault Diagnosis
abstract
Algorithm-based fault tolerance (ABPT) is a low-overhead system-level concurrent error detection and fault location scheme for multiprocessor systems. We present new methods for the design of ABFT systems. Our design procedure is applicable to a wide range of systems in which processors share data elements. A feature of our design approach is that the type of checks to be used in the final system can be controlled by the system designer. We also present some new bounds on the number of checks needed in ABFT system design.>
Bapiraju Vinnakota, Niraj K. Jha
IEEE Trans. Parallel Distributed Syst.2
1993 Behavioral Synthesis of Highly Testable Data Paths under the Non-Scan and Partial Scan Environments
abstract
Behavioral synthesis tools which only optimize area and performance can easily produce a hard-to-test architecture. In this paper, we propose a new behavioral synthesis algorithm for testability which reduces sequential loop size while minimizing area. The algorithm considers two levels of testability synthesis: synthesis for non-scan, which assumes no test strategy beforehand; and synthesis for partial scan, which uses the available scan information during resource allocation. Experimental results show that in almost all the cases our algorithm can synthesize benchmarks with a very high fault coverage in a small amount of test generation time, using the fewest registers and functional modules. Comparisons are also made with other behavioral synthesis algorithms which disregard testability in order to establish the efficacy of our approach.
Tien-Chien Lee, Niraj K. Jha, Marilyn Wolf
DAC2
1993 Synthesis of Sequential Circuits for Easy Testability Through Performance-Oriented Parallel Partial Scan
abstract
Testing of sequential circuits is greatly facilitated by using scan techniques to directly control and observe the latches. Adding scan capability after logic synthesis may incur significant area and delay overheads. For such cases, we present a technique called synthesis for parallel partial scan in which latches are incrementally selected for partial scan based on some new structural analysis criteria that we introduce. Delay overhead is minimized, in most cases made zero or nearly zero, by judiciously selecting the latches in the non-critical paths. A heuristic is used to maximally merge the scan logic with the combinational logic of the sequential circuit to minimize the area overhead. The latches used in our scheme are normal, non-scan latches. Experimental results on ISCAS '89 benchmarks resynthesized by our method indicate that we can, in general, achieve the same level of testability with fewer latches selected for partial scan, as compared to previous methods based on structural analysis.>
Sandeep Bhatia, Niraj K. Jha
ICCD2
1993 Efficient Diagnosis in Algorithm-Based Fault Tolerant Multiprocessor Systems
abstract
The conventional assumption that checks in an algorithm-based fault tolerant (ABFT) system can be invalidated due to aliasing of erroneous data elements has complicated the task of error detection and location. In this paper we show that aliasing is a very rare occurrence. We then present a simple polynomial-time diagnosis algorithm which takes advantage of this result to run much more efficiently compared to the conventional method of diagnosis. We introduce the concept of NC-detectability and NC-locatability to measure the fault tolerance of the system when check invalidation does not occur and show how to design systems with specified error detectability/NC-detectability and locatability/NC-locatability. For the data-check graphs designed using these methods, when aliasing does not occur, our diagnosis algorithm has a worst case complexity of O(s/sup 2/n/sup 2/log n), where s is the error locatability and n is the number of data elements in the system. We also consider the case where the processors which compute the checks themselves fail.>
Santhanam Srinivasan, Niraj K. Jha
ICCD2
1993 Design of Algorithm-Based Fault Tolerant Systems With In-System Checks
abstract
To improve the reliability of computeintensive applications run on multiprocessor architec tures, fault tolerance is introduced into the system with on-line detection and location of faults. This can be achieved by a low-cost scheme, called Algorithm-based fault tolerance (ABFT), which encodes data at the system level and modifies the algorithm to operate on the encoded data. The resultant encoded output data is checked for correctness by some checks. In this pa per we present an extended model for representing and designing ABFT systems. The model takes into con sideration the processors evaluating the checks. We propose a design method which considers the proces sors computing the checks to be a part of the ABFT system and guarantees concurrent error detection even in the presence of faults in these processors, unlike most methods presented earlier.
Shalini Yajnik, Niraj K. Jha
ICPP (1)2
1993 A Conditional Resource-Sharing Method for Behavior Synthesis of Highly- Testable Data Paths
abstract
Existing conditional resource sharing methods using in behavioral synthesis focus on area and performance optimization and do not consider testability. This paper extends our previous work to handle conditional branches. A hierarchical control-data flow graph (HCDFG) is used to model the system behavior. A postorder traversal of the HCDFG is employed to reduce sequential depths and loops for testability synthesis. Experimental results for the benchmarks show that our method, with no a priori test strategy assumption, can achieve higher fault coverage in shorter test generation time than an algorithm which disregards testability, and, with partial scan test assumption, can have high testability with fewer scan registers than some design-for-test methods.>
Tien-Chien Lee, Niraj K. Jha, Marilyn Wolf
ITC2
1993 Fault Detection in CVS Parity Trees with Application to Strongly Self-Checking Parity and Two-Rail Checkers
abstract
The problem of single stuck-at, stuck-open, and stuck-on fault detection in cascode voltage switch (CVS) parity trees is considered. The results are also applied to parity and two-rail checkers. It is shown that, if the parity tree consists of only differential cascode voltage switch (DCVS) EX-OR gates, then the test set consists of at most five vectors (in some cases only four vectors are required) for detecting all detectable single stuck-at, stuck-open, and stuck-on faults, independent of the number of primary inputs and the number of inputs to any EX-OR gate in the tree. If, however, only a single-ended output is desired from the tree, then the final gate will be a single-ended cascode voltage switch (SCVS) EX-OR gate, for which the test set has only eight vectors. For a strongly self-checking (SSC) CVS parity checker, the size of a test set consisting of only codewords is nine, whereas for an SSC CVS two-rail checker the size of a test set consisting of only codewords is at most five.>
Niraj K. Jha
IEEE Trans. Computers1
1993 Optimal Design of Checks for Error Detection and Location in Fault-Tolerant Multiprocessor Systems
abstract
RANDGEN, a simple and efficient general-purpose algorithm for generating arbitrary data-check (DC) graphs with a small number of checks, which satisfy a variety of properties that have been found to be useful in algorithm-based fault tolerance (ABFT) designs, is proposed. The concept of majority diagnosability is introduced in an attempt to explicitly redesign DC graphs for easy diagnosis. UNIFGEN, a variation of RANDGEN that produces DC graphs with uniform checks is examined.>
Ramesh K. Sitaraman, Niraj K. Jha
IEEE Trans. Computers2
1993 Diagnosability and Diagnosis of Algorithm-Based Fault-Tolerant Systems
abstract
Parallel processing architectures are commonly used for signal processing and other computationally intensive applications. These applications are characterized by high throughput and long processing periods. Such characteristics decrease the reliability of high-performance architectures. The erroneous data produced by faulty processors could have damaging consequences, particularly in critical real-time applications. It is therefore desirable that any erroneous data produced by the system be detected and located as quickly as possible. Algorithm-based fault tolerance (ABFT) is a low-cost system-level concurrent error detection and fault location scheme. Methods used in the analysis of multiprocessor systems using system-level diagnosis are applied to the analysis of ABFT systems. A new algorithm for analyzing an ABFT system for its fault diagnosability is developed using these methods. Based on this work, a fault diagnosis algorithm is developed for ABFT systems.>
Bapiraju Vinnakota, Niraj K. Jha
IEEE Trans. Computers2
1993 Easily testable nonrestoring and restoring gate-level cellular array dividers
abstract
The problem of design for testability of gate-level dividers is investigated using the single stuck-at fault model. Two C-testable designs are given for the nonrestoring array divider. The first design is C-testable with only eight vectors. The additional hardware required to obtain C-testability for an n*n nonrestoring array divider consists of n-1 two-input XOR gates and one control input. The second design does not require any extra circuitry or control inputs, and is C-testable with six vectors. In other words, the basic array structure itself is C-testable. Two easily testable designs for the restoring array divider are also presented. The first n*n array design is shown to be linearly testable with 2n+8 vectors. In the second design, a logic implementation that makes the design C-testable with only six vectors is used. The extra circuitry required for both designs consists of n XOR gates and one control input.>
Niraj K. Jha, Abha Ahuja
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1993 Design and synthesis of self-checking VLSI circuits
abstract
Self-checking circuits can detect the presence of both transient and permanent faults. A self-checking circuit consists of a functional circuit that produces encoded output vectors and a checker that checks the output vectors. The checker has the ability to expose its own faults as well. The functional circuit can be either combinational or sequential. A self-checking system consists of an interconnection of self-checking circuits. The advantage of such a system is that errors can be caught as soon as they occur; thus, data contamination is prevented. Methods for the cost-effective design of combinational and sequential self-checking functional circuits and checkers are examined. The area overhead for all proposed design alternatives is studied in detail.>
Niraj K. Jha, Sying-Jyan Wang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1993 Synthesis of Algorithm-Based Fault-Tolerant Systems from Dependence Graphs
abstract
Algorithm-based fault tolerance (ABFT) is a method for improving the reliability of parallel architectures used for computation-intensive tasks. A two-stage approach to the synthesis of ABFT systems is proposed. In the first stage, a system-level code is chosen to encode the data used in the algorithm. In the second stage, the optimal architecture to implement the scheme is chosen using dependence graphs. Dependence graphs are a graph-theoretic form of algorithm representation. The authors demonstrate that not all architectures are ideal for the implementation of a particular ABFT scheme. They propose new measures to characterize the fault tolerance capability of a system to better exploit the proposed synthesis method. Dependence graphs can also be used for the synthesis of ABFT schemes for non-linear problems. An example of a fault-tolerant median filter is provided to illustrate their utility for such problems.>
Bapiraju Vinnakota, Niraj K. Jha
IEEE Trans. Parallel Distributed Syst.2
1992 Behavioral synthesis for easy testability in data path scheduling
abstract
A data path scheduling algorithm to improve testability without assuming any particular test strategy is presented. A scheduling heuristic for easy testability, based on previous work on data path allocation for testability, is introduced. A mobility path scheduling algorithm to implement this heuristic while also minimizing area is developed. Experimental results on benchmark and example circuits show high fault coverage, short test generation time, and little or no area overhead.>
Tien-Chien Lee, Marilyn Wolf, Niraj K. Jha
ICCAD3
1992 Multiple Input Bridging Fault Detection in CMOS Sequential Circuits
abstract
Bridging fault testing algorithms for CMOS sequential circuits, assuming current supply monitoring, are discussed. Sequential circuits implemented both with and without scan design are considered. Experimental results are given to show the efficacy of the methods.>
Niraj K. Jha, Sying-Jyan Wang, Phillip C. Gripka
ICCD1
1992 Behavioral Synthesis for Easy Testability in Data Path Allocation
abstract
The first behavioral synthesis scheme for improving testability in data path allocation independent of test strategy is presented. The authors propose two behavioral synthesis-for-test heuristics: improve observability and controllability of registers, and reduce sequential depth between registers. Also presented are algorithms that optimize a behavior-level design using these two criteria while minimizing area. Experimental results for benchmark circuits synthesized by the author's experimental system. PHITS, show that these methods give a high fault coverage in small amounts of CPU time at a low area overhead.>
Tien-Chien Lee, Marilyn Wolf, Niraj K. Jha, John M. Acken
ICCD3
1992 A totally self-checking checker for a parallel unordered coding scheme
abstract
Bose has developed a parallel unordered coding scheme using only r checkbits for 2/sup r/ information bits. This code can detect all unidirectional errors and requires simple parallel encoding/decoding. The information symbols can be separated from the check symbols. However, the information symbols containing all-0's and all-1's need to be transformed to two other information symbols. This allows one to reduce the number of checkbits over Berger code by 1. Since information symbols containing a power-of-two number of bits are quite common, this coding scheme should become quite popular. In this paper, a modular, economical and easily testable totally self-checking (TSC) checker design for the above code is described. The TSC concept is well-known for providing concurrent error detection of transient as well as permanent faults. The design is self-testing with at most only 2r+16 codeword tests. This means that if k is the number of information bits, the size of the codeword test is only O(log/sub 2/k). This is the first known TSC checker design for this code.>
Steven W. Burns, Niraj K. Jha
VTS2
1991 Design and Synthesis of Self-Checking VLSI Circuits and Systems
abstract
Self-checking circuits and systems can detect the presence of both transient and permanent faults. The advantage of such a system is that errors can be caught as soon as they occur, and thus data contamination is prevented. Although much effort has been concentrated on the design of self-checking checkers by previous researchers, very few results have been presented for the design of self-checking functional circuits, and fewer still for the design of self-checking systems. Methods are explored for the cost-effective design of combinational and sequential functional circuits, checkers and systems.>
Niraj K. Jha, Sying-Jyan Wang
ICCD1
1991 A new transition count method for testing of logic circuits
abstract
The authors propose a transition count method for detecting faults in single- and multiple-output logic circuits. It can be extended to sequential circuits in which scan design is incorporated. This method is called double transition count (DTC) testing for single-output circuits and multiple transition count (MTC) testing for multiple-output circuits. It is shown that the detectability of faults obtained by DTC and MTC testing is the same as that obtained by conventional testing. Hence, this method does not result in any information loss even though the set of output vectors is considerably compressed. The size of a DTC or MTC test is equal to the size of the equivalent conventional test since no test vectors need to be repeated. The basic test circuitry required is very simple, consisting of only one flip flop, one OR gate, one inverter, and one switch per output.>
Konstantinos I. Diamantaras, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1991 Totally self-checking checker designs for Bose-Lin, Bose, and Blaum codes
abstract
The totally self-checking (TSC) circuit concept is well established in the area of concurrent error detection (CED). The outputs of a functional circuit are encoded and are monitored by a TSC checker. Thus errors can be detected concurrently with normal operation. Both permanent and transient faults can be detected. Some efficient systematic codes have been developed for detecting unidirectional errors in t or fewer bits (Bose-Lin code) and for detecting burst unidirectional errors (Bose and Blaum codes). Since unidirectional errors are the most common errors in VLSI circuits, such codes should find widespread use. TSC checker designs have been found for the three codes mentioned above. The designs are easily testable, relatively economical, and have a modular structure.>
Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1991 Design of robustly testable combinational logic circuits
abstract
An integrated approach to the design of combinational logic circuits in which all path delay faults and multiple line stuck-at, transistor stuck-open faults are detectable by robust tests is proposed. Robustly testable static CMOS primitive logic circuit designs are presented for any arbitrary combinational logic function. They require no special gates, and fan-in and fan-out constraints do not affect the designs. Extra controllable inputs or additional hardware to achieve testability was not used. It is demonstrated that the method guarantees the design of CMOS logic circuits in which all path delay faults are locatable.>
Sandip Kundu, Sudhakar M. Reddy, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1991 Easily testable gate-level and DCVS multipliers
abstract
Some C-testable designs of a carry-save parallel multiplier are presented. Results are given for both the gate-level implementation and the differential cascode voltage switch (DCVS) implementation. DCVS circuits are dynamic CMOS circuits which have the advantage of being protected against test set invalidation due to circuit delays. In the first gate-level design, it is assumed that the full-adders have any arbitrary, irredundant logic implementation. Such a design is C-testable with only nine test vectors, which detect all single stuck-at-faults. For a specific logic implementation of the full-adders, another design is shown to be C-testable with only six test vectors. The DCVS design is also C-testable with only six test vectors, which detect all detectable stuck-at, and stuck-open faults in the circuit. Both the hardware and delay overhead for all C-testable designs are very small. For three C-testable designs of the 32 by 32 multiplier, the hardware overhead is 2.7% or less and the delay overhead is 2.4% or less.>
Andres R. Takach, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1990 Design of robustly testable static CMOS parity trees derived from binary decision diagrams
abstract
A robustly testable design of static CMOS parity trees is presented. A test set for such a tree which cannot be invalidated in the presence of arbitrary timing skews and/or circuit delays can be derived. The constituents of the parity tree are static CMOS EXCLUSIVE-OR (EX-OR) gates, which are constructed from their corresponding binary decision diagrams (BDDs). The EX-OR gates in the tree can have any number of inputs. The robust test set detects all the single stuck-open, stuck-on, and stuck-at faults when both logic and current monitoring are done. It is shown that such implementations of parity trees are testable with a test set of size O(log/sub k/n), where n is the number of primary inputs and k is a circuit parameter.>
Niraj K. Jha, Carol Q. Tong
ICCD1
1990 Strong fault-secure and strongly self-checking domino-CMOS implementations of totally self-checking circuits
abstract
The totally self-checking (TSC) concept is well established for applications in the area of online error-indication. TSC circuits can detect both transient and permanent faults. They consist of a functional circuit with encoded inputs and outputs and a checker which monitors these outputs. The TSC concept can be generalized for the functional circuits using the strongly fault-secure (SFS) concept. Here, the concept of strongly self-checking (SSC) circuits, which is a generalization from TSC circuits, is introduced. Most of the TSC circuits presented in the literature are designed at the logic gate level using the stuck-at fault model. However, this fault model is inadequate for MOS technologies. Here, it is shown that a TSC gate-level functional circuit can be implemented in the domino-CMOS technology as an SFS circuit, while a TSC gate-level checker can be implemented as an SSC checker. For the domino-CMOS implementation the fault model is enlarged to stuck-at, stuck-open, and stuck-on faults. It is shown that domino-CMOS is much more suitable for implementation of self-checking circuits than static CMOS.>
Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1989 Design of sufficiently strongly self-checking embedded checkers for systematic and separable codes
abstract
Totally self-checking (TSC) circuits are considered. If the first erroneous output is a noncodeword then the TSC goal is said to be achieved. The author has recently defined a new class of checkers which meet the TSC goal, namely, the strongly self-checking (SSC) checkers. A TSC or SSC checker can be guaranteed to meet the TSC goal if it receives the required set of codeword tests from the functional circuit. However, in practice, it is not always possible to ensure this for embedded checkers. A design for a sufficiently strongly self-checking (SSSC) embedded checker for systematic and separable codes is presented. SSSC checkers can be used when it is impossible to design a TSC or SSC checker based on the available set of codewords from the functional circuit.>
Niraj K. Jha
ICCD1
1989 Separable codes for detecting unidirectional errors
abstract
Systematic and separable codes are widely used for error detection in VLSI circuits. Recently, systematic codes have been found for detecting t-unidirectional and burst unidirectional errors. In such codes, r checkbits are added to k information bits to form a codeword of length (k+r), and it is assumed that all 2/sup k/ possible information k-tuples are present. However, in many applications all 2/sup k/ possible k-tuples are not used. In such cases it is possible to obtain codes usually referred to as separable codes. The author presents separable codes which can detect t or all unidirectional errors.>
Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1989 A totally self-checking checker for Borden's code
abstract
It is well known that a large number of errors in VLSI circuits are unidirectional in nature. J.M. Borden (Inf. Control, vol.53, p.66-73, April 1982) has studied a code for detecting only t, not all, unidirectional errors. For a code to be utilized in a self-checking system, a self-checking checker must exist for it. The totally self-checking (TSC) checker is one such checker. Many designs for TSC checkers exist for m-out-of-n codes. However, no work has been reported on the design of a TSC checker for Borden's code. The author presents such a design. They believe that this design may lead to greater applicability of Borden's code to self-checking systems.>
Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1988 On the design of robust multiple fault testable CMOS combinational logic circuits
abstract
Tests that detect modeled faults independent of the delays in the circuit under test are called robust tests. An integrated approach to the design of combinational logic circuits in which all single stuck-open faults and path delay faults are detectable by robust tests was presented by the authors earlier. It is shown that the earlier design actually results in circuits in which all multiple stuck-at and stuck-open and multipath delay faults are robustly testable. The tests to detect such faults are presented.>
Sandip Kundu, Sudhakar M. Reddy, Niraj K. Jha
ICCAD3
1988 A new class of symmetric error correcting / unidirectional error detecting codes
abstract
A new class of t-error correcting/all or d (d>t) unidirectional error-detecting (UED) code is developed. Such codes can correct symmetric errors in the information part of the codeword. When the added checkbits also have symmetric errors, the errors are detected. Furthermore, (t+1) or more unidirectional errors can be detected by the t-IEC (information error correcting)/A-UED code, while unidirectional errors in (t+1) or more but d or fewer bits can be detected by the t-IEC/d-UED code. By giving up the capability of correcting symmetric errors in the added checkbits a very substantial saving in the length of the codeword is obtained. Such codes can offer protection against transient as well as permanent faults.>
Niraj K. Jha
ICCD1
1988 Multiple Stuck-Open Fault Detection in CMOS Logic Circuits
abstract
It is shown that a test set based on two-pattern tests, which are designed to detect single stuck-open faults, can be found that detects all multiple stuck-open faults inside any CMOS gate in the circuit. The concept is extended to three-pattern tests, which are obtained for every single stuck-open fault at the checkpoints. If a certain condition is satisfied, then it can be shown that the resulting test set can detect any multiple stuck-open fault in the circuit. Even when this condition is not fully met, a very large percentage of the multiple stuck-open faults can still be guaranteed to be detected. For the special case of fan-out-free CMOS circuits, it is shown that a single stuck-open fault test set based on two-pattern tests can always be found that has 100% multiple stuck-open fault coverage. This test can also be guaranteed to be robust in the presence of arbitrary delays.>
Niraj K. Jha
IEEE Trans. Computers1
1988 A universal test set for CMOS circuits
abstract
A universal test set for CMOS circuits is demonstrated that can be derived from the functional description of the circuit alone. It is shown that for a restricted class of CMOS circuits, the gate-level universal test set (UTS/sub g/) consisting of maximal false vectors and minimal true vectors can sensitize every detectable stuck-open fault in the circuit. A universal initialization set (UIS) is defined which can also be derived from just the functional description, and which contains initialization vectors for each of the test vectors. This set consists of maximal true vectors and minimal false vectors. It is shown that a test set on UTS/sub g/ and UIS can be guaranteed to detect every detectable stuck-open fault in both redundant and irredundant CMOS implementation of the function, even in the presence of arbitrary delays and timing-skews. The size of the test set is also investigated.>
Gopal Gupta 0001, Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1988 Testing for multiple faults in domino-CMOS logic circuits
abstract
The problem of multiple faults detection in domino-CMOS logic circuits is considered. The multiple faults can be of the stuck-open and stuck-on types. It is shown that a multiple fault in the domino-CMOS circuit can be mapped to a multiple stuck-at fault in its gate-level model. A method is given to initialize the domino-CMOS circuit and apply a multiple stuck-at fault test set based on the gate-level model of the circuit. This results in the detection of all multiple faults having detectable consistent faults. The problem of test set invalidation due to arbitrary signal delays is easily taken care of in domino-CMOS circuits, making such an implementation of a function even more attractive than a fully complementary CMOS implementation, from the testability point of view.>
Niraj K. Jha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1986 Detecting Multiple Faults in CMOS Circuits
Niraj K. Jha
ITC1
1985 Design of Testable CMOS Logic Circuits Under Arbitrary Delays
abstract
The sequential behavior of CMOS logic circuits in the presence of stuck-open faults requires that an initialization input followed by a test input be applied to detect such a fault. However, a test set based on the assumption that delays through all gates and interconnections are zero, can be invalidated in the presence of arbitrary delays in the circuit. In this paper, we will present a necessary and sufficient condition for the existence of a test set, which cannot be invalidated under arbitrary delays, for an AND-OR or OR-AND CMOS realization for any given function. We will also introduce a Hybrid CMOS realization which, for any given function, is guaranteed to have a valid test set under arbitrary delays.
Niraj K. Jha, Jacob A. Abraham
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1