EDBT 2026 Demo / reviewers in the wild / expert
Anantha P. Chandrakasan
dblp:c/AnanthaChandrakasan
· DBLP profile ↗
139ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0002-5977-2748ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 89 · 5 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 1 since 2021Computer networks · 15 · 1 since 2021Software engineering, systems software and programming languages · 11 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI WorkloadsabstractAs AI workloads drive increases in datacenter power consumption, accurate GPU power estimation is critical for proactive power management. However, existing power models face a scalability bottleneck not in the modeling techniques themselves, but in obtaining the hardware utilization inputs they require. Conventional approaches rely on either costly simulation or hardware profiling, which makes them impractical when rapid predictions are required. This work presents EnergAIzer, which addresses this scalability bottleneck by developing a lightweight solution to predict utilization inputs, reducing the estimation walltime from hours to seconds. Our key insight is that kernels in AI workloads commonly employ optimizations that create structured patterns, which analytically determine memory traffic and execution timeline. We construct a performance model using these patterns as an analytical scaffold for empirical data fitting, which also naturally exposes module-level utilization. This predicted utilization is then fed into our power model to estimate dynamic power consumption. EnergAIzer achieves $8 \%$ power errors on NVIDIA Ampere GPUs, competitive with traditional power models with elaborate cycle-level simulation or hardware profiling. We demonstrate EnergAIzer’s exploration capabilities for frequency scaling and architectural configurations, including forecasting the power of NVIDIA H100 with just $7 \%$ error. In summary, EnergAIzer provides fast and accurate power prediction for AI workloads, paving the way for power-aware design explorations. Kyungmi Lee, Zhiye Song, Xin Zhang 0025, Tamar Eilam, Anantha P. Chandrakasan |
ISPASS | 6 |
| 2026 | Securing DNN Acceleration From Off-Chip Memory Vulnerabilities With Low-Overhead Authenticated EncryptionabstractSecurity vulnerabilities in deep neural network (DNN) accelerators pose risks for high-stakes applications, with off-chip memory attacks representing a critical threat to both data confidentiality and integrity. While general-purpose processors employ comprehensive cryptographic authenticated encryption for memory security, domain-specific DNN accelerators lack adequate protection, particularly against integrity violations. To address this research gap, we present Sorbet, a DNN accelerator equipped with authenticated encryption to defend against both confidentiality and integrity attacks on off-chip memory. Integrity verification introduces complex memory access patterns in DNN accelerators, as the granularity of authentication operations often clashes with the tiling strategies used for efficient off-chip memory access. Our approach tackles this challenge with a secure memory interface (SMI) module that efficiently: 1) translates the accelerator’s tile request to the required data for cryptographic authentication and 2) aligns fetched data with the memory map of the accelerator’s on-chip buffers. Moreover, our design mitigates the area and performance overhead of cryptographic operations by adopting a lightweight cipher while maintaining security requirements against splicing and replay attacks. Our fabricated chip achieves 22% latency overhead across diverse workloads, including convolutions and multihead attentions (MHAs), which can be further reduced with larger on-chip buffer size and double-buffering. It incurs only 7.9% area and 18.3% energy overhead, which is competitive with recent DNN accelerator defenses with weaker off-chip memory protection. Kyungmi Lee, Gaurab Das, Donghyeon Han, Anantha P. Chandrakasan |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | MEGA.mini: A NPU with Novel Heterogeneous AI Processing Architecture Balancing Efficiency, Performance, and Intelligence for the Era of Generative AIabstract• NPU with a Novel big.LITTLECore Architecture to Balance 3 Key Aspects of AI Acceleration–Efficiency: > 95% computations w/ Low-precision FXP–Performance: 3 hierarchical solutions @ MEGA+mini–Intelligence: Hybrid IA (FP for < 5% outlier data) Donghyeon Han, Anantha P. Chandrakasan |
HCS | 2 |
| 2025 | Efficient Circuit Performance Prediction Using Machine Learning: From Schematic to Layout and Silicon Measurement with Minimal Data InputabstractWe present an ML-driven framework for predicting circuit performance metrics, bridging the gap between schematic and layout simulations, multi-process corner analysis, and measured silicon data. We focus on 14nm and 5nm FinFET-based ring oscillators, collecting data across varying supply voltages, temperatures, and process corners. Using three baseline ML models—XGBoost, Random Forest, and a Neural Network—we simulate real-world design scenarios where parameter fine-tuning may not always be feasible. Key tasks include predicting layout performance from schematic data, performance prediction across process corners, and predicting measured chip performance via transfer learning. Our results show that these models can achieve less than 5% mean absolute percentage error (MAPE) for power and frequency prediction while reducing required simulations by more than 2×. In migrating from 14nm to 5nm, XGBoost and Neural Network achieve high accuracy (>0.99 R2) using just 10% of 5nm simulations. This framework offers a promising approach to accelerating circuit design across technology nodes, reducing simulation costs while maintaining accuracy in predicting performance. Dimple Vijay Kochar, Maitreyi Ashok, John Cohn, Anantha P. Chandrakasan, Xin Zhang 0025 |
ISCAS | 4 |
| 2025 | Protecting the Mixed-Signal Domain: Secure ADCs for Internet of Things Devices
Maitreyi Ashok, Ruicong Chen, Taehoon Jeong, Anantha P. Chandrakasan, Hae-Seung Lee |
Proc. IEEE | 4 |
| 2025 | Efficient Circuit Performance Prediction Using Machine Learning: From Schematic to Layout and Silicon Measurement With Minimal Data InputabstractWe present an ML-driven framework for predicting circuit performance metrics, bridging the gap between schematic and layout simulations, multi-process corner analysis, and measured silicon data. We demonstrate this using 14nm and 5nm FinFET-based ring oscillators, by collecting data across varying supply voltages, temperatures, and process corners. Using three baseline ML models—XGBoost, Random Forest, and a Neural Network—we simulate real-world design scenarios where parameter fine-tuning may not always be feasible. Key tasks include predicting layout performance from schematic data, performance prediction across process corners, and fabricated chip performance. Our results show that these models can achieve less than 5% mean absolute percentage error (MAPE) for power and frequency prediction while reducing required simulations by more than$2\times $. When migrating from 14nm to 5nm, XGBoost and Neural Network achieve high accuracy (>0.99$R^{2}$) using just 10% of the otherwise required 5nm simulations. We also present an extensive robustness analysis to demonstrate that our results are not limited to a single data split or initialization. By varying random seeds across multiple runs, we evaluate the stability of each model with respect to algorithm initialization and the selection of training data subsets. This demonstrates that the observed accuracy is consistent and not the result of a specific, favorable configuration. This framework offers a promising approach to accelerating circuit design across technology nodes by reducing simulation costs while maintaining accuracy in predicting performance. Dimple Vijay Kochar, Maitreyi Ashok, John Cohn, Xin Zhang 0025, Anantha P. Chandrakasan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | Heterogeneously Integrated Nitrogen-Vacancy Sensing for Real-Time CMOS Security Threat DetectionabstractThis work proposes a prototype system for utilizing nitrogen-vacancy center-based quantum sensing for generalized threat detection systems. Changes to the operation or environment of an IC will cause differences in the magnetic field emanations, which can be detected through changes in a spin-state-dependent photocurrent within a diamond. Threat detection circuitry can be integrated within the sensitive CMOS IC itself at a high spatial resolution for real-time monitoring and spatially resolved low-overhead protections. The key contributions of this work are the CMOS and NV center system for high magnetometer sensitivity while maintaining CMOS design flexibility, the novel security application for quantum sensing, and the proposed method of heterogeneous integration for a complete system. Maitreyi Ashok, Hanfeng Wang, Ethan G. Arnault, Hamza Raniwala, Aya G. Amer, Matthew Trusheim, Dirk R. Englund, Anantha P. Chandrakasan |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2024 | HEAP: A Fully Homomorphic Encryption Accelerator with Parallelized BootstrappingabstractFully homomorphic encryption (FHE) is a cryptographic technology with the potential to revolutionize data privacy by enabling computation on encrypted data. Lately, the CKKS FHE scheme has become quite popular because it can process real numbers. However, CKKS computing is not pervasive yet because it is resource-intensive both in terms of compute and memory, and is multiple orders of magnitude slower than computing on unencrypted data. The recent algorithmic and hardware optimizations to accelerate CKKS computing are promising, but CKKS computing continues to underperform due to an expensive operation known as bootstrapping. While there have been several efforts to accelerate bootstrapping, it continues to remain the main performance bottleneck. One of the reasons for this performance bottleneck is that unlike the non-bootstrapping parts of CKKS computing the bootstrapping algorithm is inherently sequential and exhibits interdependencies among the data. To address this challenge, in this paper, we introduce HEAP an accelerator that uses a hybrid scheme-switching approach. HEAP uses the CKKS scheme for the non-bootstrapping steps, but switches to the TFHE scheme when performing the bootstrapping step of the CKKS scheme. The hybrid approach transitions to the TFHE scheme by extracting coefficients from a single RLWE ciphertext to represent multiple LWE ciphertexts. We incorporate the bootstrapping function into the TFHE BlindRotate operation and simultaneously apply the BlindRotate operation to all LWE ciphertexts. A parallelized execution of bootstrapping is then feasible because there are no data dependencies between distinct LWE ciphertexts. With our approach, we require smaller-sized bootstrapping keys leading to about $18 \times$ less amount of data to be read from the main memory for the keys. In addition, we introduce a variety of hardware optimizations in HEAP—from modular arithmetic level to NTT and BlindRotate datapath optimizations. The approach in HEAP is agnostic of the hardware and can be mapped to any system with multiple compute nodes. To evaluate HEAP, we implemented it in RTL and mapped it to a single FPGA system and an eight-FPGA system. Our comprehensive evaluation of HEAP for the bootstrapping operation shows a $15.39 \times$ improvement when compared to FAB. Similarly, evaluation of HEAP for the logistic regression model training shows $14.71 \times$ and $11.57 \times$ improvement when compared to FAB and FAB-2 implementations, respectively. Rashmi S. Agrawal 0001, Anantha P. Chandrakasan, Ajay Joshi |
ISCA | 2 |
| 2023 | FAB: An FPGA-based Accelerator for Bootstrappable Fully Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) offers protection to private data on third-party cloud servers by allowing computations on the data in encrypted form. To support general-purpose encrypted computations, all existing FHE schemes require an expensive operation known as "bootstrapping". Unfortunately, the computation cost and the memory bandwidth required for bootstrapping add significant overhead to FHE-based computations, limiting the practical use of FHE.In this work, we propose FAB, an FPGA-based accelerator for bootstrappable FHE. Prior FPGA-based FHE accelerators have proposed hardware acceleration of basic FHE primitives for impractical parameter sets without support for bootstrapping. FAB, for the first time ever, accelerates bootstrapping (along with basic FHE primitives) on an FPGA for a secure and practical parameter set. The key contribution of this work is the architecture of a balanced FAB design, which is not memory bound. In our design, we leverage recent algorithms for bootstrapping while being cognizant of the compute and memory constraints of our FPGA. In addition, we use a minimal number of functional units for computing, operate at a low frequency, leverage high data rates to and from main memory, utilize the limited on-chip memory effectively, and perform careful operation scheduling.We evaluate FAB using a single Xilinx Alveo U280 FPGA and by scaling it to a multi-FPGA system consisting of eight such FPGAs. For bootstrapping a fully-packed ciphertext, while operating at 300MHz, FAB outperforms existing state-of-the-art CPU and GPU implementations by 213× and 1.5× respectively. Our target FHE application is training a logistic regression model over encrypted data. For logistic regression model training scaled to 8 FPGAs on the cloud, FAB outperforms a CPU and GPU by 456× and 9.5× respectively, providing practical performance at a fraction of the ASIC design cost. Rashmi S. Agrawal 0001, Leo de Castro, Guowei Yang 0005, Chiraag Juvekar, Rabia Tugce Yazicigil, Anantha P. Chandrakasan, Vinod Vaikuntanathan, Ajay Joshi |
HPCA | 6 |
| 2023 | A Fully-Integrated Energy-Scalable Transformer Accelerator Supporting Adaptive Model Configuration and Word Elimination for Language Understanding on Edge DevicesabstractEfficient natural language processing on the edge is needed to interpret voice commands, which have become a standard way to interact with devices around us. Due to the tight power and compute constraints of edge devices, it is important to adapt the computation to the hardware conditions. We present a Transformer accelerator with a variable-depth adder tree to support different model dimensions, a SuperTransformer model from which Sub Transformers of various sizes can be sampled enabling adaptive model configuration, and a dedicated word elimination unit to prune redundant tokens. We achieve up to 6.9× scalability in network latency and energy between the largest and smallest Sub Transformers, under the same operating conditions. Word elimination can reduce network energy by 16%, with a 14.5% drop in F1 score. At 0.68V and 80MHz, processing a 32-length input with our custom 2-layer Transformer model for intent detection and slot filling takes 0.61ms and 1.6μJ. Zexi Ji, Hanrui Wang 0002, Miaorong Wang, Win-San Khwa, Meng-Fan Chang, Song Han 0003, Anantha P. Chandrakasan |
ISLPED | 7 |
| 2023 | MAD: Memory-Aware Design Techniques for Accelerating Fully Homomorphic EncryptionabstractCloud computing has made it easier for individuals and companies to get access to large compute and memory resources. However, it has also raised privacy concerns about the data that users share with the remote cloud servers. Fully homomorphic encryption (FHE) offers a solution to this problem by enabling computations over encrypted data. Unfortunately, all known constructions of FHE require a noise term for security, and this noise grows during computation. To perform unlimited computations on the encrypted data, we need to perform a periodic noise reduction step known as bootstrapping. This bootstrapping operation is memory-bound as it requires several GBs of data. This leads to orders of magnitude increase in the time required for operating on encrypted data as compared to unencrypted data. Rashmi S. Agrawal 0001, Leo de Castro, Chiraag Juvekar, Anantha P. Chandrakasan, Vinod Vaikuntanathan, Ajay Joshi |
MICRO | 4 |
| 2023 | SecureLoop: Design Space Exploration of Secure DNN AcceleratorsabstractDeep neural networks (DNNs) are gaining popularity in a wide range of domains, ranging from speech and video recognition to healthcare. With this increased adoption comes the pressing need for securing DNN execution environments on CPUs, GPUs, and ASICs. While there are active research efforts in supporting a trusted execution environment (TEE) on CPUs, the exploration in supporting TEEs on accelerators is limited, with only a few solutions available [18, 19, 27]. A key limitation along this line of work is that these secure DNN accelerators narrowly consider a few specific architectures. The design choices and the associated cost for securing these architectures do not transfer to other diverse architectures. Kyungmi Lee, Mengjia Yan 0001, Joel S. Emer, Anantha P. Chandrakasan |
MICRO | 4 |
| 2022 | SparseBFA: Attacking Sparse Deep Neural Networks with the Worst-Case Bit Flips on CoordinatesabstractDeep neural networks (DNNs) are shown to be vulnerable to a few carefully chosen bit flips in their parameters, and bit flip attacks (BFAs) exploit such vulnerability to degrade the performance of DNNs. In this work, we show that DNNs with high sparsity that typically result from weight pruning have a unique source of vulnerability to bit flips when their coordinates of nonzero weights are attacked. We propose SparseBFA, an algorithm that searches for a small number of bits among the coordinates of nonzero weights when the parameters of DNNs are stored using sparse matrix formats. Using SparseBFA, we find that the performance of DNNs drops to the random-guess level by flipping less than 0.00005% (1 in 2 million) of the total bits. Kyungmi Lee, Anantha P. Chandrakasan |
ICASSP | 2 |
| 2022 | A Bit-level Sparsity-aware SAR ADC with Direct Hybrid Encoding for Signed Expressions for AIoT ApplicationsabstractIn this work, we propose the first bit-level sparsity-aware SAR ADC with direct hybrid encoding for signed expressions (HESE) for AIoT applications. ADCs are typically a bottleneck in reducing the energy consumption of analog neural networks (ANNs). For a pre-trained Convolutional Neural Network (CNN) inference, a HESE SAR for an ANN can reduce the number of non-zero signed digit terms to be output, and thus enables a reduction in energy along with the term quantization (TQ). The proposed SAR ADC directly produces the HESE signed-digit representation (SDR) using two thresholds per cycle for 2-bit look-ahead (LA). A prototype in 65nm shows that the HESE SAR provides sparsity encoding with a Walden FoM of 15.2fJ/conv.-step at 45MS/s. The core area is 0.072mm2. Ruicong Chen, H. T. Kung 0001, Anantha P. Chandrakasan, Hae-Seung Lee |
ISLPED | 3 |
| 2022 | Hardware Trojan Detection Using Unsupervised Deep Learning on Quantum Diamond Microscope Magnetic Field ImagesabstractThis article presents a method for hardware trojan detection in integrated circuits. Unsupervised deep learning is used to classify wide field-of-view (4 × 4 mm 2 ), high spatial resolution magnetic field images taken using a Quantum Diamond Microscope (QDM). QDM magnetic imaging is enhanced using quantum control techniques and improved diamond material to increase magnetic field sensitivity by a factor of 4 and measurement speed by a factor of 16 over previous demonstrations. These upgrades facilitate the first demonstration of QDM magnetic field measurement for hardware trojan detection. Unsupervised convolutional neural networks and clustering are used to infer trojan presence from unlabeled data sets of 600 × 600 pixel magnetic field images without human bias. This analysis is shown to be more accurate than principal component analysis for distinguishing between field programmable gate arrays configured with trojan-free and trojan-inserted logic. This framework is tested on a set of scalable trojans that we developed and measured with the QDM. Scalable and TrustHub trojans are detectable down to a minimum trojan trigger size of 0.5% of the total logic. The trojan detection framework can be used for golden-chip-free detection, since knowledge of the chips’ identities is only used to evaluate detection accuracy. Maitreyi Ashok, Ronald L. Walsworth, Edlyn V. Levine, Anantha P. Chandrakasan |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2021 | Securing Embedded Medical Devices using Dual-Factor AuthenticationabstractThis work provides an analysis of dual-factor authentication protocol for securing low-power medical devices. The dual-factor protocol incorporates voluntary physical action-based authentication in addition to traditional cryptographic methods for adding an extra layer of security. Along with touch signals, we propose and analyze the use of electromyographic (EMG) signals obtained from hand gestures for the second factor authentication. We demonstrate the feasibility of touch and EMG signals for dual-factor authentication by prototyping the medical device using off-the-shelf components. We also develop energy models for all these protocols, and analyze their overheads compared to traditional single-factor cryptographic authentication. Saurav Maji, Utsav Banerjee, Samuel H. Fuller, Rabia Tugce Yazicigil, Anantha P. Chandrakasan |
CBMS | 5 |
| 2021 | Leaky Nets: Recovering Embedded Neural Network Models and Inputs Through Simple Power and Timing Side-Channels - Attacks and DefensesabstractWith the recent advancements in machine learning theory, many commercial embedded microprocessors use neural network (NN) models for a variety of signal processing applications. However, their associated side-channel security vulnerabilities pose a major concern. There have been several proof-of-concept attacks demonstrating the extraction of their model parameters and input data. But, many of these attacks involve specific assumptions, have limited applicability, or pose huge overheads to the attacker. In this work, we study the side-channel vulnerabilities of embedded NN implementations by recovering their parameters using timing-based information leakage and simple power analysis side-channel attacks. We demonstrate our attacks on popular microcontroller platforms over networks of different precisions, such as floating point, fixed point, and binary networks. We are able to successfully recover not only the model parameters but also the inputs for the above networks. Countermeasures against timing-based attacks are implemented and their overheads are analyzed. Saurav Maji, Utsav Banerjee, Anantha P. Chandrakasan |
IEEE Internet Things J. | 3 |
| 2021 | Emerging Terahertz Integrated Systems in SiliconabstractSilicon-based terahertz (THz) integrated circuits (ICs) have made rapid progress over the past decade. The demonstrated basic component performance, as well as the maturity of design tools and methodologies, have made it possible to build high-complexity THz integrated systems. Such implementations are undoubtedly highly attractive due to their low cost and high integration capability; however, their unique characteristics, both advantageous and disadvantageous, also call for research investigations into unconventional systematic architectures and novel THz applications. In this paper, we review the current status and future trend of silicon-based THz ICs, with the focus on state-of-the-art THz microsystems for emerging sensing and communication applications in the last few years, such as high-resolution imaging, high medium/long-term stability time keeping, high-speed wireline/wireless communications, and miniaturization of RF tags, as well as THz packaging technologies. Cheng Wang 0009, Jack W. Holloway, Muhammad Ibrahim Wasiq Khan, Mohamed I. Ibrahim, Georgios C. Dogiamis, Bradford Perkins, Mehmet Kaynak, Rabia Tugce Yazicigil, Anantha P. Chandrakasan, Ruonan Han 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 12 |
| 2020 | Efficient Post-Quantum TLS Handshakes using Identity-Based Key Exchange from LatticesabstractIdentity-Based Encryption (IBE) is considered an alternative to traditional certificate-based public key cryptography to reduce communication overheads in wireless sensor networks. In this work, we build on the well-known lattice-based DLP-IBE scheme to construct an ID-based certificateless authenticated key exchange for post-quantum Transport Layer Security (TLS) handshakes. We also propose concrete parameters for the underlying lattice computations and provide detailed implementation results. Finally, we compare the combined computation and communication cost of our ID-based certificate-less handshake with the traditional certificate-based handshake, both using lattice-based algorithms at similar postquantum security levels, and show that our ID-based handshake is 3.7× more energy-efficient, thus highlighting the advantage of ID-based key exchange for post-quantum TLS. Utsav Banerjee, Anantha P. Chandrakasan |
ICC | 2 |
| 2020 | Accelerating Post-Quantum Cryptography using an Energy-Efficient TLS Crypto-ProcessorabstractPost-quantum cryptography (PQC) is currently a growing area of research and NIST PQC Round 2 schemes are being actively analyzed and optimized for both security and efficiency. In this work, we repurpose the cryptographic accelerators in an energy-efficient pre-quantum TLS crypto-processor to implement post-quantum key encapsulation schemes SIKE, Frodo and ThreeBears and signature scheme SPHINCS+. We utilize the modular arithmetic unit inside the elliptic curve cryptography accelerator to implement SIKE, while we use the AES-256 and SHA2-256 hardware primitives to substitute SHA3-256 and SHAKE-256 computations and accelerate the other three protocols. We accelerate the most computationally expensive components of these PQC protocols in hardware, thereby achieving up to an order of magnitude improvement in energy-efficiency over software implementations. Utsav Banerjee, Siddharth Das, Anantha P. Chandrakasan |
ISCAS | 3 |
| 2020 | Self-reconfigurable micro-implants for cross-tissue wireless and batteryless connectivityabstractWe present the design, implementation, and evaluation of μmedIC, a fully-integrated wireless and batteryless micro-implanted sensor. The sensor powers up by harvesting energy from RF signals and communicates at near-zero power via backscatter. In contrast to prior designs which cannot operate across various in-body environments, our sensor can self-reconfigure to adapt to different tissues and channel conditions. This adaptation is made possible by two key innovations: a reprogrammable antenna that can tune its energy harvesting resonance to surrounding tissues, and a backscatter rate adaptation protocol that closes the feedback loop by tracking circuit-level sensor hints. Mohamed R. Abdelhamid, Ruicong Chen, Joonhyuk Cho, Anantha P. Chandrakasan, Fadel Adib |
MobiCom | 4 |
| 2018 | A nonvolatile flip-flop-enabled cryptographic wireless authentication tag with per-query key update and power-glitch attack countermeasuresabstractCounterfeiting is a major issue plaguing global supply chains. To mitigate this issue, a wireless authentication tag is presented that implements a cryptographically secure pseudorandom number generator (PRNG) and authenticated encryption modes. The tag uses Keccak, the cryptographic core of SHA3, to update keys before each protocol invocation, limiting side-channel leakage. Power-glitch attacks are mitigated through state backup on ferroelectric capacitor-based nonvolatile flip-flops with a fully integrated energy backup storage, which needs a 2.2× smaller area compared with conventional approaches. The 130 nm CMOS tag harvests wireless power through a 433 MHz inductive link and communicates with a reader by a pulse-based modulation that minimizes the wireless power dead time. Full system operation including the tag, reader, and server protocol is demonstrated in the presence of worst-case power interruption events. Chiraag Juvekar, Anantha P. Chandrakasan, Joyce Kwong |
ASP-DAC | 2 |
| 2018 | Energy-Efficient Speaker Identification with Low-Precision NetworksabstractPower-consumption in small devices is dominated by off-chip memory accesses, necessitating small models that can fit in on-chip memory. In the task of text-dependent speaker identification, we demonstrate a 16× byte-size reduction for state-of-art small-footprint LCN/CNN/DNN speaker identification models. We achieve this by using ternary quantization that constrains the weights to {-1, 0, 1}. Our model comfortably fits in the 1 MB on-chip BRAM of most off-the-shelf FPGAs, allowing for a power-efficient speaker ID implementation with 100× fewer floating point multiplications, and a 1000× decrease in estimated energy cost. Additionally, we explore the use of depth-wise separable convolutions for speaker identification, and show while significantly reducing multiplications in full-precision networks, they perform poorly when ternarized. We simulate hardware designs for inference on our model, the first hardware design targeted for efficient evaluation of ternary networks and end-to-end neural network-based speaker identification. Skanda Koppula, James R. Glass, Anantha P. Chandrakasan |
ICASSP | 3 |
| 2018 | Recode then LSB-first SAR ADC for Reducing Energy and Bit-cyclesabstractLeast Significant Bit-first (LSB-first) algorithm is suitable for low-activity signals as it reduces DAC activity and the number of bit-cycles required per conversion. However, certain code transitions degrade performance by requiring large switching energy and number of bit-cycles even when the code change over previous code is small. This paper addresses it by a new algorithm called Recode then LSB-first (RLSB-first) that reduces the switching energy required for all cases of small code change across the full range of possible previous sample codes. The energy reduction is achieved while maintaining a low number of bit-cycles per conversion. Harneet Singh Khurana, Anantha P. Chandrakasan, Hae-Seung Lee |
ISCAS | 2 |
| 2018 | GAZELLE: A Low Latency Framework for Secure Neural Network Inference
Chiraag Juvekar, Vinod Vaikuntanathan, Anantha P. Chandrakasan |
USENIX Security Symposium | 3 |
| 2017 | Low-Power On-Chip Network Providing Guaranteed Services for Snoopy Coherent and Artificial Neural Network SystemsabstractDuring the transition to packet-switched on-chip networks we lose the relative timing and ordering of requests, which are essential for shared memory coherency and the communication of spikes in hardware-based artificial neural networks. We present a bufferless network architecture that enforces a time-based sharing of multi-hop single-cycle paths, providing guaranteed services at low cost. We guarantee ordered delivery of requests, fixed network latency, and jitter-free neural spikes. In a 64-node network, we achieve a 84% lower latency and 7.5x higher throughput than SCORPIO. Full-system 36-core simulations show a 9% lower runtime than SCORPIO, with 39% lower power and 36% lower area. Bhavya K. Daya, Li-Shiuan Peh, Anantha P. Chandrakasan |
DAC | 3 |
| 2017 | eeDTLS: Energy-Efficient Datagram Transport Layer Security for the Internet of ThingsabstractIn the fast growing world of the Internet of Things (IoT), security has become a major concern. Datagram Transport Layer Security (DTLS) is considered to be one of the most suited protocols for securing the IoT. However, computation and communication overheads make it very expensive to implement DTLS on resource-constrained IoT sensor nodes. In this work, we profile the energy costs of DTLS 1.3, using experimental models for cryptographic computations and radio-frequency (RF) communications. Based on this analysis, we present eeDTLS, a low-energy variant of DTLS, that provides the same security strength as DTLS, but has lower energy requirements. By employing a combination of packet size reduction and optimized handshake computations, eeDTLS can provide up to 45% energy savings in a typical IoT use case. eeDTLS can be implemented in conjunction with any low-energy IoT RF protocol, and the proposed energy models and protocol optimizations can also be used to improve the energy efficiency of custom IoT security architectures. Utsav Banerjee, Chiraag Juvekar, Samuel H. Fuller, Anantha P. Chandrakasan |
GLOBECOM | 4 |
| 2017 | Harnessing Partial Packets in Wireless Networks: Throughput and Energy BenefitsabstractThis paper proposes a partial packet recovery scheme called packetized rateless algebraic consistency (PRAC). PRAC exploits intra- and inter-packet consistency to identify and recover erroneous packet segments, without recourse to soft physical layer (PHY) or detailed feedback information. PRAC uses a rateless linear code for data encoding and an iterative decoding process for data reconstruction. It allows, but does not rely upon, the use of any PHY forward error correction code, and requires no feedback other than a notification of completion and, in the absence of partial packets, incurs no overhead. In order to quantify PRAC's performance in terms of both throughput and energy efficiency, experiments are conducted using commercial transceivers in two different scenarios. Our implementation results reveal that PRAC offers an average throughput gain of 35% compared with a baseline ARQ scheme discarding partial packets, and 13% compared with an ideal hybrid-ARQ scheme. On high PER links, throughput is improved by 148% and 34%, respectively. In addition, PRAC reduces on average the total energy consumption of the transmitting nodes by 16%, while, on high PER links, savings can be up to 50%. Georgios Angelopoulos, Muriel Médard, Anantha P. Chandrakasan |
IEEE Trans. Wirel. Commun. | 3 |
| 2016 | Quest for high-performance bufferless NoCs with single-cycle express paths and self-learning throttlingabstractRouter buffers are the main reason for the Network-on-Chip's (NoC) scalable bandwidth, but consumes significant area and power. The SCEPTER bufferless NoC sets up single-cycle virtual express paths dynamically, allowing packets to traverse non-minimal paths without latency penalty. Using prioritization, bypassing, and throttling mechanisms, we maximize opportunities to use these paths while pushing bandwidth. For 64 and 256 nodes, we achieve 62% lower latency, 1.3x higher throughput, and 35% lower starvation over a baseline bufferless NoC for synthetic traffic. Full-system 36-core simulations show a 19% lower runtime, on-par performance to a buffered network, with 36% lower area, 33% lower power. Bhavya K. Daya, Li-Shiuan Peh, Anantha P. Chandrakasan |
DAC | 3 |
| 2016 | Enabling simultaneously bi-directional TSV signaling for energy and area efficient 3D-ICs
Sunghyun Park 0002, Alice Wang 0002, Uming Ko, Li-Shiuan Peh, Anantha P. Chandrakasan |
DATE | 5 |
| 2016 | Memory-Efficient Modeling and Search Techniques for Hardware ASR Decoders
Michael Price 0001, Anantha P. Chandrakasan, James R. Glass |
INTERSPEECH | 2 |
| 2015 | AdaptCast: An integrated source to transmission scheme for wireless sensor networksabstractThis paper introduces AdaptCast, an integrated source to transmission scheme for wireless sensor networks (WSNs) that efficiently represents collected data and increases their robustness against channel errors across a wide range of signal to noise (SNR) values in a rateless fashion. AdaptCast leverages sparsity inherent in the majority of physical signals in order to parsimoniously represent them without relying on a specific signal model. The proposed scheme does not suffer from the sudden degradation in the tradeoff between distortion and SNR of rated channel coding schemes due to its direct, relative bit importance preserving modulation mapping. In addition, it does not require continuous feedback or channel state information (CSI) as a result of its rateless operation. Apart from point-to-point transmission, AdaptCast enables efficient multicasting to a set of nodes, serving each of them at a rate commensurate to its individual channel quality. We demonstrate AdaptCast's application-independent operation by using several typical signals captured in WSNs. Based on our analysis and simulation results, considering the tradeoff between distortion and channel quality, AdaptCast performs close in a point-to-point scenario to an idealized layered transmission scheme with instantaneous CSI and offers significant benefits in multiuser settings. Georgios Angelopoulos, Muriel Médard, Anantha P. Chandrakasan |
ICC | 3 |
| 2015 | Caraoke: An E-Toll Transponder Network for Smart CitiesabstractElectronic toll collection transponders, e.g., E-ZPass, are a widely-used wireless technology. About 70% to 89% of the cars in US have these devices, and some states plan to make them mandatory. As wireless devices however, they lack a basic function: a MAC protocol that prevents collisions. Hence, today, they can be queried only with directional antennas in isolated spots. However, if one could interact with e-toll transponders anywhere in the city despite collisions, it would enable many smart applications. For example, the city can query the transponders to estimate the vehicle flow at every intersection. It can also localize the cars using their wireless signals, and detect those that run a red-light. The same infrastructure can also deliver smart street-parking, where a user parks anywhere on the street, the city localizes his car, and automatically charges his account. This paper presents Caraoke, a networked system for delivering smart services using e-toll transponders. Our design operates with existing unmodified transponders, allowing for applications that communicate with, localize, and count transponders, despite wireless collisions. To do so, Caraoke exploits the structure of the transponders' signal and its properties in the frequency domain. We built Caraoke reader into a small PCB that harvests solar energy and can be easily deployed on street lamps. We also evaluated Caraoke on four streets on our campus and demonstrated its capabilities. Omid Abari, Deepak Vasisht, Dina Katabi, Anantha P. Chandrakasan |
SIGCOMM | 4 |
| 2014 | SCORPIO: 36-core shared memory processor demonstrating snoopy coherence on a mesh interconnectabstractThis article consists of a collection of slides from the author's conference presentation on the special features, system design and architectures, processing capabilities, and targeted markets for Freescale Inc.'s SCORPIO, a 36-core shared memory processor. Chia-Hsin Owen Chen, Sunghyun Park 0002, Suvinay Subramanian, Tushar Krishna, Bhavya K. Daya, Woo-Cheol Kwon, Brett Wilkerson, John Arends, Anantha P. Chandrakasan, Li-Shiuan Peh |
Hot Chips Symposium | 9 |
| 2014 | PRAC: Exploiting partial packets without cross-layer or feedback informationabstractThis paper proposes a partial packet recovery scheme, called Packetized Rateless Algebraic Consistency (PRAC). PRAC exploits intra and inter-packet consistency to identify and recover erroneous packet segments, without recourse to cross-layer or detailed feedback information. In the absence of cross-layer coordination or detailed feedback, the prevailing methods proposed in the literature have discarded packets with errors. PRAC uses a rateless linear packet code for data encoding and an iterative decoding process consisting of a search algorithm and an algebraic consistency rule (ACR) check. It allows, but not relies upon, the use of any PHY FEC code, requires no feedback other than a notification of completion and, in the absence of partial packets, incurs no overhead. Our implementation and experimental results in a 7-node indoor testbed using wireless boards equipped with CC2500 radio transceivers reveal that PRAC offers an average throughput gain of 35% compared to a baseline ARQ scheme discarding partial packets and 13% compared to an ideal genie-aided HARQ (iHARQ) scheme. Specifically for links with high PERs, PRAC significantly enhances their robustness and its maximum throughput gain is 148% and 34% compared against the baseline and iHARQ schemes, respectively. Georgios Angelopoulos, Anantha P. Chandrakasan, Muriel Médard |
ICC | 2 |
| 2014 | Energy and area-efficient hardware implementation of HEVC inverse transform and dequantizationabstractHigh Efficiency Video Coding (HEVC) inverse transform for residual coding uses 2-D 4×4 to 32×32 transforms with higher precision as compared to H.264/AVC's 4×4 and 8×8 transforms resulting in an increased hardware complexity. In this paper, an energy and area-efficient VLSI architecture of an HEVC-compliant inverse transform and dequantization engine is presented. We implement a pipelining scheme to process all transform sizes at a minimum throughput of 2 pixel/cycle with zero-column skipping for improved throughput. We use data-gating in the 1-D Inverse Discrete Cosine Transform engine to improve energy-efficiency for smaller transform sizes. A high-density SRAM-based transpose memory is used for an area-efficient design. This design supports decoding of 4K Ultra-HD (3840×2160) video at 30 frame/sec. The inverse transform engine takes 98.1 kgate logic, 16.4 kbit SRAM and 10.82 pJ/pixel while the dequantization engine takes 27.7 kgate logic, 8.2 kbit SRAM and 1.10 pJ/pixel in 40 nm CMOS technology. Although larger transforms require more computation per coefficient, they typically contain a smaller proportion of non-zero coefficients. Due to this trade-off, larger transforms can be more energy-efficient. Mehul Tikekar, Chao-Tsung Huang, Vivienne Sze, Anantha P. Chandrakasan |
ICIP | 4 |
| 2014 | SCORPIO: A 36-core research chip demonstrating snoopy coherence on a scalable mesh NoC with in-network orderingabstractIn the many-core era, scalable coherence and on-chip interconnects are crucial for shared memory processors. While snoopy coherence is common in small multicore systems, directory-based coherence is the de facto choice for scalability to many cores, as snoopy relies on ordered interconnects which do not scale. However, directory-based coherence does not scale beyond tens of cores due to excessive directory area overhead or inaccurate sharer tracking. Prior techniques supporting ordering on arbitrary unordered networks are impractical for full multicore chip designs. Bhavya K. Daya, Chia-Hsin Owen Chen, Suvinay Subramanian, Woo-Cheol Kwon, Sunghyun Park 0002, Tushar Krishna, Jim Holt, Anantha P. Chandrakasan, Li-Shiuan Peh |
ISCA | 8 |
| 2014 | A bipolar ±40 MV self-starting boost converter with transformer reuse for thermoelectric energy harvestingabstractThis paper presents a converter for boosting the low-voltage output of thermoelectric energy harvesters to power standard CMOS circuits. The converter can start up from a fully de-energized state off a bipolar ±40 mV input and can harvest net positive energy from voltages as low as ±30 mV in steady state. A single transformer is multiplexed between an oscillator that is used during startup and a flyback converter that is used during steady-state operation. During steady-state operation, the converter is automatically shut off if the input power is found to be too low. Simulation results on the converter designed in a 0.35 μm CMOS process demonstrate a peak steady-state conversion efficiency of 68% at an output voltage of 5.5 V and input voltage range between 30 mV and 500 mV in magnitude. Nachiket V. Desai, Yogesh K. Ramadass, Anantha P. Chandrakasan |
ISLPED | 3 |
| 2014 | Memory-Hierarchical and Mode-Adaptive HEVC Intra Prediction Architecture for Quad Full HD Video DecodingabstractThis paper presents a high-throughput and areaefficient VLSI architecture for intra prediction in the emerging high efficiency video coding standard. Three design techniques are proposed to address the complexity systematically: 1) a hierarchical memory deployment that stores neighboring samples in 4.9 Kb of static RAM (SRAM) instead of 43.2-k gates of registers and increases throughput by processing reference samples in registers; 2) a mode-adaptive scheduling scheme for all prediction units, which provides at least 2 samples/cycle throughput while using low-throughput SRAM and can achieve 2.46 samples/cycle on the average based on the experimental results; and 3) resource sharing for multipliers and the readout circuits of reference sample registers, which can save 2.5-k gates. These techniques can efficiently reduce area by 40% but induce more power because of additional signal transitions. Signal-gating circuits are then applied to reduce 69% of SRAM power and 32% of logic power, which cost only 1.0-k gates. When synthesized at 200 MHz with 40-nm process, the proposed architecture needs only 27.0-k gates and 4.9 Kb of single-port SRAM. The layout core area is 0.036 mm2, and the power consumption is 2.11 mW in the postlayout simulation. The corresponding performance can support quad full high-definition (HD) (3840 × 2160) video decoding at 30 frames/s. Chao-Tsung Huang, Mehul Tikekar, Anantha P. Chandrakasan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | SMART: a single-cycle reconfigurable NoC for SoC applicationsabstractAs technology scales, SoCs are increasing in core counts, leading to the need for scalable NoCs to interconnect the multiple cores on the chip. Given aggressive SoC design targets, NoCs have to deliver low latency, high bandwidth, at low power and area overheads. In this paper, we propose Single-cycle Multi-hop Asynchronous Repeated Traversal (SMART) NoC, a NoC that reconfigures and tailors a generic mesh topology for SoC applications at runtime. The heart of our SMART NoC is a novel low-swing clockless repeated link circuit embedded within the router crossbars, that allows packets to potentially bypass all the way from source to destination core within a single clock cycle, without being latched at any intermediate router. Our clockless repeater link has been proven in silicon in 45nm SOI. Results show that at 2GHz, we can traverse 8mm within a single cycle, i.e. 8 hops with 1mm cores. We implement the SMART NoC to layout and show that SMART NoC gives 60% latency savings, and 2.2X power savings compared to a baseline mesh NoC. Chia-Hsin Owen Chen, Sunghyun Park 0002, Tushar Krishna, Suvinay Subramanian, Anantha P. Chandrakasan, Li-Shiuan Peh |
DATE | 5 |
| 2013 | 40.4fJ/bit/mm low-swing on-chip signaling with self-resetting logic repeaters embedded within a mesh NoC in 45nm SOI CMOSabstractMesh NoCs are the most widely-used fabric in high-performance many-core chips today. They are, however, becoming increasingly power-constrained with the higher on-chip bandwidth requirements of high-performance SoCs. In particular, the physical datapath of a mesh NoC consumes significant energy. Low-swing signaling circuit techniques can substantially reduce the NoC datapath energy, but existing low-swing circuits involve huge area footprints, unreliable signaling or considerable system overheads such as an additional supply voltage, so embedding them into a mesh datapath is not attractive. In this paper, we propose a novel low-swing signaling circuit, a self-resetting logic repeater, to meet these design challenges. The SRLR enables single-ended low-swing pulses to be asynchronously repeated, and hence, consumes less energy than differential, clocked low-swing signaling. To mitigate global process variations while delivering high energy efficiency, three circuit techniques are incorporated. Fabricated in 45nm SOI CMOS, our 10mm SRLR-based low-swing datapath achieves 6.83Gb/s/µm bandwidth density with 40.4fJ/bit/mm energy at 4.1Gb/s data rate at 0.8V. Sunghyun Park 0002, Masood Qazi, Li-Shiuan Peh, Anantha P. Chandrakasan |
DATE | 4 |
| 2013 | Experimental study of the interplay of channel and network coding in low power sensor applicationsabstractIn this paper, we evaluate the performance of random linear network coding (RLNC) in low data rate indoor sensor applications operating in the ISM frequency band. We also investigate the results of its synergy with forward error correction (FEC) codes at the PHY-layer in a joint channel-network coding (JCNC) scheme. RLNC is an emerging coding technique which can be used as a packet-level erasure code, usually implemented at the network layer, which increases data reliability against channel fading and severe interference, while FEC codes are mainly used for correction of random bit errors within a received packet. The hostile wireless environment that low power sensors usually operate in, with significant interference from nearby networks, motivates us to consider a joint coding scheme and examine the applicability of RLNC as an erasure code in such a coding structure. Our analysis and experiments are performed using a custom low power sensor node, which integrates on-chip a low-power 2.4 GHz transmitter and an accelerator implementing a multi-rate convolutional code and RLNC, in a typical office environment. According to measurement results, RLNC of code rate 4/8 can provide an effective SNR improvement of about 3.4 dB, outperforming a PHY-layer FEC code of the same code rate, at a PER of 10-2. In addition, RLNC performs very well when used in conjunction with a PHY-layer FEC code as a JCNC scheme, offering an overall coding gain of 5.6 dB. Georgios Angelopoulos, Arun Paidimarri, Anantha P. Chandrakasan, Muriel Médard |
ICC | 3 |
| 2013 | HEVC interpolation filter architecture for quad full HD decodingabstractIn this paper, an area-efficient and high-throughput interpolation filter architecture is presented for the latest video coding standard, High Efficiency Video Coding. A unified filter design is first proposed for the 8-tap luma and 4-tap chroma filters to optimize area, which uses only 13 adders. And a 2D filter architecture is then devised with an adaptive scheduling which supports all symmetric prediction partitions with a throughput of at least two samples/cycle. Experimental results also show that this architecture can achieve 2.58 samples/cycle on the average. The total gate count is 45.2k when synthesized at 200MHz with 40nm process, and the corresponding performance can support at least 3840×2160 videos at 30 fps. Chao-Tsung Huang, Chiraag Juvekar, Mehul Tikekar, Anantha P. Chandrakasan |
VCIP | 4 |
| 2013 | Technique for Efficient Evaluation of SRAM Timing FailureabstractThis brief presents a technique to evaluate the timing variation of static random access memory (SRAM). Specifically, a method called loop flattening, which reduces the evaluation of the timing statistics in the complex highly structured circuit to that of a single chain of component circuits, is justified. Then, to very quickly evaluate the timing delay of a single chain, a statistical method based on importance sampling augmented with targeted high-dimensional spherical sampling can be employed. The overall methodology has shown 650× or greater speedup over the nominal Monte Carlo approach with 10.5% accuracy in probability. Examples based on both the large-signal and small-signal SRAM read path are discussed, and a detailed comparison with state-of-the-art accelerated statistical simulation techniques is given. Masood Qazi, Mehul Tikekar, Lara Dolecek, Devavrat Shah, Anantha P. Chandrakasan |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2012 | Self-aware computing in the Angstrom processorabstractAddressing the challenges of extreme scale computing requires holistic design of new programming models and systems that support those models. This paper discusses the Angstrom processor, which is designed to support a new Self-aware Computing (SEEC) model. In SEEC, applications explicitly state goals, while other systems components provide actions that the SEEC runtime system can use to meet those goals. Angstrom supports this model by exposing sensors and adaptations that traditionally would be managed independently by hardware. This exposure allows SEEC to coordinate hardware actions with actions specified by other parts of the system, and allows the SEEC runtime system to meet application goals while reducing costs (e.g., power consumption). Henry Hoffmann, Jim Holt, George Kurian, Eric Lau, Martina Maggio, Jason E. Miller, Sabrina M. Neuman, Mahmut E. Sinangil, Yildiz Sinangil, Anant Agarwal, Anantha P. Chandrakasan, Srini Devadas |
DAC | 11 |
| 2012 | Approaching the theoretical limits of a mesh NoC with a 16-node chip prototype in 45nm SOIabstractIn this paper, we present a case study of our chip prototype of a 16-node 4x4 mesh NoC fabricated in 45nm SOI CMOS that aims to simultaneously optimize energy-latency-throughput for unicasts, multicasts and broadcasts. We first define and analyze the theoretical limits of a mesh NoC in latency, throughput and energy, then describe how we approach these limits through a combination of microarchitecture and circuit techniques. Our 1.1V 1GHz NoC chip achieves 1-cycle router-and-link latency at each hop and energy-efficient router-level multicast support, delivering 892Gb/s (87.1% of the theoretical bandwidth limit) at 531.4mW for a mixed traffic of unicasts and broadcasts. Through this fabrication, we derive insights that help guide our research, and we believe, will also be useful to the NoC and multicore research community. Sunghyun Park 0002, Tushar Krishna, Chia-Hsin Owen Chen, Bhavya K. Daya, Anantha P. Chandrakasan, Li-Shiuan Peh |
DAC | 5 |
| 2012 | Hardware-aware motion estimation search algorithm development for high-efficiency video coding (HEVC) standardabstractThis work presents a hardware-aware search algorithm for HEVC motion estimation. Implications of several decisions in search algorithm are considered with respect to their hardware implementation costs (in terms of area and bandwidth). Proposed algorithm provides 3X logic area in integer motion estimation, 16% on-chip reference buffer area and 47X maximum off-chip bandwidth savings when compared to HM-3.0 fast search algorithm. Mahmut E. Sinangil, Anantha P. Chandrakasan, Vivienne Sze, Minhua Zhou |
ICIP | 2 |
| 2012 | Memory cost vs. coding efficiency trade-offs for HEVC motion estimation engineabstractThis paper presents a comparison between various High Efficiency Video Coding (HEVC) motion estimation configurations in terms of coding efficiency and memory cost in hardware. An HEVC motion estimation hardware model that is suitable to implement HEVC reference software (HM) search algorithm is created and memory area and data bandwidth requirements are calculated based on this model. 11 different motion estimation configurations are considered. Supporting smaller block sizes is shown to impose significant memory cost in hardware although the coding gain achieved through supporting them is relatively smaller. Hence, depending on target encoder specifications, the decision can be made not to support certain block sizes. Specifically, supporting only 64x64, 32x32 and 16x16 block sizes provide 3.2X on-chip memory area, 26X on-chip bandwidth and 12.5X off-chip bandwidth savings at the expense of 12% bit-rate increase when compared to the anchor configuration supporting all block sizes. Mahmut E. Sinangil, Anantha P. Chandrakasan, Vivienne Sze, Minhua Zhou |
ICIP | 2 |
| 2012 | The Effect of Random Dopant Fluctuations on Logic Timing at Low VoltageabstractIn order to achieve ultra-low power (ULP), ICs are being designed for VDD≤ 0.5 V. At these low voltages, random dopant fluctuations (RDFs) result in a stochastic component of logic delay that can be comparable to the global corner delay. Moreover, the probability density function (PDF) of this stochastic delay can be highly non-Gaussian. In order to predict the statistical impact of RDF-induced local variations on logic timing, it is necessary to incorporate these effects into a timing closure methodology. This paper presents a computationally efficient methodology for stochastic characterization of standard cell li- braries at low voltage, where the cell delay is a nonlinear function of the transistor random variables (RVs), and the resulting cell delay has a non-Gaussian PDF. It also presents a computation- ally efficient methodology for computing any point on the PDF of a timing path (TP) delay, in the case where cell delays are non-Gaussian. The method is called nonlinear operating point analysis of local variation (NLOPALV). The general NLOPALV theory is developed. It is applied to cell library characterization, and the accuracy of the NLOPALV approach is validated by comparison to Monte Carlo simulation. NLOPALV is also applied to timing path analysis on a 28 nm DSP IC. The approach has been implemented using commercial CAD tools, and integrated into a commercial IC design flow. The NLOPALV approach gives timing results that are within 5% accuracy compared to Monte Carlo analysis at VDD= 0.5 V. This compares to errors on the order of 50% when the Gaussian approximation is used. Rahul Rithe, Sharon Chou, Jie Gu 0007, Alice Wang 0002, Satyendra Datla, Gordon Gammie, Anantha P. Chandrakasan |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2011 | Joint algorithm-architecture optimization of CABAC to increase speed and reduce area costabstractTo address the increasing demand for higher resolution and frame rates, processing speed (i.e. performance) and area cost need to be considered in the development of next generation video coding. Accordingly, both algorithm and architecture should be taken into account during video codec design. This paper proposes joint optimization of both the algorithm and architecture to ensure that high coding efficiency can be achieved in conjunction with high processing speed and low area cost. Specifically, it presents two optimizations that can be performed on Context-based Adaptive Binary Arithmetic Coding (CABAC), a form of entropy coding in H.264/AVC. First, subinterval reordering is proposed for the arithmetic de coder to increase the processing speed by 14 to 22% with no cost to coding efficiency. Second, modification of the motion vector difference (mvd) context selection is proposed to reduce memory requirements (i.e. area cost) by 50% with negligible coding efficiency impact (≤0.02%). These joint algorithm and architecture optimizations are non-standard com pliant and thus are well suited to be used in High Efficiency Video Coding (HEVC), the successor to H.264/AVC. Vivienne Sze, Anantha P. Chandrakasan |
ICASSP | 2 |
| 2011 | Reduction of Variation-Induced Energy Overhead in Multi-Core ProcessorsabstractCore-to-core variability in future many-core chip multi-processors (CMPs) negatively impacts energy. Under-performing cores necessitate increasing the system voltage to maintain homogeneous core performance, introducing an energy overhead. Multiple supply voltages can be used to mitigate the impact of delay variation in CMPs. In this paper, we carefully analyze the use of a local search algorithm to pick near-optimal supply voltages while meeting a fixed performance target. With two system voltages, we prove our algorithm selects the global optimum and in the more general multiple voltage case we develop quantitative bounds. Using a custom simulation methodology on a real processor core, we show that two system voltages provide the most incremental benefit, reducing the energy overhead relative to a single voltage by 59-75% and total energy by 6-16%. Additionally, the worst 5-15% of cores in such systems necessitate increasingly larger amounts of incremental energy for a constant incremental performance gain. Therefore, turning off or disabling these cores is beneficial to a joint performance-energy metric. Nigel Drego, Anantha P. Chandrakasan, Duane S. Boning, Devavrat Shah |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Loop flattening & spherical sampling: Highly efficient model reduction techniques for SRAM yield analysisabstractThe impact of process variation in deep-submicron technologies is especially pronounced for SRAM architectures which must meet demands for higher density and higher performance at increased levels of integration. Due to the complex structure of SRAM, estimating the effect of process variation accurately has become very challenging. In this paper, we address this challenge in the context of estimating SRAM timing variation. Specifically, we introduce a method called loop flattening that demonstrates how the evaluation of the timing statistics in the complex, highly structured circuit can be reduced to that of a single chain of component circuits. To then very quickly evaluate the timing delay of a single chain, we employ a statistical method based on importance sampling augmented with targeted, high-dimensional, spherical sampling. Overall, our methodology provides an accurate estimation with 650X or greater speed-up over the nominal Monte Carlo approach. Masood Qazi, Mehul Tikekar, Lara Dolecek, Devavrat Shah, Anantha P. Chandrakasan |
DATE | 5 |
| 2010 | Non-linear Operating Point Statistical Analysis for Local Variations in logic timing at low voltageabstractFor CMOS feature size of 65 nm and below, local (or intra-die or within-die) variations in transistor Vt contribute stochastic variation in logic delay that is a large percentage of the nominal delay. Moreover, when circuits are operated at low voltage (Vdd ¿ 0.5 V), the standard deviation of gate delay becomes comparable to nominal delay, and the Probability Density Function (PDF) of the gate delay is highly non-Gaussian. This paper presents a computationally efficient algorithm for computing the PDF of logic Timing Path (TP) delay, which results from local variations. This approach is called Non-linear Operating Point Analysis for Local Variations (NLOPALV). The approach is implemented using commercial STA tools and integrated into the standard CAD flow using custom scripts. Timing paths from a 28 nm commercial DSP are analyzed using the proposed technique and the performance is observed to be within 5% accuracy compared to SPICE based Monte-Carlo analysis. Rahul Rithe, Jie Gu 0007, Alice Wang 0002, Satyendra Datla, Gordon Gammie, Anantha P. Chandrakasan |
DATE | 7 |
| 2010 | Technologies for Ultradynamic Voltage ScalingabstractEnergy efficiency of electronic circuits is a critical concern in a wide range of applications from mobile multi-media to biomedical monitoring. An added challenge is that many of these applications have dynamic workloads. To reduce the energy consumption under these variable computation requirements, the underlying circuits must function efficiently over a wide range of supply voltages. This paper presents voltage-scalable circuits such as logic cells, SRAMs, ADCs, and dc-dc converters. Using these circuits as building blocks, two different applications are highlighted. First, we describe an H.264/AVC video decoder that efficiently scales between QCIF and 1080p resolutions, using a supply voltage varying from 0.5 V to 0.85 V. Second, we describe a 0.3 V 16-bit micro-controller with on-chip SRAM, where the supply voltage is generated efficiently by an integrated dc-dc converter. Anantha P. Chandrakasan, Denis C. Daly, Daniel F. Finchelstein, Joyce Kwong, Yogesh K. Ramadass, Mahmut E. Sinangil, Vivienne Sze, Naveen Verma |
Proc. IEEE | 1 |
| 2009 | A high throughput CABAC algorithm using syntax element partitioningabstractEnabling parallel processing is becoming increasingly necessary for video decoding as performance requirements continue to rise due to growing resolution and frame rate demands. It is important to address known bottlenecks in the video decoder such as entropy decoding, specifically the highly serial Context-based Adaptive Binary Arithmetic Coding (CABAC) algorithm. Concurrency must be enabled with minimal cost to coding efficiency, power, area and delay. This work proposes a new CABAC algorithm for the next generation standard in which binary symbols are grouped by syntax elements and assigned to different partitions which can be decoded in parallel. Furthermore, since the distribution of binary symbols changes with quantization, an adaptive binary symbol allocation scheme is proposed to maximize throughput. Application of this next generation CABAC algorithm on five 720p sequences shows a throughput increase of up to 3x can be achieved with negligible impact on coding efficiency (0.06% to 0.37%), which is a 2 to 4x reduction in coding penalty compared with H.264/AVC and entropy slices. Area cost is also reduced by 2x. This increased throughput can be traded-off for low power consumption in mobile applications. Vivienne Sze, Anantha P. Chandrakasan |
ICIP | 2 |
| 2009 | Low-Power Impulse UWB Architectures and CircuitsabstractUltra-wide-band (UWB) communication has a variety of applications ranging from wireless USB to radio frequency (RF) identification tags. For many of these applications, energy is critical due to the fact that the radios are situated on battery-operated or even batteryless devices. Two custom low-power impulse UWB systems are presented in this paper that address high- and low-data-rate applications. Both systems utilize energy-efficient architectures and circuits. The high-rate system leverages parallelism to enable the use of energy-efficient architectures and aggressive voltage scaling down to 0.4 V while maintaining a rate of 100 Mb/s. The low-rate system has an all digital transmitter architecture, 0.65 and 0.5 V radio-frequency (RF) and analog circuits in the receiver, and no RF local oscillators, allowing the chipset to power on in 2 ns for highly duty-cycled operation. Anantha P. Chandrakasan, Fred S. Lee, David D. Wentzloff, Vivienne Sze, Brian P. Ginsburg, Patrick P. Mercier, Denis C. Daly, Raúl Blázquez |
Proc. IEEE | 1 |
| 2009 | Multicore Processing and Efficient On-Chip Caching for H.264 and Future Video DecodersabstractPerformance requirements for video decoding will continue to rise in the future due to the adoption of higher resolutions and faster frame rates. Multicore processing is an effective way to handle the resulting increase in computation. For power-constrained applications such as mobile devices, extra performance can be traded-off for lower power consumption via voltage scaling. As memory power is a significant part of system power, it is also important to reduce unnecessary on-chip and off-chip memory accesses. This paper proposes several techniques that enable multiple parallel decoders to process a single video sequence; the paper also demonstrates several on-chip caching schemes. First, we describe techniques that can be applied to the existing H.264 standard, such as multiframe processing. Second, with an eye toward future video standards, we propose replacing the traditional raster-scan processing with an interleaved macroblock ordering; this can increase parallelism with minimal impact on coding efficiency and latency. The proposed architectures allowNparallel hardware decoders to achieve a speedup of up to a factor ofN. For example, ifN=3, the proposed multiple frame and interleaved entropy slice multicore processing techniques can achieve performance improvements of 2.64times and 2.91times, respectively. This extra hardware performance can be used to decode higher definition videos. Alternatively, it can be traded-off for dynamic power savings of 60% relative to a single nominal-voltage decoder. Finally, on-chip caching methods are presented that significantly reduce off-chip memory bandwidth, leading to a further increase in performance and energy efficiency. Data-forwarding caches can reduce off-chip memory reads by 53%, while using a last-frame cache can eliminate 80% of the off-chip reads. The proposed techniques were validated and benchmarked using full-system Verilog hardware simulations based on an existing decoder; they should also be applicable to most other decoder architectures. The metrics used to evaluate the ideas in this paper are performance, power, area, memory efficiency, coding efficiency, and input latency. Daniel F. Finchelstein, Vivienne Sze, Anantha P. Chandrakasan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | The design of a low power carbon nanotube chemical sensor systemabstractThis paper presents a hybrid CNT/CMOS chemical sensor system that comprises of a carbon nanotube sensor array and a CMOS interface chip. The full system, including the sensor, consumes 32 muW at 1.83 kS/s readout rate, accomplished through an extensive use of CAD tools and a model-based architecture optimization. A redundant use of CNT sensors in the frontend increases the reliability of the system. Taeg Sang Cho, Kyeong-Jae Lee, Anantha P. Chandrakasan |
DAC | 4 |
| 2008 | The mixed signal optimum energy point: voltage and parallelismabstractAn energy optimization is proposed that addresses the non-trivial digital contribution to power and impact on performance in high-speed mixed-signal circuits. Parallel energy and behavioral models are used to quantify architectural tradeoffs across the analog/digital boundary. An interleaved ADC is optimized as a case study to demonstrate this approach. The chosen operating point of 36 channels and 700mV operation gives a 3× improvement in energy compared to the seed of the model. The model matches closely the measured results of an ADC testchip implemented in a 65nm CMOS process. Brian P. Ginsburg, Anantha P. Chandrakasan |
DAC | 2 |
| 2008 | Breaking the simulation barrier: SRAM evaluation through norm minimizationabstractWith process variation becoming a growing concern in deep submicron technologies, the ability to efficiently obtain an accurate estimate of failure probability of SRAM components is becoming a central issue. In this paper we present a general methodology for a fast and accurate evaluation of the failure probability of memory designs. The proposed statistical method, which we call importance sampling through norm minimization principle, reduces the variance of the estimator to produce quick estimates. It builds upon the importance sampling, while using a novel norm minimization principle inspired by the classical theory of Large Deviations. Our method can be applied for a wide class of problems, and our illustrative examples are the data retention voltage and the read/write failure tradeoff for 6T SRAM in 32 nm technology. The method yields computational savings on the order of 10000x over the standard Monte Carlo approach in the context of failure probability estimation for SRAM considered in this paper. Lara Dolecek, Masood Qazi, Devavrat Shah, Anantha P. Chandrakasan |
ICCAD | 4 |
| 2008 | Parallel CABAC for low power video codingabstractWith the growing presence of high definition video content on battery-operated handheld devices such as camera phones, digital still cameras, digital camcorders, and personal media players, it is becoming ever more important that video compression be power efficient. A popular form of entropy coding called Context-Based Adaptive Binary Arithmetic Coding (CABAC) provides high coding efficiency but has limited throughput. This can lead to high operating frequencies resulting in high power dissipation. This paper presents a novel parallel CABAC scheme which enables a throughput increase of N-fold (depending on the degree parallelism), reducing the frequency requirement and expected power consumption of the coding engine. Experiments show that this new scheme (with N=2) can deliver ∼2x throughput improvement at a cost of 0.76% average increase in bit-rate or equivalently a decrease in average PSNR of 0.025dB on five 720p resolution video clips when compared with H.264/AVC. Vivienne Sze, Anantha P. Chandrakasan, Madhukar Budagavi, Minhua Zhou |
ICIP | 2 |
| 2008 | Ultra-low-power UWB for sensor network applicationsabstractLong distance, low data-rate UWB communication for sensor network applications requires a highly energy efficient transceiver combined with circuit and system-level optimizations to maximize range. A custom pulsed-UWB transceiver chipset in 90 nm CMOS is presented that targets these aggressive specifications. The transceiver efficiently communicates at data rates from 0-to-16.7 Mbps in three 550 MHz-wide channels in the 3.1 to 5 GHz band by using pulse position modulation (PPM). The transmitter uses an all-digital architecture and calibration technique to synthesize pulses with programmable width and center frequency. The non-coherent receiver operates at 0.65 V and performs channel selection Altering, energy detection, and bit-slicing. As FCC regulations limit the maximum transmit power of UWB communication, a run-length limiting technique is presented to reduce energy requirements when maximizing range at low data rates. Patrick P. Mercier, Denis C. Daly, Manish Bhardwaj, David D. Wentzloff, Fred S. Lee, Anantha P. Chandrakasan |
ISCAS | 6 |
| 2007 | Delay-Based BPSK for Pulsed-UWB CommunicationabstractThis paper proposes a practical and effective modulation technique applicable to pulsed-UWB systems that mimics the desirable, continuous spectrum of a BPSK signal without requiring an inversion in the signal path. This technique can be used for scrambling the spectrum of a PPM signal, or as a replacement for BPSK signaling. It has been implemented in an all-digital, delay line based UWB transmitter in 90 nm CMOS. An analysis of the spectral characteristics of the modulation technique is given, as well as simulation and measured results. David D. Wentzloff, Anantha P. Chandrakasan |
ICASSP (3) | 2 |
| 2007 | A 0.4-V UWB baseband processorabstractA 0.4-V UWB digital baseband processor has been fabricated in a standard-VT 90-nm CMOS technology. The base-band processor operates at an ultra-low supply voltage to reduce energy consumption and utilizes a highly parallelized architecture to meet throughput constraints. While ultra-low voltage operation is usually limited to low energy, low performance applications, this work examines how it can be applied to low energy, high performance applications. Measured results for a 20-pJ/bit 100-Mbps UWB baseband processor are presented. Architectural techniques and design methodologies for reducing additional complexity due to parallelism are discussed. Vivienne Sze, Anantha P. Chandrakasan |
ISLPED | 2 |
| 2006 | An Energy Efficient Sub-Threshold Baseband Processor Architecture for Pulsed Ultra-Wideband CommunicationsabstractThis paper describes how parallelism in the digital baseband processor can reduce the energy required to receive ultra-wideband (UWB) packets. The supply voltage of the digital baseband is lowered so that the correlator operates near its minimum energy point resulting in a 68% energy reduction across the entire baseband. This optimum supply voltage occurs below the threshold voltage, placing the circuit in the sub-threshold region. The correlator and the rest of the baseband must be parallelized to maintain throughput at this reduced voltage. While sub-threshold operation is traditionally used for low energy, low frequency applications such as wrist-watches, this paper examines how sub-threshold operation can be applied to low energy, high performance applications. The correlators are further parallelized for a 31x reduction in the synchronization time, which along with duty-cycling, lowers the energy per packet by 43% for a 500 byte packet. Simulation results for a 100 Mbps UWB baseband processor are described Vivienne Sze, Raúl Blázquez, Manish Bhardwaj, Anantha P. Chandrakasan |
ICASSP (3) | 4 |
| 2006 | Sub-threshold design: the challenges of minimizing circuit energyabstractIn this paper, we identify the key challenges that oppose sub-threshold circuit design and describe fabricated chips that verify techniques for overcoming the challenges. Benton H. Calhoun, Alice Wang 0002, Naveen Verma, Anantha P. Chandrakasan |
ISLPED | 4 |
| 2006 | Variation-driven device sizing for minimum energy sub-threshold circuitsabstractSub-threshold operation is a compelling approach for energy-constrained applications, but increased sensitivity to variation must be mitigated. We explore variability metrics and the variation sensitivity of stacked device topologies. We show that upsizing is necessary to achieve robustness at reduced voltages and propose a design methodology to meet yield constraints. The need for upsizing imposes an energy overhead, influencing the optimal supply voltage to minimize energy. Finally, we characterize performance variability by summing delay distributions of each stage in an arbitrary critical path and achieve results accurate to within 10% of Monte Carlo simulation. Joyce Kwong, Anantha P. Chandrakasan |
ISLPED | 2 |
| 2005 | Direct Conversion Pulsed UWB Transceiver ArchitectureabstractUltra-wideband (UWB) communication is an emerging wireless technology that promises high data rates over short distances and precise locationing. The large available bandwidth and the constraint of a maximum power spectral density drives a unique set of system challenges. This paper addresses these challenges using two UWB transceivers and a discrete prototype platform. Raúl Blázquez, Fred S. Lee, David D. Wentzloff, Brian P. Ginsburg, Johnna Powell, Anantha P. Chandrakasan |
DATE | 6 |
| 2005 | Energy Efficiency of the IEEE 802.15.4 Standard in Dense Wireless Microsensor Networks: Modeling and Improvement PerspectivesabstractWireless microsensor networks, which have been the topic of intensive research in recent years, are now emerging in industrial applications. An important milestone in this transition has been the release of the IEEE 802.15.4 standard that specifies interoperable wireless physical and medium access control layers targeted to sensor node radios. In this paper, we evaluate the potential of an 802.15.4 radio for use in an ultra low power sensor node operating in a dense network. Starting from measurements carried out on the off-the-shelf radio, effective radio activation and link adaptation policies are derived. It is shown that, in a typical sensor network scenario, the average power per node can be reduced down to 211 /spl mu/W. Next, the energy consumption breakdown between the different phases of a packet transmission is presented, indicating which part of the transceiver architecture can most effectively be optimized in order to further reduce the radio power, enabling self-powered wireless microsensor networks. Bruno Bougard, Francky Catthoor, Denis C. Daly, Anantha P. Chandrakasan, Wim Dehaene |
DATE | 4 |
| 2005 | Architectures for energy-aware impulse UWB communicationsabstractUltra-wideband (UWB) signaling is an emerging technology that promises high data rates at a very low power cost. Due to the complex characteristics of the UWB channel, the signal processing required to achieve high data rates even at low distances is very large. In this paper, the complexity trade-offs associated with three elements of the digital baseband: the correlators bank, the RAKE receiver, and the Viterbi based MLSE equalizer, are presented for a practical transceiver. Raúl Blázquez, Anantha P. Chandrakasan |
ICASSP (5) | 2 |
| 2005 | Design Considerations for Ultra-Low Energy Wireless Microsensor NodesabstractThis tutorial paper examines architectural and circuit design techniques for a microsensor node operating at power levels low enough to enable the use of an energy harvesting source. These requirements place demands on all levels of the design. We propose architecture for achieving the required ultra-low energy operation and discuss the circuit techniques necessary to implement the system. Dedicated hardware implementations improve the efficiency for specific functionality, and modular partitioning permits fine-grained optimization and power-gating. We describe modeling and operating at the minimum energy point in the subthreshold region for digital circuits. We also examine approaches for improving the energy efficiency of analog components like the transmitter and the ADC. A microsensor node using the techniques we describe can function in an energy-harvesting scenario. Benton H. Calhoun, Denis C. Daly, Naveen Verma, Daniel F. Finchelstein, David D. Wentzloff, Alice Wang 0002, Seong-Hwan Cho, Anantha P. Chandrakasan |
IEEE Trans. Computers | 8 |
| 2004 | Timing, energy, and thermal performance of three-dimensional integrated circuitsabstractWe examine the performance of custom circuits in an emerging technology known as three-dimensional integration. By combining multiple device layers with a high-density inter-layer interconnect, 3D integration of a given circuit is expected to provide better timing and energy performance relative to a single-wafer implementation of the same circuit. In this paper, we show that by using our performance-driven design tool for 3D ICs, the interconnect energy dissipation of standard-cell circuits can be reduced by 24% to 42% using two to five device layers respectively. Similarly, the interconnect energy-delay product can be reduced by 30% to 50%.At the same time, thermal performance in 3D ICs is expected to be a critical issue. By incorporating thermal management and analysis into our placement tool, we may investigate the thermal scalability of 3D integration. We find that the thermal performance actually can be improved with the use of a modest number of additional device layers. Also, we show that the absolute die temperature can be controlled through the use of extra silicon. Shamik Das, Anantha P. Chandrakasan, Rafael Reif |
ACM Great Lakes Symposium on VLSI | 2 |
| 2004 | Traceback-enhanced MAP decoding algorithmabstractSoft-input soft-output algorithms are the principal component of the iterative decoding used in turbo codes and other 'turbo' feedback schemes. To enable efficient implementation, especially on energy constrained platforms such as portable devices, it is crucial to reduce the computational complexity to a minimum. We propose an enhancement to the MAX-LOG-MAP algorithm by adding a traceback operation similar to that used in the Viterbi algorithm, and devise a new efficient way to initialize the start state of the traceback. This enhancement is effective for each decoding iteration, and provides saving on top of existing techniques such as early termination and memory optimizations. It reduces the computational complexity by an additional 15%, without incurring any performance penalty. Curt Schurgers, Anantha P. Chandrakasan |
ICASSP (4) | 2 |
| 2004 | Characterizing and modeling minimum energy operation for subthreshold circuitsabstractSubthreshold operation is emerging as an energy-saving approach to many new applications. This paper examines energy minimization for circuits operating in the subthreshold region. We show the dependence of the optimum V DD for a given technology on design characteristics and operating conditions. Solving equations for total energy provides an analytical solution for the optimum V DD and V T to minimize energy for a given frequency in subthreshold operation. SPICE simulations of a 200K transistor FIR filter confirm the analytical solution and the dependence of the minimum energy operating point on important parameters. Benton H. Calhoun, Anantha P. Chandrakasan |
ISLPED | 2 |
| 2004 | Calibration of Rent's rule models for three-dimensional integrated circuitsabstractIn this paper, we determine the accuracy of Rahman's interconnect prediction model for three-dimensional (3-D) integrated circuits. Utilizing this model, we calculate the wiring requirement for a set of benchmark standard-cell circuits. We then obtain placed and routed wirelength figures for these circuits using 3-D standard-cell placement and global-routing tools we have developed. We find that the Rahman model predicts wirelengths accurately (to within 20% of placement and of routing, on average), and suggest some areas for minor improvement to the model. Shamik Das, Anantha P. Chandrakasan, Rafael Reif |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Design tools for 3-D integrated circuitsabstractWe present a set of design tools for 3-D Integration. Using these tools - a 3-D standard-cell placement tool, global routing tool, and layout editor - we have targeted existing standard-cell circuit netlists for fabrication using wafer bonding. We have analyzed the performance of several circuits using these tools and find that 3-D integration provides significant benefits. For example, relative to single-die placement, we observe on average 28% to 51% reduction in total wire length. Shamik Das, Anantha P. Chandrakasan, Rafael Reif |
ASP-DAC | 2 |
| 2003 | Scaling into Ambient Intelligence
Twan Basten, Luca Benini, Anantha P. Chandrakasan, Menno Lindwer, Jie Liu 0001, Rex Min, Feng Zhao 0001 |
DATE | 3 |
| 2003 | Power-aware architectures and circuits for FPGA-based signal processingabstractThis work showcases a power-aware system design methodology for DSP applications on reconfigurable hardware platforms. In particular, an enhanced FPGA architecture is proposed and analyzed for a deep submicron process technology. These enhancements reduce Configurable Logic Block (CLB) usage for distributed arithmetic implementations of signal processing applications by 50% or more thereby reducing the load on interconnect resources. Multi-Threshold CMOS (MTCMOS) circuit design techniques are aggressively applied to reduce subthreshold leakage using an auto power-down feature for unused logic. Results show a 14x reduction in leakage current for unused CLBs or CLBs in deep sleep mode. CLBs in active mode see up to 2.8x steady-state power reduction. A testchip demonstrating these techniques in 0.13 micron technology has been sent out for fabrication. Frank Honoré, Benton H. Calhoun, Anantha P. Chandrakasan |
FPGA | 3 |
| 2003 | Coarse acquisition for ultra wideband digital receiversabstractUltra wideband (UWB) radio is a new wireless technology that uses sub-nanosecond pulses to transmit information, resulting in a bandwidth greater than 1 GHz. The problem of synchronizing a receiver with the incoming signal grows in complexity as the signal bandwidth increases. This paper addresses coarse synchronization in UWB receivers. It analyzes how the design of the correlation process affects the time to achieve synchronization, highlighting the importance of the probability of false alarm in its performance. Raúl Blázquez, Puneet P. Newaskar, Anantha P. Chandrakasan |
ICASSP (4) | 3 |
| 2003 | Design methodology for fine-grained leakage control in MTCMOSabstractMulti-threshold CMOS is a popular technique for reducing standby leakage power with low delay overhead. MTCMOS designs typically use large sleep devices to reduce standby leakage at the block level. We provide a formal examination of sneak leakage paths and a design methodology that enables gate-level insertion of sleep devices for sequential and combinational circuits. A fabricated 0.13 /spl mu/m, dual V/sub T/ test chip employs this methodology to implement a low-power FPGA core with gate-level sleep FETs and over 8/spl times/ measured standby current reduction. The methodology allows local sleep regions that reduce leakage in active configurable logic blocks (CLBs) by up to 2.2/spl times/ (measured) for some CLB configurations. Benton H. Calhoun, Frank Honoré, Anantha P. Chandrakasan |
ISLPED | 3 |
| 2003 | Energy-aware architectures for a real-valued FFT implementationabstractEnergy-aware design is highly desirable for systems that encounter a wide diversity of operating scenarios. This is in contrast to traditional low power design for the worst case scenario, which may not be globally energy efficient. Energy-aware design focuses on enabling architectures which scale down energy as quality requirements are relaxed. A new energy-scalable system design methodology is proposed for a Real-Valued FFT processor which supports variable bit precision (8 and 16-bit precision) and variable FFT length (128- 512 point). Two energy-aware architectures, Ensemble of Point Solutions method and Reuse of Point Solutions method, are described and evaluated. Simulated and measured results show a 66% energy savings for 8-bit datapath and 52% savings for 128-point FFT length over a non-scalable approach. Alice Wang 0002, Anantha P. Chandrakasan |
ISLPED | 2 |
| 2003 | Energy reduction in VLSI computation modules: an information-theoretic approachabstractWe consider the problem of reduction of computation cost by introducing redundancy in the number of ports as well as in the input and output sequences of computation modules. Using our formulation, the classical "communication scenario" is the case when a computation module has to recompute the input sequence at a different location or time with high fidelity and low bit-error rates. We then consider communication with different computational cost objective than that given by bit-error rate. An example is communication over deep submicrometer very-large scale integration (VLSI) buses where the expected energy consumption per communicated information bit is the cost of computation. We treat this scenario using tools from information theory and establish fundamental bounds on the achievable expected energy consumption per bit in deep submicrometer VLSI buses as a function of their utilization. Some of our results also shed light on coding schemes that achieve these bounds. We then prove that the best tradeoff between the expected energy consumption per bit and bus utilization can be achieved using codes constructed from typical sequences of Markov stationary ergodic processes. We use this observation to give a closed-form expression for the best tradeoff between the expected energy consumption per bit and the utilization of the bus. This expression, in principle, can be computed using standard numerical methods. The methodology developed here naturally extends to more general computation scenarios. Paul P. Sotiriadis, Vahid Tarokh, Anantha P. Chandrakasan |
IEEE Trans. Inf. Theory | 3 |
| 2003 | Wiring requirement and three-dimensional integration technology for field programmable gate arraysabstractIn this paper, analytical models for predicting interconnect requirements in field-programmable gate arrays (FPGAs) are presented, and opportunities for three-dimensional (3-D) implementation of FPGAs are examined. The analytical models for two-dimensional FPGAs are calibrated by routing and placement experiments with benchmark circuits and extended to 3-D FPGAs. Based on system-level modeling, we find that in FPGAs with more than 20 K four-input look-up tables, the reduction in channel width, interconnect delay and power dissipation can be over 50% by 3-D implementation. Arifur Rahman, Shamik Das, Anantha P. Chandrakasan, Rafael Reif |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2003 | Instruction level and operating system profiling for energy exposed softwareabstractEnergy conscious software design can significantly improve the energy efficiency of a portable system. A software energy estimation technique using instruction class profiling is presented. The technique is shown to have an estimation error of less than 3% with trivial runtime overhead, based on a set of application programs evaluated on the StrongARM SA-1100 and Hitachi SH-4 microprocessors. A technique to isolate the switching and leakage energy components of software is outlined. The energy overhead of a real-time operating system is also profiled. The overall impact of system-level software energy management is quantified using the MIT /spl mu/AMPS system as an application example. Amit Sinha, Nathan Ickes, Anantha P. Chandrakasan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2002 | Subthreshold leakage modeling and reduction techniquesabstractAs technology scales, subthreshold leakage currents grow exponentially and become an increasingly large component of total power dissipation. CAD tools to help model and manage subthreshold leakage currents will be needed for developing ultra low power and high performance integrated circuits. This paper gives an overview of current research to control leakage currents, with an emphasis on areas where CAD improvements will be needed. The first part of the paper explores techniques to model subthreshold leakage currents at the device, circuit, and system levels. Next, circuit techniques such as source biasing, dual Vt partitioning, MTCMOS, and VTCMOS are described. These techniques reduce leakage currents during standby states and minimize power consumption. This paper also explores ways to reduce total active power by limiting leakage currents and optimally trading off between dynamic and leakage power components. James T. Kao, Siva G. Narendra, Anantha P. Chandrakasan |
ICCAD | 3 |
| 2002 | Bounding the Lifetime of Sensor Networks Via Optimal Role AssignmentsabstractA key challenge in ad-hoc, data-gathering, wireless sensor networks is achieving a lifetime of several years using nodes that carry merely hundreds of joules of stored energy. We explore the fundamental limits of energy-efficient collaborative data-gathering by deriving upper bounds on the lifetime of increasingly sophisticated sensor networks. Manish Bhardwaj, Anantha P. Chandrakasan |
INFOCOM | 2 |
| 2002 | A framework for energy-scalable communication in high-density wireless networksabstractPower-aware communication is essential for maximizing the life-time of energy-constrained wireless devices. Applications running on such devices can cooperatively reduce communication energy by trading communication latency, reliability, or range for energy savings. We introduce a framework that exposes these high level trade-offs to a power-aware communication subsystem featuring variable-strength convolutional coding, an adjustable power amplifier, and a voltage-scaled processor. An application programming interface (API) exposes an application's minimum quality constraints on the communication. These constraints are translated into energy-efficient parameter settings for the communication hardware. We apply our framework to improved communication energy models and measurements from a wireless microsensor node to effect over an order of magnitude of energy scalability. Rex Min, Anantha P. Chandrakasan |
ISLPED | 2 |
| 2002 | Full-chip sub-threshold leakage power prediction model for sub-0.18 µm CMOSabstractThe driving force for the semiconductor industry growth has been the elegant scaling nature of CMOS technology. In future CMOS technology generations, supply and threshold voltages will have to continually scale to sustain performance increase, control switching power dissipation, and maintain reliability. These continual scaling requirements on supply and threshold voltages pose several technology and circuit design challenges. With threshold voltage scaling sub-threshold leakage power is expected to become a significant portion of the total power in future CMOS systems. Therefore, it becomes crucial to predict sub-threshold leakage power of such systems. In this paper, we present a sub-threshold leakage power prediction model that takes into account within-die threshold voltage variation. Statistical measurements of 32-bit microprocessors in 0.18 mm CMOS confirms that the mean error of the model to be 4%. Comparisons of this model to two other existing models that do not take within-die threshold voltage variation into account are also presented. Siva G. Narendra, Vivek De, Shekhar Borkar, Dimitri A. Antoniadis, Anantha P. Chandrakasan |
ISLPED | 5 |
| 2002 | Energy scalable system designabstractWe introduce the notion of energy-scalable system-design. The principal idea is to maximize computational quality for a given energy constraint at all levels of the system hierarchy. The desirable energy-quality (E-Q) characteristics of systems are discussed. E-Q behavior of algorithms is considered and transforms that significantly improve scalability are analyzed using three distinct categories of commonly used signal-processing algorithms on the StrongARM SA-1100 processor as examples (viz., filtering, frequency domain transforms and classification). Scalability hooks in hardware are analyzed using similar examples on the Pentium III processor and a scalable programming methodology is proposed. Design techniques for true energy scalable hardware are also demonstrated using filtering as an example. Amit Sinha, Alice Wang 0002, Anantha P. Chandrakasan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2002 | A bus energy model for deep submicron technologyabstractWe present a comprehensive mathematical analysis of the energy dissipation in deep submicron technology buses. The energy estimation is based on an elaborate bus model that includes distributed and lumped parasitic elements that appear as technology scales. The energy drawn from the power supply during the transition of the bus is evaluated in a closed form. The notion of the transition activity of an individual line is generalized to that of the transition activity matrix of the bus. The transition activity matrix is used for statistical estimation of the power dissipation in deep submicron technology buses. Paul P. Sotiriadis, Anantha P. Chandrakasan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | An application-specific protocol architecture for wireless microsensor networksabstractNetworking together hundreds or thousands of cheap microsensor nodes allows users to accurately monitor a remote environment by intelligently combining the data from the individual nodes. These networks require robust wireless communication protocols that are energy efficient and provide low latency. We develop and analyze low-energy adaptive clustering hierarchy (LEACH), a protocol architecture for microsensor networks that combines the ideas of energy-efficient cluster-based routing and media access together with application-specific data aggregation to achieve good performance in terms of system lifetime, latency, and application-perceived quality. LEACH includes a new, distributed cluster formation technique that enables self-organization of large numbers of nodes, algorithms for adapting clusters and rotating cluster head positions to evenly distribute the energy load among all the nodes, and techniques to enable distributed signal processing to save communication resources. Our results show that LEACH can improve system lifetime by an order of magnitude compared with general-purpose multihop approaches. Wendi B. Heinzelman, Anantha P. Chandrakasan, Hari Balakrishnan |
IEEE Trans. Wirel. Commun. | 2 |
| 2001 | Reducing bus delay in submicron technology using codingabstractIn this paper we study the delay associated with transmission of data through busses. Previous work in this area has presented models for delay assuming a distributed model or a lumped capacitive coupling between wires. In this paper we extend the Elmore delay to account for a distributed model with distributed coupling component and an arbitrary number of lines driven by independent sources. The effect of data patterns is taken into account allowing us to estimate the delay on a sample by sample basis instead of making a worst case assumption. Using this detailed wire delay model, we propose a technique to speed up the communication through a data bus using coding. The idea is to encode the data being transmitted through the bus with the goal of eliminating certain types of transitions that require a large delay. We show that by using proper encoding techniques, the bus can be sped up by a factor of 2. Paul P. Sotiriadis, Anantha P. Chandrakasan |
ASP-DAC | 2 |
| 2001 | JouleTrack - A Web Based Tool for Software Energy ProfilingabstractA software energy estimation methodology is presented that avoids explicit characterization of instruction energy consumption and pre-dicts energy consumption to within 3% accuracy for a set of bench-mark programs evaluated on the StrongARM SA-1100 and Hitachi SH-4 microprocessors. The tool, JouleTrack, is available as an online resource and has various estimation levels. It also isolates the switch-ing and leakage components of the energy consumption. Amit Sinha, Anantha P. Chandrakasan |
DAC | 2 |
| 2001 | Energy efficient protocols for low duty cycle wireless microsensor networksabstractEmerging distributed wireless microsensor networks will enable the reliable and fault tolerant monitoring of the environment. Such microsensors are required to operate for years from a small energy source, while maintaining a reliable communication link to the base station. The design of energy-aware communication protocols can have a dramatic impact on the network lifetime for such applications. A detailed communication energy model, obtained from measurements, is introduced that incorporates the non-ideal behavior of the physical layer electronics. This includes the start-up energy cost of the RF transceiver, which dominates energy dissipation for short packet sizes. Using this model, various communication layer protocols are explored for asymmetrical sensor networks such as machine monitoring. The paper also proposes the use of a variable bandwidth allocation scheme that exploits spatial distribution of sensors. Seong-Hwan Cho, Anantha P. Chandrakasan |
ICASSP | 2 |
| 2001 | Energy efficient system partitioning for distributed wireless sensor networksabstractA scheme for efficient system partitioning of computation in wireless sensor networks is presented. Local computation of the sensor data in wireless networks can be highly energy-efficient, because redundant communication costs can be reduced. It is important to develop energy-efficient signal processing algorithms to be run at the sensor nodes. This paper presents a technique to optimize system energy by parallelizing computation through the network and by exploiting underlying hooks for power management. By parallelizing computation, the voltage supply level and clock frequency of the nodes can be lowered, which reduces energy dissipation. A 60% energy reduction for a sensor application of source localization is demonstrated. The results are generalized for finding optimal voltage and frequency operating points that lead to minimum system energy dissipation. Alice Wang 0002, Anantha P. Chandrakasan |
ICASSP | 2 |
| 2001 | Upper bounds on the lifetime of sensor networksabstractWe ask a fundamental question concerning the limits of energy efficiency of sensor networks-what is the upper bound on the lifetime of a sensor network that collects data from a specified region using a certain number of energy-constrained nodes? The answer to this question is valuable for two main reasons. First, it allows calibration of real world data-gathering protocols and an understanding of factors that prevent these protocols from approaching fundamental limits. Secondly, the dependence of lifetime on factors like the region of observation, the source behavior within that region, basestation location, number of nodes, radio path loss characteristics, efficiency of node electronics and the energy available on a node, is exposed. This allows architects of sensor networks to focus on factors that have the greatest potential impact on network lifetime. By employing a combination of theory and extensive simulations of constructed networks, we show that in all data gathering scenarios presented, there exist networks which achieve lifetimes equal to or >95% of the derived bounds. Hence, depending on the scenario, our bounds are either tight or near-tight. Manish Bhardwaj, Timothy Garnett, Anantha P. Chandrakasan |
ICC | 3 |
| 2001 | Energy Efficient Real-Time SchedulingabstractReal-time scheduling on processors that support dynamic voltage and frequency scaling is analyzed. The Slacked Earliest Deadline First (SEDF) algorithm is proposed and it is shown that the algorithm is optimal in minimizing processor energy consumption and maximum lateness. An upper bound on the processor energy savings is also derived. Real-time scheduling of periodic tasks is also analyzed and optimal voltage and frequency allocation for a given task set is determined that guarantees schedulability and minimizes energy consumption. Amit Sinha, Anantha P. Chandrakasan |
ICCAD | 2 |
| 2001 | Scaling of stack effect and its application for leakage reductionabstractArticle Share on Scaling of stack effect and its application for leakage reduction Authors: Siva Narendra Microsystems Technology Laboratories, Massachusetts Institute of Technology, Cambridge, MA and Microprocessor Research Laboratories, Intel Corporation, Hillsboro, OR Microsystems Technology Laboratories, Massachusetts Institute of Technology, Cambridge, MA and Microprocessor Research Laboratories, Intel Corporation, Hillsboro, ORView Profile , Vivek De Microprocessor Research Laboratories, Intel Corporation, Hillsboro, OR Microprocessor Research Laboratories, Intel Corporation, Hillsboro, ORView Profile , Dimitri Antoniadis Microsystems Technology Laboratories, Massachusetts Institute of Technology, Cambridge, MA Microsystems Technology Laboratories, Massachusetts Institute of Technology, Cambridge, MAView Profile , Anantha Chandrakasan Microsystems Technology Laboratories, Massachusetts Institute of Technology, Cambridge, MA Microsystems Technology Laboratories, Massachusetts Institute of Technology, Cambridge, MAView Profile , Shekhar Borkar Microprocessor Research Laboratories, Intel Corporation, Hillsboro, OR Microprocessor Research Laboratories, Intel Corporation, Hillsboro, ORView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 195–200https://doi.org/10.1145/383082.383132Online:06 August 2001Publication History 140citation1,436DownloadsMetricsTotal Citations140Total Downloads1,436Last 12 Months24Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Siva G. Narendra, Vivek De, Dimitri A. Antoniadis, Anantha P. Chandrakasan, Shekhar Borkar |
ISLPED | 4 |
| 2001 | Analysis and implementation of charge recycling for deep sub-micron busesabstractArticle Share on Analysis and implementation of charge recycling for deep sub-micron buses Authors: Paul Sotiriadis Department of EECS, Massachusetts Inst. of Technology Department of EECS, Massachusetts Inst. of TechnologyView Profile , Theodoros Konstantakopoulos Department of EECS, Massachusetts Inst. of Technology Department of EECS, Massachusetts Inst. of TechnologyView Profile , Anantha Chandrakasan Department of EECS, Massachusetts Inst. of Technology Department of EECS, Massachusetts Inst. of TechnologyView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 364–369https://doi.org/10.1145/383082.383184Online:06 August 2001Publication History 7citation204DownloadsMetricsTotal Citations7Total Downloads204Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Paul P. Sotiriadis, Theodoros Konstantakopoulos, Anantha P. Chandrakasan |
ISLPED | 3 |
| 2001 | Energy efficient Modulation and MAC for Asymmetric RF Microsensor SystemsabstractArticle Share on Energy efficient Modulation and MAC for Asymmetric RF Microsensor Systems Authors: Andrew Wang Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MA Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MAView Profile , SeongHwan Cho Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MA Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MAView Profile , Charles Sodini Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MA Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MAView Profile , Anantha Chandrakasan Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MA Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MAView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 106–111https://doi.org/10.1145/383082.383105Online:06 August 2001Publication History 133citation1,499DownloadsMetricsTotal Citations133Total Downloads1,499Last 12 Months8Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Andrew Y. Wang, Seong-Hwan Cho, Charles G. Sodini, Anantha P. Chandrakasan |
ISLPED | 4 |
| 2001 | Operating System and Algorithmic Techniques for Energy Scalable Wireless Sensor Networks
Amit Sinha, Anantha P. Chandrakasan |
Mobile Data Management | 2 |
| 2001 | Physical layer driven protocol and algorithm design for energy-efficient wireless sensor networksabstractThe potential for collaborative, robust networks of microsensors has attracted a great deal of research attention. For the most part, this is due to the compelling applications that will be enabled once wireless microsensor networks are in place; location-sensing, environmental sensing, medical monitoring and similar applications are all gaining interest. However, wireless microsensor networks pose numerous design challenges. For applications requiring longterm, robust sensing, such as military reconnaissance, one important challenge is to design sensor networks that have long system lifetimes. This challenge is especially difficult due to the energyconstrained nature of the devices. In order to design networks that have extremely long lifetimes, we propose a physical layer driven approach to designing protocols and algorithms. We first present a hardware model for our wireless sensor node and then introduce the design of physical layer aware protocols, algorithms, and applications that minimize energy consumption of the system. Our approach prescribes methods that can be used at all levels of the hierarchy to take advantage of the underlying hardware. We also show how to reduce energy consumption of non-ideal hardware through physical layer aware algorithms and protocols. 1 Eugene Shih, Seong-Hwan Cho, Nathan Ickes, Rex Min, Amit Sinha, Alice Wang 0002, Anantha P. Chandrakasan |
MobiCom | 7 |
| 2001 | Quantifying and enhancing power awareness of VLSI systemsabstractAn increasingly important figure-of-merit of a VLSI system is "power awareness," which is its ability to scale power consumption in response to changing operating conditions. These changes might be brought about by the time-varying nature of inputs, desired output quality, or just environmental conditions. Regardless of whether they were engineered for being power aware, systems display variations in power consumption as conditions change. This implies, by the definition above, that all systems are naturally power aware to some extent. However, one would expect that some systems are "more" power aware than others. Equivalently, we should be able to re-engineer systems to increase their power awareness. In this paper, we attempt to quantitatively define power awareness and how such awareness can be enhanced using a systematic technique. We illustrate this technique by applying it to VLSI systems at several levels of the system hierarchy - multipliers, register files, digital filters, dynamic voltage-scaled processors, and data-gathering wireless networks. It is seen that, as a result, the power awareness of these preceding systems can be significantly enhanced leading to increases in battery lifetimes in the range of 60-200%. Manish Bhardwaj, Rex Min, Anantha P. Chandrakasan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2001 | Vibration-to-electric energy conversionabstractA system is proposed to convert ambient mechanical vibration into electrical energy for use in powering autonomous low power electronic systems. The energy is transduced through the use of a variable capacitor. Using microelectromechanical systems (MEMS) technology, such a device has been designed for the system. A low-power controller IC has been fabricated in a 0.6-/spl mu/m CMOS process and has been tested and measured for losses. Based on the tests, the system is expected to produce 8 /spl mu/W of usable power. In addition to the fabricated programmable controller, an ultra low-power delay locked loop (DLL)-based system capable of autonomously achieving a steady-state lock to the vibration frequency is described. Scott E. Meninger, José Oscar Mur-Miranda, Rajeevan Amirtharajah, Anantha P. Chandrakasan, Jeffrey H. Lang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2000 | An Energy Efficient Reconfigurable Public-Key Cryptograhpy Processor Architecture
James Goodman 0001, Anantha P. Chandrakasan |
CHES | 2 |
| 2000 | A methodology for modeling the effects of systematic within-die interconnect and device variation on circuit performanceabstractWe present a methodology to study the impact of spatial pattern dependent variation on circuit performance and implement the technique in a CAD framework. We investigate the effects of interconnect CMP and poly CD device variation on interconnect delay and clock skew in both aluminum and copper interconnect technology. Our results indicate that interconnect CMP variation strongly affects interconnect delay, while poly CD variation has a large impact on clock skew in a 1 GHz design. Given this circuit impact, CAD tools in the future must account for such systematic within-die variations. Vikas Mehrotra, Shiou Lin Sam, Duane S. Boning, Anantha P. Chandrakasan, Rakesh Vallishayee, Sani R. Nassif |
DAC | 4 |
| 2000 | Energy-scalable algorithms and protocols for wireless microsensor networksabstractWireless microsensor networks lend themselves to trade-offs in energy and quality. In these networks, the individual sensor data per se are not necessarily important to the end user. Rather, it is the combined knowledge of all the sensors that describes what is occurring in the environment. By allowing the algorithms and protocols to adapt the quality of this description, with a corresponding change in energy dissipation, sensor networks can be flexible to the end-user's requirements. In this paper, we provide models for predicting quality and energy and show the advantages of trading off these two parameters. By ensuring that the system operates at a minimum energy for each quality point, the system can achieve both flexibility and energy efficiency, allowing the end-user to maximize system lifetime. Wendi B. Heinzelman, Amit Sinha, Alice Wang 0002, Anantha P. Chandrakasan |
ICASSP | 4 |
| 2000 | Bus Energy Minimization by Transition Pattern Coding (TPC) in Deep Submicron TechnologiesabstractThe energy dissipation associated with driving long wires accounts for a significant fraction of the overall system energy. This is particularly the case with the increasing importance of the inter-wire parasitic capacitance in deep sub-micron technology. A closed form solution for estimating the energy dissipation of a data bus is presented that uses an elaborate parasitic wire model. This includes the distributed RLC effects of wires as well as the coupling between wires. We also propose a general class of coding techniques to reduce energy dissipation for data transmission by trading-off between computation and communication costs. An algorithm is presented to design efficient coding strategies to minimize energy. When the effects of interwire capacitance are taken into account, the best coding strategy is not to simply minimize transitions - an approach followed by previous research. Instead, Transition Pattern Coding (TPC) modifies the transition profile to minimize energy, and in many cases higher transition activity can result in lower energy. Results show that up to a factor of 2 reduction in energy. Paul P. Sotiriadis, Anantha P. Chandrakasan |
ICCAD | 2 |
| 2000 | OpenDesign: An Open User-Configurable Project Environment for Collaborative Design and Execution on the InternetabstractOpenDesign is an open user-configurable project environment that supports distributed collaborative design and execution on the Internet. The environment is created by configuring a generic client for a specific project. This is in contrast to an implementation of a project-specific client-server architecture. This paper introduces the OpenDesign environment in the contest of a design process and project-specific tasks. An OpenDesign task is defined as execution of one or more CAD point tools, whereas a task flow is a dependency graph of tasks and/or other task flows. Challenges arise when, within a single project, (1) tasks must be executed on remote hosts under different file systems, (2) data must be accessed, moved, modified, and archived with consistency, (3) tasks and task flows are assigned to more than one designer, and (4) designers are physically dispersed. In collaboration with peer institutions, a number of demo design projects demonstrate the features and the opportunities with the OpenDesign environment. Hemang Lavana, Franc Brglez, Robert B. Reese, Gangadhar Konduri, Anantha P. Chandrakasan |
ICCD | 5 |
| 2000 | Algorithmic transforms for efficient energy scalable computationabstractWe introduce the notion of energy scalable computation on general purpose processors. The principle idea is to maximize computational qualityfor a given energy constraint. Teh desirable energy-quality behavior of algorithms is discussed. subsequently the energy-quality scalability of three distinct categories of commonly used signal processing algorithms (viz. filtering, frequency domain transforms and classification) are analyzed on the StrongARM SA-1100 processor and transformations are described which obtain significant improvements in the energy-quality scalability of the algorithm. Amit Sinha, Alice Wang 0002, Anantha P. Chandrakasan |
ISLPED | 3 |
| 2000 | Special issue on low-power RF systems
Anantha P. Chandrakasan |
Proc. IEEE | 1 |
| 2000 | High-efficiency multiple-output DC-DC conversion for low-voltage systemsabstractThis versatile power converter controller provides dual outputs at a fixed switching frequency and can regulate either output voltage or target system delay (using an external L-C filter). In the voltage regulation mode, the output voltage is monitored with an analog-digital (A/D) converter, and the feedback compensation network is implemented digitally. The generation of the pulsewidth modulation (PWM) signal is done with a hybrid delay line/counter approach, which saves power and area relative to previous implementations. Power devices are included on chip to create the two independently regulated output PWM signals. The key features of this design are its low-power dissipation, reconfigurability, use of either delay or voltage feedback, and multiple outputs. Abram P. Dancy, Rajeevan Amirtharajah, Anantha P. Chandrakasan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1999 | Design and Implementation of a Scalable Encryption Processor with Embedded Variable DC/DC ConverterabstractThis work describes the design and implementation of an energy-efficient, scalable encryption processor that utilizes variable voltage supply techniques and a highefficiency embedded variable output DC/DC converter.The resulting implementation dissipates 134nJ/bit @ V DD = 2.5V, when encrypting at its maximum rate of 1Mb/s using a maximum datapath width of 512 bits.The embedded converter achieves an efficiency of 96% at this peak load.The processor is 2-3 orders of magnitude more energy efficient than optimized assembly code running on a low-power processor such as the StrongARM. James Goodman 0001, Anantha P. Chandrakasan, Abram P. Dancy |
DAC | 2 |
| 1999 | A Framework for Collaborative and Distributed Web-Based DesignabstractThe increasing complexity and geographical separation of design data, tools and teams has created a need for a collaborative and distributed design environment. In this paper we present a framework that enables collaborative and distributed Web-based CAD, in which the designers can collaborate on a design and e ciently utilize existing design tools on the Internet. The framework includes a Java-based hierarchical collaborative schematic/block editor with interfaces to distributed Web tools and cell libraries, infrastructure to store and manipulate design objects, and protocols for tool communication, message passing and collaboration. 1 Gangadhar Konduri, Anantha P. Chandrakasan |
DAC | 2 |
| 1999 | Power scalable processing using distributed arithmeticabstractA recent trend in low power design has been the employment of reduced precision processing methods for decreasing arithmetic activity and average power dissipation.Such designs can trade off power and arithmetic precision as system requirements change.This work explores the potential of Distributed Arithmetic (DA) computation structures for low power precisionon-demand computation.We present two proof-ofconcept VLSI implementations whose power dissipation changes according to the precision of the computation performed. Rajeevan Amirtharajah, Thucydides Xanthopoulos, Anantha P. Chandrakasan |
ISLPED | 3 |
| 1999 | Vibration-to-electric energy conversionabstractA system is proposed to convert ambient mechanical vibration into electrical energy for use in powering autonomous low-power electronic systems.The energy is transduced through the use of a variable capacitor, which has been designed with MEMS (microelectromechanical systems) technology.A low-power controller IC has been fabricated in a 0.6pm CMOS process and has been tested and measured for losses.Based on the tests, the system is expected to produce SpW of usable power. Scott E. Meninger, José Oscar Mur-Miranda, Rajeevan Amirtharajah, Anantha P. Chandrakasan, Jeffrey H. Lang |
ISLPED | 4 |
| 1999 | A low power variable length decoder for MPEG-2 based on nonuniform fine-grain table partitioningabstractVariable length coding is a widely used technique in digital video compression systems. Previous work related to variable length decoders (VLDs) was primarily aimed at high throughput applications, but the increased demand for portable multimedia systems has made power a very important factor. In this paper, a data-driven variable length decoding architecture is presented, which exploits the signal statistics of variable length codes to reduce power. The approach uses fine-grain lookup table (LUT) partitioning to reduce switched capacitance based on codeword frequency. The complete VLD for MPEG-2 has been fabricated and consumes 530 /spl mu/W at 1.35 V with a video rate of 48-M discrete cosine transform samples/s using a 0.6-/spl mu/m CMOS technology. More than an order of magnitude power reduction is demonstrated without performance loss compared to a conventional parallel decoding scheme with a single LUT. Seong-Hwan Cho, Thucydides Xanthopoulos, Anantha P. Chandrakasan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1998 | MTCMOS Hierarchical Sizing Based on Mutual Exclusive Discharge PatternsabstractMulti-threshold CMOS is a popular circuit style that will provide high performance and low power operation. Optimally sizing the gating sleep transistor to provide adequate performance is difficult because the overall delay characteristics are strongly dependent on the discharge patterns of internal gates. This paper proposes a methodology for sizing the sleep transistor for a large module based on mutual exclusive discharge patterns of internal blocks. This algorithm can be applied at all levels of a circuit hierarchy, where the internal blocks can represent transistors, cells within an array, or entire modules. This methodology will give an upper bound for the sleep transistor size required to meet any performance constraint. James T. Kao, Siva G. Narendra, Anantha P. Chandrakasan |
DAC | 3 |
| 1998 | A reconfigurable dual output low power digital PWM power converterabstractMost work to date on power reduction has focused at the component level, not at the system level. In this paper, we propose a framework for describing the power behavior of system-level designs. The model consists of a set of resources, an environmental workload specification, and a power management policy, which serves as the heart of the system model. We map this model to a simulation-based framework to obtain an estimate of the system's power dissipation. Accompanying this, we propose an algorithm to optimize power management policies. The optimization algorithm can be used in a tight loop with the estimation engine to derive new power-management policy algorithms for a given system-level description. We tested our approach by applying it to a real-life low-power portable design, achieving a power estimation accuracy of ∼10%, and a 23% reduction in power after policy optimization. Abram P. Dancy, Anantha P. Chandrakasan |
ISLPED | 2 |
| 1998 | Special Section on Low-Power Electronics and Design
Anantha P. Chandrakasan, Edwin H.-M. Sha |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1998 | Low power scalable encryption for wireless systems
James Goodman 0001, Anantha P. Chandrakasan |
Wirel. Networks | 2 |
| 1998 | A low power, low bandwidth protocol for remote wireless terminals
George Hadjiyiannis, Anantha P. Chandrakasan, Srini Devadas |
Wirel. Networks | 2 |
| 1997 | Transistor Sizing Issues and Tool For Multi-Threshold CMOS TechnologyabstractMulti-threshold CMOS is an increasingly popular circuit approach that enables high performance and low power operation. However, no methodologies have been developed to size the high V t sleep transistor in an intelligent manner that trades off area and performance. In fact, many attempts at sizing the sleep transistor without close consideration of input vector patterns or internal structures can lead to large overestimates or large underestimates in sleep transistor sizing. This paper describes some of the issues involved in sizing transistors for MTCMOS and also introduces a variable breakpoint switch level simulator that can rapidly calculate delay in MTCMOS circuits as functions of design variables such as V dd , V t , and sleep transistor sizing. 1. BACKGROUND Power consumption in conventional CMOS circuits can be attributed to switching power, leakage power, and short circuit power. Switching power is usually the dominant term and is given by the well known formula: P switchi... James T. Kao, Anantha P. Chandrakasan, Dimitri A. Antoniadis |
DAC | 2 |
| 1997 | Architectural Exploration Using Verilog-Based Power Estimation: A Case Study of the IDCTabstractWe describe an architectural design space exploration methodologythat minimizes the energy dissipation of digital circuits.The centerpiece of our methodology is a Verilog-based power estimationtool, Pythia, that blends the accuracy of low-level circuitsimulators such as powermill with the speed of high level powerestimators geared to design exploration. Pythia takes into accountvoltage-dependent capacitive nonlinearities and supports runtimeadaptation of supply voltage. It employs a hybrid modelingaproach in which low-level simulation of logic gates and flip-flopscan be combined with high level macromodels for memory structureswhere the energy per access is not as sensitive to input datastatistics. The speed and accuracy of Pythia has enabled a detailedcase study of two different approaches for the computation of theInverse Discrete Cosine Transform, an integral component of theMPEG video coding algorithm. One approach uses conventionalmethods and the other exploits signal statistics to dynamically minimizethe average number of operations required and the operatingsupply voltage. Thucydides Xanthopoulos, Yoshifumi Yaoi, Anantha P. Chandrakasan |
DAC | 3 |
| 1997 | Network driven motion estimation for portable video terminalsabstractMotion estimation has been shown to help significantly in the compression of video sequences. However, since most motion estimation algorithms require a large amount of computation, it is undesirable to use them in power constrained applications, such as battery operated wireless video terminals. This paper presents an approach to reducing the power dissipation of wireless video terminals in a networked environment by exploiting the predictability of object motion. Since the location of an object in the current frame can be predicted from its location in previous frames, it is possible to optimally partition the motion estimation computation between battery operated portable devices and high powered compute servers on the wired network. This can achieve a reduction in the number of operations performed at the encoder for motion estimation by over two orders of magnitude while introducing minimal degradation to the decoded video compared with full search encoder-based motion estimation. Wendi B. Heinzelman, Anantha P. Chandrakasan |
ICASSP | 2 |
| 1997 | Low power design without compromise (panel)abstractNo abstract available. Jim Burr, Anantha P. Chandrakasan, Fari Assaderaghi, Francky Catthoor, Frank Fox, Dave Greenhill, Deo Singh, Jim Sproch |
ISLPED | 2 |
| 1997 | Network-driven motion estimation for wireless video terminalsabstractThis paper describes a new compression algorithm, termed network-driven motion estimation (NDME), which reduces the power dissipation of wireless video devices in a networked environment by exploiting the predictability of object motion. Since the location of an object in the current frame can often be predicted accurately from its location in previous frames, it is possible to optimally partition the motion estimation computation between the portable devices and high powered compute servers on the wired network. In network-driven motion estimation, a remote high-powered resource at the base-station (or on the wired network), predicts the motion vectors of the current frame from the motion vectors of the previous frames. The base-station sends these predicted motion vectors to a portable video encoder, where motion compensation proceeds as usual. Network-driven motion estimation adaptively adjusts the coding algorithm based on the amount of motion in the sequence, using motion prediction to code portions of the video sequence which contain a large amount of motion and conditional replenishment to code portions of the sequence which contain little scene motion. This algorithm achieves a reduction in the number of operations performed at the encoder for motion estimation by over two orders of magnitude while introducing minimal degradation to the decoded video compared with full search encoder-based motion estimation. Wendi B. Heinzelman, Anantha P. Chandrakasan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1997 | Embedded power supply for low-power DSPabstractThe use of dynamically adjustable power supplies as a method to lower power dissipation in DSP is analyzed. Power can be reduced substantially without sacrificing performance in fixed-throughput applications by slowing the clock and lowering supply voltage instead of idling when computational workload varies. This can yield a typical power savings of 30-50%. If latency can be tolerated, buffering data and averaging processing rate can yield power reductions of an order of magnitude in some applications. Continuous variation of the supply voltage can be approximated by very crude quantization and dithering: a four-level controller is sufficient to get within a few percent of the optimal power savings. Significant savings are possible only if the voltage can be changed on the same time scale as the variations in workload. A chip has been fabricated and tested to verify the closed-loop functionality of a variable voltage system. The controller takes only 0.4 mm/sup 2/ and draws a maximum of 1 mW at 2 V with a 40 MHz clock. The control framework developed is applicable to generic DSP applications. Vadim Gutnik, Anantha P. Chandrakasan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1996 | Design Considerations and Tools for Low-voltage Digital System DesignabstractAggressive voltage scaling to 1V and below through technology, circuit, and architecture optimization has been proven to be the key to ultra low-power design.The key technology trends for low-voltage operation are presented including low-threshold devices, multiple-threshold devices, and SOI and bulk CMOS based variable threshold devices.The requirements on CAD tools that allow designers to choose and optimize various technology, circuit, and system parameters are also discussed. Anantha P. Chandrakasan, Isabel Y. Yang, Carlin Vieri, Dimitri A. Antoniadis |
DAC | 1 |
| 1996 | A binary block matching architecture with reduced power consumption and silicon area requirementabstractMotion estimation is essential for reducing bit rate by exploiting the temporal redundancy existent in image sequences. MPEG, one of the current standards for video coding, specifies the use of block matching (BM) for motion estimation. Conventional block matching is based on the mean-absolute difference (MAD) distortion metric, which requires a large number of 8-bit arithmetic computations and thereby limits wider usage. This paper examines the possibility of using a binary distortion metric based on contour data to reduce the silicon area and power consumption of the block matching chip by a factor of 5 or more. Our simulation results indicate that the performance of the proposed binary system for a restricted class of sequences, such as "head-and-shoulder" images, is very close to that of conventional gray level methods. Detailed design of the binary block matching (BBM) chip is currently underway. Potential applications include low power, portable video devices and machine vision applications such as stereo matching and template matching. Marcelo M. Mizuki, Ujjaval Y. Desai, Ichiro Masaki, Anantha P. Chandrakasan |
ICASSP | 4 |
| 1996 | Data driven signal processing: an approach for energy efficient computingabstractThe computational switching activity of digital CMOS circuits can be dynamically minimized by designing algorithms that exploit signal statistics. This results in processors that have time-varying power requirements and perform computation on demand. An approach is presented to minimize the energy dissipation per data sample in variable-load DSP systems by adaptively minimizing the power supply voltage for each sample using a variable switching speed processor. In general, using buffering and filtering, the computation can be spread over multiple samples averaging the workload and lowering energy further. It is also shown that four levels of voltage quantization combined with dithering is sufficient to closely emulate arbitrary voltage levels. Anantha P. Chandrakasan, Vadim Gutnik, Thucydides Xanthopoulos |
ISLPED | 1 |
| 1996 | Multiple constant multiplications: efficient and versatile framework and algorithms for exploring common subexpression eliminationabstractMany applications in DSP, telecommunications, graphics, and control have computations that either involve a large number of multiplications of one variable with several constants, or can easily be transformed to that form. A proper optimization of this part of the computation, which we call the multiple constant multiplication (MCM) problem, often results in a significant improvement in several key design metrics, such as throughput, area, and power. However, until now little attention has been paid to the MCM problem. After defining the MCM problem, we introduce an effective problem formulation for solving it where first the minimum number of shifts that are needed is computed, and then the number of additions is minimized using common subexpression elimination. The algorithm for common subexpression elimination is based on an iterative pairwise matching heuristic. The power of the MCM approach is augmented by preprocessing the computation structure with a new scaling transformation that reduces the number of shifts and additions. An efficient branch and bound algorithm for applying the scaling transformation has also been developed. The flexibility of the MCM problem formulation enables the application of the iterative pairwise matching algorithm to several other important and common high level synthesis tasks, such as the minimization of the number of operations in constant matrix-vector multiplications, linear transforms, and single and multiple polynomial evaluations. All applications are illustrated by a number of benchmarks. Miodrag Potkonjak, Mani Srivastava 0001, Anantha P. Chandrakasan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1996 | Predictive system shutdown and other architectural techniques for energy efficient programmable computationabstractWith the popularity of portable devices such as personal digital assistants and personal communicators, as well as with increasing awareness of the economic and environmental costs of power consumption by desktop computers, energy efficiency has emerged as an important issue in the design of electronic systems. While power efficient ASIC's with dedicated architectures have addressed the energy efficiency issue for niche applications such as DSP, much of the computation continues to be implemented as software running on programmable processors such as microprocessors, microcontrollers, and programmable DSP's. Not only is this true for general purpose computation on personal computers and workstations, but also for portable devices, application-specific systems etc. In fact, firmware and embedded software executing on RISC and DSP processor cores that are embedded in ASIC's has emerged as a leading implementation methodology for speech coding, modem functionality, video compression, communication protocol processing etc. This paper describes architectural techniques for energy efficient implementation of programmable computation, particularly focussing on the computation needed in portable devices where event-driven user interfaces, communication protocols, and signal processing play a dominant role. Two key approaches described here are predictive system shutdown and extended voltage scaling. Results indicate that a large reduction in power consumption can be achieved over current day solutions with little or no loss in system performance. Mani Srivastava 0001, Anantha P. Chandrakasan, Robert W. Brodersen |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1995 | Synthesis and selection of DCT algorithms using behavioral synthesis-based algorithm space explorationabstractNumerous fast algorithms for the discrete cosine transform (DCT) have been proposed in image and video processing literature. Until recently, it has been difficult to compare different DCT algorithms and select one which is best suited for implementation under a given set of design goals and constraints. In this paper, we propose an approach for design space exploration at the algorithm and behavioral levels using high level synthesis tools. In particular, we study and compare the following nine DCT algorithms: Lee's, Wang's, DIT, DFT, QR, Givens, Arai, MCM, and direct algorithm. The main conclusion of this study is that the best choice among fast DCT algorithms depends on a particular set of design goals and constraints. Another important conclusion is that for almost all sets of implementation goals and constraints more than an order of magnitude improvement can be achieved using algorithm and behavioral design space exploration. Miodrag Potkonjak, Anantha P. Chandrakasan |
ICIP | 2 |
| 1995 | Minimizing power consumption in digital CMOS circuitsabstractAn approach is presented for minimizing power consumption for digital systems implemented in CMOS which involves optimization at all levels of the design. This optimization includes the technology used to implement the digital circuits, the circuit style and topology, the architecture for implementing the circuits and at the highest level the algorithms that are being implemented. The most important technology consideration is the threshold voltage and its control which allows the reduction of supply voltage without significant impact on logic speed. Even further supply reductions can be made by the use of an architecture-based voltage scaling strategy, which uses parallelism and pipelining, to tradeoff silicon area and power reduction. Since energy is only consumed when capacitance is being switched power can be reduced by minimizing this capacitance through operation reduction choice of number representation, exploitation of signal correlations, resynchronization to minimize glitching, logic design, circuit design, and physical design. The low-power techniques that are presented have been applied to the design of a chipset for a portable multimedia terminal that supports pen input, speech I/O and full-motion video. The entire chipset that performs protocol conversion, synchronization, error correction, packetization, buffering, video decompression and D/A conversion operates from a 1.1 V supply and consumes less than 5 mW.> Anantha P. Chandrakasan, Robert W. Brodersen |
Proc. IEEE | 1 |
| 1995 | Optimizing power using transformationsabstractThe increasing demand for portable computing has elevated power consumption to be one of the most critical design parameters. A high-level synthesis system, HYPER-LP, is presented for minimizing power consumption in application specific datapath intensive CMOS circuits using a variety of architectural and computational transformations. The synthesis environment consists of high-level estimation of power consumption, a library of transformation primitives, and heuristic/probabilistic optimization search mechanisms for fast and efficient scanning of the design space. Examples with varying degree of computational complexity and structures are optimized and synthesized using the HYPER-LP system. The results indicate that more than an order of magnitude reduction in power can be achieved over current-day design methodologies while maintaining the system throughput; in some cases this can be accomplished while preserving or reducing the implementation area.> Anantha P. Chandrakasan, Miodrag Potkonjak, Renu Mehra, Jan M. Rabaey, Robert W. Brodersen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1994 | Efficient Substitution of Multiple Constant Multiplications by Shifts and Additions Using Iterative Pairwise MatchingabstractMany numerically intensive applications have computations that involve a large number of multiplications of one variable with several constants.A proper optimization of this part of the computation, which we call the multiple constant multiplication (MCM) problem, often results in a significant improvement in several key design metrics.After defining the MCM problem, we formulate it as a special case of common subexpression elimination.The algorithm for common subexpression elimination is based on an iterative pairwise matching heuristic.The flexibility of the MCM problem formulation enables the application of the iterative pairwise matching algorithm to several other important high level synthesis tasks.All applications are illustrated by a number of benchmarks. Miodrag Potkonjak, Mani Srivastava 0001, Anantha P. Chandrakasan |
DAC | 3 |
| 1994 | Research challenges in wireless multimediaabstractThe near future will bring the fusion of four rapidly evolving technologies: high speed networking and associated services, wireless communications, scaled integrated circuit technology, and multimedia-based applications. These new technologies will enable the access of multimedia data from network servers at any time and any place by light weight, low cost wireless terminals. Robert W. Brodersen, Thomas D. Burd, Fred L. Burghardt, Andrew J. Burstein, Anantha P. Chandrakasan, Roger Doering, Shankar Narayanaswamy, Trevor Pering, Brian C. Richards, Thomas E. Truman, Jan M. Rabaey |
PIMRC | 5 |
| 1992 | HYPER-LP: a system for power minimization using architectural transformationsabstractAn automated high-level synthesis system, HYPER-LP, for minimizing power consumption in application-specific datapath-intensive CMOS circuits using a variety of architectural and computational transformations is presented. The sources of power consumption are reviewed, and the effects of architectural transformations on the various power components are presented. The synthesis environment consists of high-level estimation of power consumption, a library of transformation primitives (local and global), and heuristic/probabilistic optimization search mechanisms for fast and efficient scanning of the design space. Examples with varying degree of computational complexity and structures are optimized and synthesized. The results indicate that an order of magnitude reduction in power can be achieved over current-day design methodologies while maintaining the system throughput; in some cases, this can be accomplished while preserving or reducing the implementation area.> Anantha P. Chandrakasan, Miodrag Potkonjak, Jan M. Rabaey, Robert W. Brodersen |
ICCAD | 1 |