Jan M. Rabaey

dblp:r/JMRabaey · DBLP profile ↗
← Back
154ranked-venue papers
27as first author
12since 2021 · last 2025
0000-0001-6290-4855ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 97 · 21 first-author · 5 since 2021Computer networks · 21 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-authorSoftware engineering, systems software and programming languages · 8 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 EMG Acquisition and Processing for Hand Movement Decoding on Embedded Systems: State of the Art and Challenges
abstract
The electromyography (EMG) signal is particularly useful in monitoring muscle activity, and it can be acquired noninvasively on the skin surface. Thanks to these key characteristics, EMG-based human–machine interfaces (HMIs) for prosthetic myocontrol, as well as gesture recognition, are becoming widespread. A key challenge in this context is to design embedded systems to process EMG signals and generate motor commands with miniaturized, unobtrusive, and low-power devices, reliably and in real time, at a relatively low cost to provide continuous monitoring without causing stigma or discomfort. This article presents an in-depth review of the current status and future research challenges in systems and circuits for EMG acquisition and processing. We start by illustrating the sensor interfaces and acquisition systems required for signal analysis to provide efficient and effective ways of understanding the signal and its nature. We, then, focus on conventional state-of-the-art (SoA) EMG gesture recognition algorithms as well as novel architectures that tackle EMG processing challenges, i.e., hyperdimensional computing (HDC), blind source separation (BSS), and spiking neural networks (SNNs). Finally, we discuss open challenges, such as EMG variability, natural control, and efficient computation, to bring the myocontrol completely out of the laboratory, filling the gap between research prototypes and real-world applications.
Simone Benatti, Elisa Donati, Ali Moin, Marcello Zanghieri, Mattia Orlandi, Alessio Burrello, Fiorenzo Artoni, Silvestro Micera, Luca Benini, Jan M. Rabaey
Proc. IEEE10
2024 Efficient Design of a Hyperdimensional Processing Unit for Multi-Layer Cognition
abstract
The methodology used to design and optimize the very first general-purpose hyperdimensional (HD) processing unit capable of executing a broad spectrum of HD workloads (called “HPU”) is presented. HD computing is a brain-inspired computational paradigm that uses the principles of high-dimensional mathematics to perform cognitive tasks. While considerable efforts have been spent toward realizing efficient HD processors, all of these targeted specific application domains, most often pattern classification. In contrast, the HPU design addresses the multiple layers of a cognitive process. A structured methodology identifies the kernel HD computations recurring at each of these layers, and maps them onto a unified and parameterized architectural model. The effectiveness in terms of runtime and energy consumption of the approach is evaluated. The results show that the resulting HPU efficiently processes the full range of HD algorithms, and far outperforms baseline implementations on a GPU.
Mohamed Ibrahim 0002, Youbin Kim, Jan M. Rabaey
DATE3
2023 Shared Control of Assistive Robots through User-intent Prediction and Hyperdimensional Recall of Reactive Behavior
abstract
There is increasing interest in shared control for assistive robotics with adaptable levels of supervised autonomy. In this work, we present a user-adaptive multi-layer shared control scheme for control of assistive devices. The system leverages the advantages of brain-inspired hyperdimensional computing (HDC) for classification & recall of reactive robotic behavior including high performance, computational efficiency and intelligent sensor fusion, to execute actuation based on the user's goal while alleviating the burden of fine control. Using a multi-modal dataset of activities of daily living, we first recognize the user's most recent behaviors, then predict the user's next action based on their habitual action sequences, and finally, determine actuation through HDC recall-based shared control which intelligently deliberates between the predicted action and sensor feedback-based autonomy. In this work, we independently implement each layer to achieve >92% accuracy and then integrate the layers and discuss the combined performance and methods to reduce accumulated error.
Alisha Menon, Laura Isabel Galindez Olascoaga, Vamshi Balanaga, Anirudh Natarajan, Jennifer Ruffing, Ryan Ardalan, Jan M. Rabaey
ICRA7
2023 Accelerating Hyperdimensional Computing with Vector Machines
abstract
Hyperdimensional Computing (HDC) is a computationally efficient method of performing highly-accurate classification by encoding information into very wide binary vectors with simple binary operations. In this work, we explore methods of accelerating the encoding process, demonstrated on a RISC-V processor. First, we propose a bit-serial word-parallel approach to accelerate the spatial encoder, the slowest HDC block, and demonstrate its promise with a 12.6 x speedup over prior methods. Then, we describe methods to vectorize each HDC block. Implementation on a vector accelerator achieves a 12.2 x speedup and 7.1 x reduction in energy/prediction. Finally, we gain an additional 20% improvement in energy efficiency by finding the optimal balance between vector lanes and execution time, overall demonstrating the significant speed and energy improvements that a vector processor can provide for HDC.
Alisha Menon, Meek Simbule, Harrison Liew, Adriel Tan, Daniel Sun 0005, Jan M. Rabaey
ISCAS6
2023 End-to-End Planner for Self-Reconfigurable Modular Robots Collaborative Objects Manipulation, Transport and Handover to Human Application
abstract
Collaborative object manipulation and transport with self-reconfigurable modular robots can take a major role in improving modularity and adaptability of smart-home and factory-like environments. Controlling modules to achieve efficient behaviours is challenging due to the high number of degrees of freedom in the system and the physical constraints. We present an end-to-end planner that discovers collaborative behaviours for modules to manipulate and transport objects to bring them to a human defined place. Our approach is based on a centralized planner using stochastic best-first search with a custom heuristic and pruning strategy. We use Quadratic Programming to define multi-robot controller to evaluate action feasibility for transitions between the search tree nodes with respect to important constraints of the system (collisions, joint and torque limits). The controller can be design to be aware of human reachable space for object handover and use it as a measure to asses closeness to the goal node. Results show that the proposed method can effectively coordinate the actions of multiple robots, leading to an emerging efficient manipulation and transport of objects with variable shapes and weight to within human reachable space. This work brings self-reconfigurable modular robots one step closer to assistive human-robot interaction applications or smart logistics.
Aurélien Morel, Anastasia Bolotnikova, Celinna Ju, Jan M. Rabaey, Auke Jan Ijspeert
RO-MAN4
2023 Generalized Key-Value Memory to Flexibly Adjust Redundancy in Memory-Augmented Networks
abstract
Memory-augmented neural networks enhance a neural network with an external key-value (KV) memory whose complexity is typically dominated by the number of support vectors in the key memory. We propose a generalized KV memory that decouples its dimension from the number of support vectors by introducing a free parameter that can arbitrarily add or remove redundancy to the key memory representation. In effect, it provides an additional degree of freedom to flexibly control the tradeoff between robustness and the resources required to store and compute the generalized KV memory. This is particularly useful for realizing the key memory on in-memory computing hardware where it exploits nonideal, but extremely efficient nonvolatile memory devices for dense storage and computation. Experimental results show that adapting this parameter on demand effectively mitigates up to 44% nonidealities, at equal accuracy and number of devices, without any need for neural network retraining.
Denis Kleyko, Geethan Karunaratne, Jan M. Rabaey, Abu Sebastian, Abbas Rahimi
IEEE Trans. Neural Networks Learn. Syst.3
2022 On the Role of Hyperdimensional Computing for Behavioral Prioritization in Reactive Robot Navigation Tasks
abstract
Hyperdimensional computing (HDC) is a brain-inspired computing paradigm that operates on pseudo-random hypervectors, an information-rich, hardware-efficient representation that is robust to noise and facilitates learning with limited training data. This work explores how robot navigation tasks can leverage the high-capacity hypervector representation to enable behavioral prioritization through a weighted encoding of heterogeneous sensor information. Experiments over 100 trials in each of the 100 randomly generated obstacle maps demonstrate that the proposed weighted sensor encoding scheme boosts the success rate of the navigation task by over 30% compared to an unweighted sensor encoding. A hybrid scheme using the HDC weighted scheme at the input of a deep feed-forward neural network achieves the highest success rate. The hybrid scheme furthermore is more robust when reducing the HDC dimension by 50%. However, the simple HDC implementation remains the most hardware efficient, making it desirable for resource-constrained systems.
Alisha Menon, Anirudh Natarajan, Laura Isabel Galindez Olascoaga, Youbin Kim, Braeden C. Benedict, Jan M. Rabaey
ICRA6
2022 Vector Symbolic Architectures as a Computing Framework for Emerging Hardware
abstract
(also known as Hyperdimensional Computing). This framework is well suited for implementation in stochastic, emerging hardware and it naturally expresses the types of cognitive operations required for Artificial Intelligence (AI). We demonstrate in this article that the field-like algebraic structure of Vector Symbolic Architectures offers simple but powerful operations on high-dimensional vectors that can support all data structures and manipulations relevant to modern computing. In addition, we illustrate the distinguishing feature of Vector Symbolic Architectures, "computing in superposition," which sets it apart from conventional computing. It also opens the door to efficient solutions to the difficult combinatorial search problems inherent in AI applications. We sketch ways of demonstrating that Vector Symbolic Architectures are computationally universal. We see them acting as a framework for computing with distributed representations that can play a role of an abstraction layer for emerging computing hardware. This article serves as a reference for computer architects by illustrating the philosophy behind Vector Symbolic Architectures, techniques of distributed computing with them, and their relevance to emerging computing hardware, such as neuromorphic computing.
Denis Kleyko, Mike Davies 0002, Edward Paxon Frady, Pentti Kanerva, Spencer J. Kent, Bruno A. Olshausen, Evgeny Osipov, Jan M. Rabaey, Dmitri A. Rachkovskij, Abbas Rahimi, Friedrich T. Sommer
Proc. IEEE8
2021 Generalized Learning Vector Quantization for Classification in Randomized Neural Networks and Hyperdimensional Computing
abstract
Machine learning algorithms deployed on edge devices must meet certain resource constraints and efficiency requirements. Random Vector Functional Link (RVFL) networks are favored for such applications due to their simple design and training efficiency. We propose a modified RVFL network that avoids computationally expensive matrix operations during training, thus expanding the network's range of potential applications. Our modification replaces the least-squares classifier with the Generalized Learning Vector Quantization (GLVQ) classifier, which only employs simple vector and distance calculations. The GLVQ classifier can also be considered an improvement upon certain classification algorithms popularly used in the area of Hyperdimensional Computing. The proposed approach achieved state-of-the-art accuracy on a collection of datasets from the UCI Machine Learning Repository - higher than previously proposed RVFL networks. We further demonstrate that our approach still achieves high accuracy while severely limited in training iterations (using on average only 21% of the least-squares classifier computational costs).
Cameron Diao, Denis Kleyko, Jan M. Rabaey, Bruno A. Olshausen
IJCNN3
2021 Of Brains and Computers
abstract
The human brain - which we consider to be the prototypal biological computer - in its current incarnation is the result of more than a billion years of evolution. Its main functions have always been to regulate the internal milieu and to help the organism/being to survive and reproduce. With growing complexity, the brain has adapted a number of design principles that serve to maximize its efficiency in performing a broad range of tasks. The physical computer, on the other hand, had only 200 years or so to evolve, and its perceived function was considerably different and far more constraint - that is to solve a set of mathematical functions. This however is rapidly changing. One may argue that the functions of brains and computers are converging. If so, the question arises if the underlaying design principles will converge or cross-breed as well, or will the different underlaying mechanisms (physics versus biology) lead to radically different solutions.
Jan M. Rabaey
ISPD1
2021 Millimetro: mmWave retro-reflective tags for accurate, long range localization
abstract
This paper presents Millimetro, an ultra-low-power tag that can be localized at high accuracy over extended distances. We develop Millimetro in the context of autonomous driving to efficiently localize roadside infrastructure such as lane markers and road signs, even if obscured from view, where visual sensing fails. While RF-based localization offers a natural solution, current ultra-low-power localization systems struggle to operate accurately at extended ranges under strict latency requirements. Millimetro addresses this challenge by re-using existing automotive radars that operate at mmWave frequency where plentiful bandwidth is available to ensure high accuracy and low latency. We address the crucial free space path loss problem experienced by signals from the tag at mmWave bands by building upon Van Atta Arrays that retro-reflect incident energy back towards the transmitting radar with minimal loss and low power consumption. Our experimental results indoors and outdoors demonstrate a scalable system that operates at a desirable range (over 100 m), accuracy (centimeter-level), and ultra-low-power (< 3 uW).
Elahe Soltanaghai, Akarsh Prabhakara, Artur Balanuta, Matthew G. Anderson, Jan M. Rabaey, Swarun Kumar, Anthony Rowe 0001
MobiCom5
2021 Adaptive Body Area Networks Using Kinematics and Biosignals
abstract
The increasing penetration of wearable and implantable devices necessitates energy-efficient and robust ways of connecting them to each other and to the cloud. However, the wireless channel around the human body poses unique challenges such as a high and variable path-loss caused by frequent changes in the relative node positions as well as the surrounding environment. An adaptive wireless body area network (WBAN) scheme is presented that reconfigures the network by learning from body kinematics and biosignals. It has very low overhead since these signals are already captured by the WBAN sensor nodes to support their basic functionality. Periodic channel fluctuations in activities like walking can be exploited by reusing accelerometer data and scheduling packet transmissions at optimal times. Network states can be predicted based on changes in observed biosignals to reconfigure the network parameters in real time. A realistic body channel emulator that evaluates the path-loss for everyday human activities was developed to assess the efficacy of the proposed techniques. Simulation results show up to 41% improvement in packet delivery ratio (PDR) and up to 27% reduction in power consumption by intelligent scheduling at lower transmission power levels. Moreover, experimental results on a custom test-bed demonstrate an average PDR increase of 20% and 18% when using our adaptive EMG- and heart-rate-based transmission power control methods, respectively. The channel emulator and simulation code is made publicly available at https://github.com/a-moin/wban-pathloss.
Ali Moin, Arno Thielens, Álvaro Araujo, Alberto L. Sangiovanni-Vincentelli, Jan M. Rabaey
IEEE J. Biomed. Health Informatics5
2020 Heartbeat-Based Synchronization Scheme for the Human Intranet: Modeling and Analysis
abstract
Sharing a common clock signal among the nodes is crucial for communication in synchronized networks. This work presents a heartbeat-based synchronization scheme for body-worn nodes. The principles of this coordination technique combined with a puncture-based communication method are introduced. Theoretical models of the hardware blocks are presented, outlining the impact of their specifications on the system. Moreover, we evaluate the synchronization efficiency in simulation and compare with a duty-cycled receiver topology. Improvement in power consumption of at least 26% and tight latency control are highlighted at no cost on the channel availability.
Robin Benarrouch, Ali Moin, Flavien Solt, Antoine Frappé, Andreia Cathelin, Andreas Kaiser, Jan M. Rabaey
ISCAS7
2020 Hyperdimensional Computing for Blind and One-Shot Classification of EEG Error-Related Potentials
Abbas Rahimi, Artiom Tchouprina, Pentti Kanerva, José del R. Millán, Jan M. Rabaey
Mob. Networks Appl.5
2020 QuantHD: A Quantization Framework for Hyperdimensional Computing
abstract
Brain-inspired hyperdimensional (HD) computing models cognition by exploiting properties of high dimensional statistics-high-dimensional vectors, instead of working with numeric values used in contemporary processors. A fundamental weakness of existing HD computing algorithms is that they require to use floating point models in order to provide acceptable accuracy on realistic classification problems. However, working with floating point values significantly increases the HD computation cost. To address this issue, we proposed QuantHD, a novel framework for quantization of HD computing model during training. QuantHD enables HD computing to work with a low-cost quantized model (binary or ternary model) while providing a similar accuracy as the floating point model. We accordingly propose an FPGA implementation which accelerates HD computing in both training and inference phases. We evaluate QuantHD accuracy and efficiency on various real-world applications, and observe that QuantHD can achieve on average 17.2% accuracy improvement as compared to the existing binarized HD computing algorithms which provide a similar computation cost. In terms of efficiency, QuantHD FPGA implementation can achieve on average 42.3× and 4.7× (34.1× and 4.1×) energy efficiency improvement and speedup during inference (training) as compared to the state-of-the-art HD computing algorithms.
Mohsen Imani, Samuel Bosch, Sohum Datta, Sharadhi Ramakrishna, Sahand Salamat, Jan M. Rabaey, Tajana Rosing
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2020 Human-Centric Computing
abstract
With the world around us rapidly becoming smarter, an extremely relevant question is how “we humans” are going to cope with the onslaught of information coming at us. One plausible answer is to use similar technologies to evolve ourselves and to equip us with the necessary tools to interact with and to become an essential part of the smart world. Various wearable devices have been or are being developed for this purpose. However, their potential to create a whole new set of human experiences is still largely unexplored. To be more effective, functionality cannot be centralized and needs to be distributed to capture the right information at the right place. This requires a human Intranet, a platform that allows multiple distributed input-output and information processing functions to coalesce and form a single application. In addition, it needs the capabilities to understand, interpret, reason, and act on the obtained data under diverse and changing conditions, and to do so in concert with the human body and its computer, the brain. To this effect, this article explores the concept of human-centric computing, an approach that aspires to create a symbiotic convergence between biological and physical computing.
Jan M. Rabaey
IEEE Trans. Very Large Scale Integr. Syst.1
2019 Efficient Biosignal Processing Using Hyperdimensional Computing: Network Templates for Combined Learning and Classification of ExG Signals
abstract
Recognizing the very size of the brain's circuits, hyperdimensional (HD) computing can model neural activity patterns with points in a HD space, that is, with HD vectors. Key examined properties of HD computing include: a versatile set of arithmetic operations on HD vectors, generality, scalability, analyzability, one-shot learning, and energy efficiency. These make it a prime candidate for efficient biosignal processing where signals are noisy and nonstationary, training data sets are not huge, individual variability is significant, and energy-efficiency constraints are tight. Purely based on native HD computing operators, we describe a combined method for multiclass learning and classification of various ExG biosignals such as electromyography (EMG), electroencephalography (EEG), and electrocorticography (ECoG). We develop a full set of HD network templates that comprehensively encode body potentials and brain neural activity recorded from different electrodes into a single HD vector without requiring domain expert knowledge or ad hoc electrode selection process. Such encoded HD vector is processed as a single unit for fast one-shot learning, and robust classification. It can be interpreted to identify the most useful features as well. Compared to state-of-the-art counterparts, HD computing enables online, incremental, and fast learning as it demands less than a third as much training data as well as less preprocessing.
Abbas Rahimi, Pentti Kanerva, Luca Benini, Jan M. Rabaey
Proc. IEEE4
2018 An EMG Gesture Recognition System with Flexible High-Density Sensors and Brain-Inspired High-Dimensional Classifier
abstract
EMG-based gesture recognition shows promise for human-machine interaction. Systems are often afflicted by signal and electrode variability which degrades performance over time. We present an end-to-end system combating this variability using a large-area, high-density sensor array and a robust classification algorithm. EMG electrodes are fabricated on a flexible substrate and interfaced to a custom wireless device for 64-channel signal acquisition and streaming. We use brain-inspired high-dimensional (HD) computing for processing EMG features in one-shot learning. The HD algorithm is tolerant to noise and electrode misplacement and can quickly learn from few gestures without gradient descent or back-propagation. We achieve an average classification accuracy of 96.64% for five gestures, with only 7% degradation when training and testing across different days. Our system maintains this accuracy when trained with only three trials of gestures; it also demonstrates comparable accuracy with the state-of-the-art when trained with one trial.
Ali Moin, Andy Zhou, Abbas Rahimi, Simone Benatti, Alisha Menon, Senam Tamakloe, Jonathan Ting, Natasha Yamamoto, Yasser Khan, Fred L. Burghardt, Luca Benini, Ana Claudia Arias, Jan M. Rabaey
ISCAS13
2018 Towards TRUE human-centric computation
Jan M. Rabaey
Comput. Commun.1
2018 Classification and Recall With Binary Hyperdimensional Computing: Tradeoffs in Choice of Density and Mapping Characteristics
abstract
Hyperdimensional (HD) computing is a promising paradigm for future intelligent electronic appliances operating at low power. This paper discusses tradeoffs of selecting parameters of binary HD representations when applied to pattern recognition tasks. Particular design choices include density of representations and strategies for mapping data from the original representation. It is demonstrated that for the considered pattern recognition tasks (using synthetic and real-world data) both sparse and dense representations behave nearly identically. This paper also discusses implementation peculiarities which may favor one type of representations over the other. Finally, the capacity of representations of various densities is discussed.
Denis Kleyko, Abbas Rahimi, Dmitri A. Rachkovskij, Evgeny Osipov, Jan M. Rabaey
IEEE Trans. Neural Networks Learn. Syst.5
2017 Optimized Design of a Human Intranet Network
abstract
We address the design space exploration of wireless body area networks for wearable and implantable technologies, a task that is increasingly challenging as the number and variety of devices per person grow. Our method efficiently decomposes the problem into smaller subproblems by coordinating specialized analysis and optimization techniques. We leverage mixed integer linear programming to generate candidate network configurations based on coarse energy estimations. Accurate discrete-event simulation is used to check the feasibility of the proposed configurations under reliability constraints and guide the search to achieve fast convergence. Numerical results show that our application-specific approach substantially reduces the exploration time with respect to generic optimization techniques and helps provide clear identification of promising solutions.
Ali Moin, Pierluigi Nuzzo 0002, Alberto L. Sangiovanni-Vincentelli, Jan M. Rabaey
DAC4
2017 A Systems Approach to Computing in Beyond CMOS Fabrics: Invited
abstract
No abstract available.
Ameya Patil 0001, Naresh R. Shanbhag, Lav R. Varshney, Eric Pop, H.-S. Philip Wong, Subhasish Mitra, Jan M. Rabaey, Jeffrey A. Weldon, Lawrence T. Pileggi, Sasikanth Manipatruni, Dmitri E. Nikonov, Ian A. Young
DAC7
2017 Exploring Hyperdimensional Associative Memory
abstract
Brain-inspired hyperdimensional (HD) computing emulates cognition tasks by computing with hypervectors as an alternative to computing with numbers. At its very core, HD computing is about manipulating and comparing large patterns, stored in memory as hypervectors: the input symbols are mapped to a hypervector and an associative search is performed for reasoning and classification. For every classification event, an associative memory is in charge of finding the closest match between a set of learned hypervectors and a query hypervector by using a distance metric. Hypervectors with the i.i.d. components qualify a memory-centric architecture to tolerate massive number of errors, hence it eases cooperation of various methodological design approaches for boosting energy efficiency and scalability. This paper proposes architectural designs for hyperdimensional associative memory (HAM) to facilitate energy-efficient, fast, and scalable search operation using three widely-used design approaches. These HAM designs search for the nearest Hamming distance, and linearly scale with the number of dimensions in the hypervectors while exploring a large design space with orders of magnitude higher efficiency. First, we propose a digital CMOS-based HAM (D-HAM) that modularly scales to any dimension. Second, we propose a resistive HAM (R-HAM) that exploits timing discharge characteristic of nonvolatile resistive elements to approximately compute Hamming distances at a lower cost. Finally, we combine such resistive characteristic with a currentbased search method to design an analog HAM (A-HAM) that results in faster and denser alternative. Our experimental results show that R-HAM and A-HAM improve the energy-delay product by 9.6× and 1347× compared to D-HAM while maintaining a moderate accuracy of 94% in language recognition.
Mohsen Imani, Abbas Rahimi, Deqian Kong, Tajana Rosing, Jan M. Rabaey
HPCA5
2017 Selection and Aggregation of Location Information Provisioning Services
abstract
Aggregation of location estimates from multiple services for provisioning of location information enhances the accuracy and robustness of the final location information. Recent Internet of Things (IoT)-based localization service architectures therefore envision a "manager" for selecting and invoking provisioning services and, in the later step, aggregating the received information and providing it to location-based applications. The selection of provisioning services should take into account the accuracy and latency requirements from the applications and accuracy, latency, and power consumption characteristics of provisioning services. However, it is yet unclear how such selection should be made. We propose two algorithms for the selection of provisioning services aiming at meeting latency and subsequently accuracy requirements from the applications, one subject to minimizing per-request power consumption, while the other subject to a per-time bucket power minimization. In the considered examples, we show that the per-time bucket optimization achieves around 25% better performance in terms of power consumption, while trading-off accuracy satisfaction.
Filip Lemic, Vlado Handziski, Mladen Miksa, Jan M. Rabaey, John Wawrzynek, Adam Wolisz
ICCCN4
2017 SLSR: A flexible middleware localization service architecture
abstract
Location information of mobile devices is a foundational input to location-based services and a valuable source of context information in wireless networks. To maximize the value, we need location information that is accurate, robust, and promptly and seamlessly available. Unfortunately, individual localization services seldom satisfy all these requirements. For achieving that vision, a set of challenges has to be addressed, pertaining to handover, fusion, and integration of different sources of location information. Current approaches for integration of individual localization services are either not specific enough or are limited in scope and lack flexibility. In the following, we provide a detailed design and a prototypical implementation of the Standardized Localization Service (SLSR), a middleware architecture for achieving those goals. We instantiate the service in an office environment and perform exhaustive performance benchmarking in a testbed specifically designed for supporting such experimentation. Our results characterize the effects of different functional components envisioned in the SLSR on its performance. Our results also quantify the accuracy benefits of fusion of representative sources of location information.
Filip Lemic, Vlado Handziski, Ivan Azcarate, John Wawrzynek, Jan M. Rabaey, Adam Wolisz
IPIN5
2017 Reliable Next-Generation Cortical Interfaces for Chronic Brain-Machine Interfaces and Neuroscience
abstract
This review focuses on recent directions stemming from work by the authors and collaborators in the emerging field of neurotechnology. Neurotechnology has the potential to provide a greater understanding of the structure and function of the complex neural circuits in the brain, as well as impacting the field of brain-machine interfaces (BMI). We envision ultralow-power wireless neural interface systems that are life-lasting, fully integrated, and that supports bidirectional data flow with high bandwidth. Moreover, we believe in the importance of building neural interface technology that is truly tetherless, has a very small recording footprint, and little to no mechanical coupling between the sensor and the external world. We believe these developments will impact both neuroscience and neurology, revealing fundamental insight about how the nervous system functions in health and disease.
Michel M. Maharbiz, Rikky Muller, Elad Alon, Jan M. Rabaey, Jose M. Carmena
Proc. IEEE4
2017 Guest Editorial: Alternative Computing and Machine Learning for Internet of Things
abstract
The impending Internet of Things (IoT) wave is promising to affect every aspect of our daily lives, ranging from smart things to smart buildings, smart cities, and smart environments. A lot of attention has been devoted to the tsunami of data produced by IoT, and the related means of extracting useful actionable information from it, spawning efforts in Big Data processing and machine learning. Yet, all of this does little to address the need for IoT to capture, interpret, and act on this wall of (noisy) information at the right time, at the right place, and in the right form. Conventional computing systems are a poor match to the needs of this emerging massively distributed real-time system. Hence, alternative computing techniques present an attractive alternative, trading off computational resolution for significant gains in quality-of-service energy efficiency and robustness. This observation is based on the conjecture that most applications related to IoT have an inherent error resilience and are evolutionary (that is, learning-based). Alternative computing strategies may be conceived at every level of the design hierarchy, starting from the device level with novel 3-D nonvolatile memory/logic combinations, or at the architectural level by shifting away from the traditional von Neumann architecture to different computing paradigms such as neuromorphic and/or stochastic computation all the way up to the algorithmic and data representation levels.
Farshad Firouzi, Bahareh J. Farahani, Andrew B. Kahng, Jan M. Rabaey, Natasha Balac
IEEE Trans. Very Large Scale Integr. Syst.4
2016 Toward standardized localization service
abstract
Localization services available on today's mobile devices are proprietary and leverage a limited set of sources of location information. Integration of new location estimation methods is therefore cumbersome, requiring adaptation to the specific interfaces of the proprietary location service. In addition, location-based applications are tightly interwoven with the location service that is typically provided by the operating system, hence these applications require significant restructuring to be able to run with another location service. To address these problems, we propose a modular localization service architecture that consists of location-based applications, an integrated location service enabling a fusion of so-called elementary location services, and resources for generating location information. A unified style of interaction among these components is enabled by a set of well-defined Application Programming Interfaces (APIs). The practicability and advantages of the proposed design is demonstrated by outlining how the APIs can be realized using modern types of component interactions.
Filip Lemic, Vlado Handziski, Nitesh Mor, Jan M. Rabaey, John Wawrzynek, Adam Wolisz
IPIN4
2016 A Robust and Energy-Efficient Classifier Using Brain-Inspired Hyperdimensional Computing
abstract
The mathematical properties of high-dimensional (HD) spaces show remarkable agreement with behaviors controlled by the brain. Computing with HD vectors, referred to as "hypervectors," is a brain-inspired alternative to computing with numbers. Hypervectors are high-dimensional, holographic, and (pseudo)random with independent and identically distributed (i.i.d.) components. They provide for energy-efficient computing while tolerating hardware variation typical of nanoscale fabrics. We describe a hardware architecture for a hypervector-based classifier and demonstrate it with language identification from letter trigrams. The HD classifier is 96.7% accurate, 1.2% lower than a conventional machine learning method, operating with half the energy. Moreover, the HD classifier is able to tolerate 8.8-fold probability of failure of memory cells while maintaining 94% accuracy. This robust behavior with erroneous memory cells can significantly improve energy efficiency.
Abbas Rahimi, Pentti Kanerva, Jan M. Rabaey
ISLPED3
2016 On the Total Power Capacity of Regular-LDPC Codes With Iterative Message-Passing Decoders
abstract
Motivated by recently derived fundamental limits on total (transmit + decoding) power for coded communication with VLSI decoders, this paper investigates the scaling behavior of the minimum total power needed to communicate over AWGN channels as the target bit-error-probability tends to zero. We focus on regular-LDPC codes and iterative message-passing decoders. We analyze scaling behavior under two VLSI complexity models of decoding. One model abstracts power consumed in processing elements (node model), and another abstracts power consumed in wires which connect the processing elements (wire model). We prove that a coding strategy using regular-LDPC codes with Gallager-B decoding achieves order-optimal scaling of total power under the node model. However, we also prove that regular-LDPC codes and iterative message-passing decoders cannot meet existing fundamental limits on total power under the wire model. Furthermore, if the transmit energy-per-bit is bounded, total power grows at a rate that is worse than uncoded transmission. Complementing our theoretical results, we develop detailed physical models of decoding implementations using post-layout circuit simulations. Our theoretical and numerical results show that approaching fundamental limits on total power requires increasing the complexity of both the code design and the corresponding decoding algorithm as communication distance is increased or error-probability is lowered.
Karthik Ganesan 0001, Pulkit Grover, Jan M. Rabaey, Andrea J. Goldsmith
IEEE J. Sel. Areas Commun.3
2015 The human intranet: where swarms and humans meet
Jan M. Rabaey
DATE1
2014 Novel Class of Energy-Efficient Very High-Speed Conditional Push-Pull Pulsed Latches
abstract
In this paper, a new class of pulsed latches is introduced and experimentally assessed in 65-nm CMOS. Its conditional push-pull pulsed latch topology is based on a push- pull final stage driven by two split paths with a conditional pulse generator. Two circuit implementations of the concept are discussed, with their main difference being in the pulse generator, which can be either shared (CSP3L) or not (CP3L). Measurements show that the proposed topology is very fast, as it outperforms the well-known transmission gate pulsed latch (TGPL) [1] by 1.5×-2×; hence the proposed pulsed latch has the highest performance ever reported. The proposed pulsed latch is also shown to significantly improve the energy efficiency compared to the state of the art. Indeed, a 2.3× improvement in ED3 product (energy × delay3) over TGPL was found for designs targeting minimum ED3. For designs targeting minimum ED, a 1.3× improvement was found in ED product. This comes at the cost of a 1.15×-1.35× cell area penalty, which translates into an overall area increase well below 1% in typical systems. Measurements on 256 replicas confirm that the above benefits are kept in the presence of variations. Accordingly, the proposed class of pulsed latches goes beyond the current state of the art and is well suited for VLSI systems that require both high performance and energy efficiency.
Elio Consoli, Gaetano Palumbo, Jan M. Rabaey, Massimo Alioto
IEEE Trans. Very Large Scale Integr. Syst.3
2013 Panel: the heritage of Mead & Conway: what has remained the same, what was missed, what has changed, what lies ahead
Marco Casale-Rossi, Alberto L. Sangiovanni-Vincentelli, Luca P. Carloni, Bernard Courtois, Hugo De Man, Antun Domic, Jan M. Rabaey
DATE7
2013 Data communication and power system for wireless neural recording
abstract
An RFID reader collecting neural data from implanted electrodes of an electro-corticography ECoG system is proposed. The system also is able to power an implant using a class E power amplifier (PA) while it receives data by an asynchronous demodulation. A standard 65nm CMOS TSMC technology is used. Simulations reveal an average power consumption of the overall system of 19 mW with on a 1.2V supply a 300MHz carrier frequency. The data transmission is 1MHz and a bit error rate (BER) of 0.3% is evaluated.
Daniela De Venuto, Jan M. Rabaey
ETFA2
2013 Energy detection technique for ultra-low power high sensitivity wake-up receiver
abstract
An energy detection technique for improving the sensitivity of an ultra-low power wake-up receiver is presented. This technique uses an envelope detector, an integrator and a comparator to reduce the excess noise introduced by wide bandwidth of IF gain stage. Theoretical analysis shows that with 10us integration time, it can effectively remove more than 15 dB of excess noise, bring the minimum required SNR for 10-3energy detection error rate down as low as -4dB. With an expected 18dB noise figure due to low current budget, the receiver could achieve -85dBm sensitivity. The performance can be further improved to -95dBm by increasing integration time to 1000us which leads to a trade-off between achievable sensitivity and wake-up latency.
Jan M. Rabaey
ISCAS2
2012 Choosing "green" codes by simulation-based modeling of implementations
abstract
How do we design an error correcting code and a corresponding decoding implementation to minimize not just the transmit power, but the sum of transmit and decoding power? Recent interest in this question has led to new fundamental results that show the traditional approach of designing the code and the decoder implementation in isolation can be suboptimal. However, joint design of codes and their corresponding decoder implementations can be hard simply because of the sheer number of possibilities for both, and the human effort often required in optimizing the decoder implementation for a given code. In this paper, we suggest taking a middle-path between analyzing theoretical models of decoding and building decoder implementations. Based on circuit simulations of power consumption of decoders for simple regular LDPC codes, we develop circuit models for the decoding power for larger and more complex (but still regular) LDPC codes. These models are then used to search for the best code and corresponding decoder (within a limited set) for a given communication distance and error probability.
Karthik Ganesan 0001, Pulkit Grover, Andrea J. Goldsmith, Jan M. Rabaey
GLOBECOM5
2011 Powering and communicating with mm-size implants
abstract
This paper deals with system level design considerations for mm-size implantable electronic devices with wireless connectivity. In particular, it focuses on neural sensors as one application requiring such miniature interfaces. Common to all these implants is the need for power supply and a wireless interface. Wireless power transfer via electromagnetic fields is identified as a promising option for powering such devices. Design methodologies, system level trade-offs, as well as limitations of power supply systems based on electromagnetic coupling are discussed in detail. Further, various wireless data communication architectures are evaluated for their feasibility in the application. Reflective impulse radios are proposed as an alternative scheme for enabling highly scalable data transmission at <;1pJ/bit. Finally, design considerations for the corresponding reader system are addressed.
Jan M. Rabaey, Michael Mark, Christopher Sutardja, Chongxuan Tang, Suraj Gowda, Mark Wagner, Dan Werthimer
DATE1
2011 A low-leakage parallel CRC generator for ultra-low power applications
abstract
Unlike static CMOS circuits, the standby energy in sense amplifier-based pass transistor logic (SAPTL) circuits can be decoupled from its performance, allowing separate optimization strategies for leakage and speed. In this paper, a 64-byte parallel cyclic-redundancy check (CRC) generator is designed and implemented using asynchronous 90nm SAPTL circuits, with a simulated minimum energy point that is 25% lower than the equivalent complementary static CMOS implementation. The low leakage operation results in (a) an 7.9X reduction in measured energy when VDDis reduced from 1V to 0.3V at α = 0.1 and (b) a 10% reduction in measured delay with a stack forward body bias of 0.4V, with no corresponding increase in energy.
Louis P. Alarcón, Tsung-Te Liu, Jan M. Rabaey
ISCAS3
2011 Linearity analysis of CMOS passive mixer
abstract
The analysis of distortion behavior in a CMOS passive mixer is presented. We use a simple device model with continuous equation to characterize the MOSFET switch in different operating regions. The nonlinear behaviors of passive mixers based on different switch topologies are investigated with power series analysis, which demonstrates that the CMOS switch based approach exhibits a better linearity performance due to the distortion cancellation of even order nonlinearity. The calculated conversion gain and linearity performances are compared with simulation results.
Tsung-Te Liu, Jan M. Rabaey
ISCAS2
2011 Digital energy detection for OOK demodulation in ultra-low power radios
abstract
A new technique for performing energy detection for ultra-low power OOK-modulated radio receivers is presented. This technique uses oversampling and digital averaging to reduce the excess noise contributions of wideband gain following the channel selection stage in these radios and helps to bring the noise performance closer to the expected value from communications theory. The proposed sampler effectively removes 10 dB of excess noise from such systems, bringing the minimum required IF SNR to achieve 10-3 bit error rate down as low as 0 dB. In addition, the oversampling technique provides added resolution in time to perform synchronization and clock extraction for the rest of the radio system. The energy detector was implemented in a 65 nm digital CMOS process, and consumes less than 10 μW under its nominal operating conditions providing a 200 kbps data rate.
Jesse Richmond, Jan M. Rabaey
ISCAS2
2010 EDA challenges and options: investing for the future
abstract
As the overall economy and semiconductor industry emerges from one of the worst recessions in years, it is time to take stock of EDA challenges and its future. This panel will focus on which challenges will surge and dominate EDA over the course of next several years and which challenges we can sell short.
Ruchir Puri, William H. Joyner, Raj Jammy, Ahmed Jerraya, Jan M. Rabaey, Walden C. Rhines, Leon Stok
DAC5
2010 Always energy-optimal microscopic wireless systems
abstract
Wireless sensor-and-control systems era starting to make inroads in medical health care. Monitoring of vital signs, advanced imaging, and innovative treatments and prosthetics are emerging. The challenge with all these systems is that they have to be ¿energy-frugal¿, that is that they have to get by with the energy they scavenge from the environment around them. At the same time, they need to be adaptive to the time-varying needs of the systems they observe or control. In this talk we will discuss how we could simultaneously explore the lower bounds of energy dissipation while at the same time providing ¿hugely-scalable¿ performance.
Jan M. Rabaey
DATE1
2010 A 2.2mW CMOS LNA for 6-8.5GHz UWB receivers
abstract
This paper presents an ultra-wideband (UWB) low noise amplifier (LNA) consuming 2.2-mW core dc power for 6-8.5GHz wireless applications. A common-gate input stage is cascaded with a common-source second stage to perform input impedance matching and wideband stagger-tuning amplification, while the current-reuse topology minimizes the dc power dissipation. The design method used to achieve flat-gain response is presented. A detailed analysis gives insight into the issue on the input impedance and suggests a solution. Implemented in a 90-nm CMOS process, the measurement results show power gain of 13.35+/-0.55 dB, input third intercept point (IIP3) of-6.2 dBm, and noise figure of 5-6.5 dB. The silicon die with 0.22-mm2active area allows the design to be adopted for highly integrated low-cost CMOS applications.
Chang-Ching Wu, Xuening Sun, Alberto L. Sangiovanni-Vincentelli, Jan M. Rabaey
ISCAS4
2010 Ultralow-Power Design in Near-Threshold Region
abstract
Operation in the subthreshold region most often is synonymous to minimum-energy operation. Yet, the penalty in performance is huge. In this paper, we explore how design in the moderate inversion region helps to recover some of that lost performance, while staying quite close to the minimum-energy point. An energy-delay modeling framework that extends over the weak, moderate, and strong inversion regions is developed. The impact of activity and design parameters such as supply voltage and transistor sizing on the energy and performance in this operational region is derived. The quantitative benefits of operating in near-threshold region are established using some simple examples. The paper shows that a 20% increase in energy from the minimum-energy point gives back ten times in performance. Based on these observations, a pass-transistor based logic family that excels in this operational region is introduced. The logic family operates most of its logic in the above-threshold mode (using low-threshold transistors), yet containing leakage to only those in subthreshold. Operation below minimum-energy point of CMOS is demonstrated. In leakage-dominated ultralow-power designs, time-multiplexing will be shown to yield not only area, but also energy reduction due to lower leakage. Finally, the paper demonstrates the use of ultralow-power design techniques in chip synthesis.
Dejan Markovic, Cheng C. Wang, Louis P. Alarcón, Tsung-Te Liu, Jan M. Rabaey
Proc. IEEE5
2009 EDA in flux: should I stay or should I go?
abstract
A crisis is a terrible thing to waste. Quotes like this are often heard by experts in the industry and academia but what does this mean to me? How should I change my professional interests? How should I evolve my career? How is EDA going to evolve? The panel represents multiple points of views on these questions. Four experts will review the current job situation in EDA and provide a historical perspective for previous recessions. The panel will further discuss views on new directions for the Electronics market, how EDA should evolve, and options for career development during the slow down of the industry.
Eshel Haritan, Andreas Kuehlmann, Tina Jones, John Epperheimer, Jan M. Rabaey, Rahul Razdan, Naveen Gupta
DAC5
2009 A 13.2 mW 1.9 GHz Interpolative BAW-based VCO for Miniaturized RF Frequency Synthesis
abstract
This paper presents a continuous frequency tuned bulk acoustic wave (BAW) resonator based oscillator without varactors. The frequency of oscillation is derived by interpolating between the resonant frequencies of two different resonators. Changing the interpolation factors allows a continuous tuning of the oscillation frequency between these two resonant frequencies. A proof of concept circuit was implemented in 0.4 mum CMOS and achieved a measured tuning range of more than 3 MHz with a phase noise between -134 and -109.5 dBc/Hz at 1 MHz offset while running at 3.3 V and consuming 4 mA of current. Additionally, an extended version of this circuit utilizing a bank of resonators to enhance the tuning range while keeping the power constant is introduced.
Michael Mark, Jan M. Rabaey
ISCAS2
2009 Asynchronous Computing in Sense Amplifier-Based Pass Transistor Logic
abstract
This paper presents the design and implementation of a low-energy asynchronous logic topology using sense amplifier-based pass transistor logic (SAPTL). The SAPTL structure can realize very low energy computation by using low-leakage pass transistor networks at low supply voltages. The introduction of asynchronous operation in SAPTL further improves energy-delay performance without a significant increase in hardware complexity. We show two different self-timed approaches: 1) the bundled data and 2) the dual-rail handshaking protocol. The proposed self-timed SAPTL architectures provide robust and efficient asynchronous computation using a glitch-free protocol to avoid possible dynamic timing hazards. Simulation and measurement results show that the self-timed SAPTL with dual-rail protocol exhibits energy-delay characteristics better than synchronous and bundled data self-timed approaches in 90-nm CMOS.
Tsung-Te Liu, Louis P. Alarcón, Matthew D. Pierson, Jan M. Rabaey
IEEE Trans. Very Large Scale Integr. Syst.4
2008 A brand new wireless day
abstract
Summary form only given. The wireless communications field has experienced a truly amazing growth since the early 1990's. Wireless connectivity slowly but surely has become pervasive. One would expect that by now this revolution must be losing some steam, but the truth is far from that. If anything, it is gathering even more speed. In the coming decades, introduction of innovative wireless technologies will enable a broad range of exciting applications to come to fruition, and reshape the way we interact with our daily living environment. Underlying it all is a three-tiered environment consisting of a large number of huge data and compute centers, billions of mobile compute and computation devices, and potentially trillions of tiny sensors and actuators. Making this happen will require some important wireless roadblocks to be either overcome or circumvented. A short list of those includes spectrum scarcity, reliability, complexity, security and obviously power. In this presentation, a number of innovative and even revolutionary solutions to address these will be discussed. Examples are collaborative cognitive networks, wireless in the mm-wave region of the spectrum, and miniature wireless. Each of these approaches pushes some part of the design technology to its limits, and may even require a totally novel approach towards design, all this while semiconductor technology is trying to cope with the uncertainty of design in the nanometer regime. One thing is for sure - the wireless designer of the next decade is bound for some very exciting times.
Jan M. Rabaey
ASP-DAC1
2008 Content Management and Replication in the SNSP: A Distributed Service-Based OS for Sensor Networks
abstract
This paper presents data management and replication techniques used by the SNSP, a distributed operating system for sensor networks. Three replication schemes, a deterministic, a distributed probabilistic and an adaptive observation based scheme are compared. The first two were adapted from related work and the latter was developed for the SNSP. Results indicate that the distributed scheme has the best performance in terms of total cost per data access as well as increased data availability. When access cost without overhead is considered, the adaptive algorithm performs the best, but its overhead is higher because data items are replicated independently of accesses.
Jana van Greunen, Jan M. Rabaey
CCNC2
2008 PicoCube: a 1 cm3 sensor node powered by harvested energy
abstract
The PicoCube is a 1 cm3 sensor node using harvested energy as its source of power. Operating at an average of only 6uW for a tirepressure application, the PicoCube represents a modular and integrated approach to the design of nodes for wireless sensor networks. It combines advanced ultra-low power circuit techniques with system-level power management. A simple packaging approach allows the modules comprising the node to fit into 1 cm3 in a reliable fashion.
Yuen-Hui Chee, Mike Koplow, Michael Mark, Nathan Pletcher, Michael D. Seeman, Fred L. Burghardt, Dan Steingart, Jan M. Rabaey, Paul K. Wright, Seth R. Sanders
DAC8
2008 Next generation wireless-multimedia devices: who is up for the challenge?
abstract
Yesterday's cell phones have rapidly evolved into versatile multi-media computers heavily loaded with a wide spectrum of technologies to support many functions and use modes. Designing and verifying such complex devices becomes increasingly challenging due to the need to: incorporate larger number of functions and diverse use modes, increased bandwidth, provide support for multimedia devices with high resolution, offer new methods for user interactions, and all of this at a fixed power budget and with high reliability. In addition, teams need to keep up with the convergence of wireless radio algorithms and be able to support more functionality moving into the software layer.
Juan C. Rey, Andreas Kuehlmann, Jan M. Rabaey, Cormac Conroy, Ted Vucurevich, Ikuya Kawasaki, Tuna B. Tarim
DAC3
2008 More Moore: foolish, feasible, or fundamentally different?
abstract
Moore’s law has been a foundation of modern electronics, sustained primarily by scaling. But can this continue despite the serious problems of litho, variability, device physics, and cost? This panel looks at several possibilities. Perhaps Moore’s law will muddle through, as it has so far, with a combination of tools, process, and design. But even if technically possible, Moore’s law is in practice driven by economics, and economics might turn against further scaling. Also, we’ve all seen how performance of single cores has topped out, despite scaling. Might this be a fundamental problem with planar technologies, prompting the need to go 3-D to get further performance increases? Or might CMOS itself give way to other technologies, allowing Moore’s law yet another respite? Compare and contrast for yourself these four very different visions of the future of your job, your industry, and your personal gadgets.
Robert C. Aitken, Jerry Bautista, Wojciech Maly, Jan M. Rabaey
ICCAD4
2008 Computing at the Crossroads (And What Does it Mean to Verification and Test?)
abstract
Today, we are interpreting computation as the execution of complex algorithms that are executed in sequential fashion and are bound to deliver deterministic answers. A number of factors are conspiring to fundamentally change that model. First, with scaling of technology to the nanoscale dimensions, it is quite certain that the underlying hardware platform will be all but deterministic (given effects such as variability and error susceptibility). Second, the emergence of distributed computation impacts the type of algorithms that are favored. Finally, many of the interesting problems to be tackled lay in the domain of perception and cognition, and most of these challenges tend to be statistical in nature. It is hence quite plausible that the nature of computation and the underlying hardware platforms will become statistical. These trends will have a profound effect on the way we verify and test designs. It is paramount that we start to explore what all of this means today if we want to be prepared for what tomorrow will bring.
Jan M. Rabaey
ITC1
2008 Analysis of Interference Effects in MB-OFDM UWB Systems
abstract
The inter-modulation and cross-modulation products of interferences introduced by nonlinearities of the receiver could significantly degrade system performance, and should be properly estimated when determining system design specifications. In MB-OFDM UWB systems, the traditional two- tone technique is still widely used to estimate nonlinear effects. However, this technique is not accurate enough. In this paper, we analyze the interference effects and propose a statistical approach to estimate the interference distortion products accurately. The analytical expressions for various interference scenarios are derived and then validated by simulation. Based on this analysis, we demonstrate how to adjust the two-tone technique to provide accurate distortion estimations for MB-OFDM UWB systems.
Jan M. Rabaey, Alberto L. Sangiovanni-Vincentelli
WCNC2
2007 Design without Borders - A Tribute to the Legacy of A. Richard Newton
Jan M. Rabaey
DAC1
2007 Design Without Borders
abstract
Summary form only given. Electrical engineers have learned how to build amazingly complex systems by assembling transistors, wires, and passive components into intricate networks. While solidly founded in semiconductor physics, pure engineering has made possible the design of multi-billion transistor chips in a repetitive, reliable and cost-effective way. A comprehensive "design methodology" was developed based on modularization, hierarchy and abstraction. Today this story is repeating itself. Physicists, chemists and biologists are exploring entirely different components such as molecules, atoms, and enzymes. Systems built from those will most probably impact our lives and society in a profound way. Outcomes will influence the ways we build mechanical structures, do computing, make drugs, generate energy and take care of our environment. Yet, while the basic components are dramatically different from our silicon devices, the basic strategy for building very complex systems from them remains unchanged. The art of design, as was developed in the silicon era, is just as applicable to these nano- or bio-constructions. Design methodology is a legacy that will live long after Moore's law has come to a halt. To quote the late Richard Newton, "The Future is BDA (Bio Design Automation) ".
Jan M. Rabaey
DSD1
2007 Short Distance Wireless, Dense Networks, and Their Opportunities
abstract
Summary form only given. The availability of wireless transceivers transmitting over ranges from few microns to less than half a meter opens the door for a wide range of exciting new applications, ranging from seamless system assembly, smart surfaces, healthcare monitoring and intelligent machinery and components. However, the implementation challenges in terms of size and power for most of these applications are pushing the limits. Fortunately, by exploring the wide range of options offered to the designer, extremely small and virtually zero-power transceivers are feasible. This paper discusses the opportunities, challenges and options of short distance wireless, and illustrates the proposed techniques with several design examples. In addition, the challenges that emerge when trying to embed these nodes into very dense networks are explored. Special consideration is given to the issues of distributed synchronization, localization and robust communication.
Jan M. Rabaey, Yuen-Hui Chee, Luca De Nardis, Simone Gambini, Davide Guermandi, Michael Mark, Nathan Pletcher
DSD1
2007 Fundamental Redundancy Versus Power Trade-Off in Standby SRAM
abstract
We study the problem of reducing power during data-retention in a standby static random access memory (SRAM). For successful data-retention, the supply voltage of an SRAM cell should be greater than a critical data retention voltage (DRV). Due to circuit parameter variations, the DRV for different cells on the same chip exhibits variation with a distribution having diminishing tail. For reliable data retention, the existing low-power design uses a worst-case technique in which a standby supply voltage that is larger than the highest DRV among all cells in an SRAM is used. Instead, our approach uses aggressive voltage reduction and counters the ensuing unreliability through a fault-tolerant memory architecture. The main results of this work are as follows: (i) We establish fundamental bounds on the power reduction in terms of the DRV-distribution using techniques from information theory. For the DRV-distribution of test-chip in (Qin, H, et al., 2006), we show that 49% power reduction with respect to (w.r.t.) the worst-case is a fundamental lower bound while 40% power reduction w.r.t. the worst-case is achievable with a practical combinatorial scheme, (ii) We study the power reduction as a function of the block-length for low-latency codes since most applications using SRAM are latency constrained. We propose a reliable memory architecture based on the Hamming code for the next test-chip implementation with a predicted power reduction of 33% while accounting for coding overheads.
Animesh Kumar, Huifang Qin, Prakash Ishwar, Jan M. Rabaey, Kannan Ramchandran
ICASSP (2)4
2007 Fundamental Bounds on Power Reduction during Data-Retention in Standby SRAM
abstract
The authors study leakage-power reduction in standby random access memories (SRAMs) during data-retention. An SRAM cell requires a minimum critical supply voltage (DRV) above which it preserves the stored-bit reliably. Due to process-variations, the intra-chip DRV exhibits variation with a distribution having a diminishing tail. In order to minimize leakage power while preserving data reliably, existing low-power design methods use a worst-case standby supply voltage. This worst-case voltage is larger than the highest DRV among all cells in an SRAM. In contrast, the approach uses aggressive voltage reduction and counters the ensuing unreliability by an error-control code based memory architecture. Using this approach, we explore fundamental trade-offs between power reduction and redundancy present in the SRAM. The authors establish fundamental bounds on the power reduction in terms of the DRV-distribution using techniques from information theory and algebraic coding theory. For an experimental test-chip DRV-distribution in the 90nm CMOS technology, the authors show that 49% power reduction with respect to (w.r.t.) the worst-case is a fundamental lower bound while 40% power reduction w.r.t. the worst-case is achievable by using a practical algebraic coding scheme. The authors also study the power reduction as a function of the block-length for low-latency codes since most applications using SRAM are latency constrained. The authors propose a reliable low-power memory architecture based on the Hamming code for the next test-chip implementation with a predicted power reduction of 33% while accounting for coding overheads
Animesh Kumar, Huifang Qin, Prakash Ishwar, Jan M. Rabaey, Kannan Ramchandran
ISCAS4
2007 Beyond Sensor Networks: ZUMA Middleware
abstract
Most wireless sensor network (WSN) proposals are unrelated to multimedia and other smart home infrastructures. This paper combines these issues and presents a solution where the Kilavi sensor network platform can be used as middleware to provide information for ZUMA. Kilavi is a centralized protocol designed for communication between high capacity management point and low capacity sensor nodes in building environment. ZUMA is a centralized future smart-home platform that interconnects all kinds devices in the home environment. This paper mainly focuses on sensor network operation issues. The discussion is on the benefits that the centralized architecture can provide in typical sensor network operations that are node discovery, routing and data dissemination, power control, and security. To support the discussion, the paper gives numbers that illustrate the benefits that centralization has over typical decentralized WSN solutions in home environment.
Mikael N. K. Soini, Jana van Greunen, Jan M. Rabaey, Lauri Sydänheimo
WCNC3
2006 Is "Network" the next "Big Idea" in design?
abstract
As the complexity of nowadays systems continues to grow, we are moving away from creating individual components from scratch, toward methodologies that emphasize composition of re-usable components via the network paradigm. Complex component interactions can create a range of amazing behaviors, some useful, some unwanted, some even dangerous. To manage them, a "science" for network design is evolving, applicable in some surprising areas. In this paper, we consider a few application domains and discus the design challenges involved from a methodology standpoint. From large-scale hardware/software systems, to dynamically adaptive sensor networks, and network-on-chip architectures, these ideas find wide application
Radu Marculescu, Jan M. Rabaey, Alberto L. Sangiovanni-Vincentelli
DATE2
2006 Research accelerator for multiple processors
David A. Patterson 0001, Arvind 0001, Krste Asanovic, Derek Chiou, James C. Hoe, Christoforos E. Kozyrakis, Shih-Lien Lu, Mark Oskin, Jan M. Rabaey, John Wawrzynek
Hot Chips Symposium9
2006 Wireless in the home - opportunities and challenges
Jan M. Rabaey
Hot Chips Symposium1
2006 An RF ToF Based Ranging Implementation for Sensor Networks
abstract
Localization is important for self-configuring sensor networks. One of the core tasks of localization is ranging (estimating distances) to reference points. In this paper a radio signal based Time of Flight (ToF) measuring ranging system for wireless sensor networks is proposed, designed and prototyped. The prototype measurement error is within -0.5m to 2m while operating at 100Msps sampling rate and using a 50MHz signal in the 2.4 GHz ISM band. The system accuracy is limited by the sampling rate and can be linearly improved with increasing rates. This RF method is more cost effective than acoustic signal based ranging schemes, as it does not require ultrasonic transducers. The system is multipath resilient and can coexist with 2.4Ghz band devices such as 802.11b/g networks.
Tufan C. Karalar, Jan M. Rabaey
ICC2
2006 The Energy-per-Useful-Bit Metric for Evaluating and Optimizing Sensor Network Physical Layers
abstract
To become truly ubiquitous, sensor network nodes must achieve ultra low power consumption. This paper proposes the energy-per-useful-bit (EPUB) metric for evaluating and comparing sensor network physical layers. EPUB includes the energy consumption of both the transmitter and receiver, and amortizes the energy consumption during the synchronization preamble over the number of data bits in the packet. Using EPUB, we compare six existing sensor network PHYs. Next, we optimize the PHY according to EPUB. We conclude that the EPUB of sensor network PHYs can be reduced by increasing data rate, lowering carrier frequency, and using simple modulation schemes such as OOK to reduce synchronization overhead
M. Josie Ammer, Jan M. Rabaey
SECON2
2006 L. Embedding Mixed-Signal Design in Systems-on-Chip
abstract
With semiconductor technology feature size scaling below 100 nm, mixed-signal design faces some important challenges, caused among others by reduced supply voltages, process variation, and declining intrinsic device gains. Addressing these challenges requires innovative solutions, at the technology, circuit, architecture, and design-methodology level. We present some of these solutions, including a structured platform-based design methodology to enable a meaningful exploration of the broad design space and to classify potential solutions in terms of the relevant metrics.
Jan M. Rabaey, Fernando De Bernardinis, Ali M. Niknejad, Borivoje Nikolic, Alberto L. Sangiovanni-Vincentelli
Proc. IEEE1
2006 Overcoming untuned radios in wireless networks with network coding
abstract
The drive toward the implementation and massive deployment of wireless sensor networks calls for ultralow-cost and low-power nodes. While the digital subsystems of the nodes are still following Moore's Law, there is no such trend regarding the performance of analog components. This work proposes a fully integrated architecture of both digital and analog components (including local oscillator) that offers significant reduction in cost, size, and overall power consumption of the node. Even though such a radical architecture cannot offer the reliable tuning of standard designs, it is shown that by using random network coding, a dense network of such nodes can achieve throughput linear in the number of channels available for communication. Moreover, the ratio of the achievable throughput of the untuned network to the throughput of a tuned network with perfect coordination is shown to be close to 1/e. This work uses network coding to leverage the fact that throughput equal to the max-flow in a graph is achievable even if the topology is not know a priori. However, the challenge here is finding the max-flow of the random graph corresponding to the network.
Dragan Petrovic, Kannan Ramchandran, Jan M. Rabaey
IEEE Trans. Inf. Theory3
2005 Are we ready for system-level synthesis?
abstract
Electronic system-level (ESL) design automation has been identified by Dataquest as the next productivity boost for the semiconductor industry. We have put together a distinguished panel of experts to discuss if we are ready for system-level synthesis.
Jason Cong, Tony Ma, Ivo Bolsens, Phil Moorby, Jan M. Rabaey, John Sanguinetti, Kazutoshi Wakabayashi, Yoshi Watanabe
ASP-DAC5
2005 Design at the end of the silicon roadmap
abstract
Scaling of silicon integrated technology into the deep sub-100 nm space brings with it a number of formidable challenges to the designer. Issues such as design complexity, power dissipation, process variability and reliability are challenging the traditional design methodologies. In this presentation, it is conjectured that the only viable long-term solution to these challenges is to drastically revise the way we do design, and a roadmap of potential solutions is presented. Ultimately, these innovative design solutions will help to pave the way to the post-silicon era.
Jan M. Rabaey
ASP-DAC1
2005 Wireless platforms: GOPS for cents and MilliWatts
abstract
In recent years, data communication has overtaken voice as the main force behind the growth in wireless. With this has come a proliferation of standards ranging from wide area networks at one end of the spectrum to personal area networks on the other end. The opportunities offered by this truly ubiquitous connectivity are tremendous, and are leading to revolutionary chances in the way computer, communication, and consumer systems operate and interact.Providing the necessary flexibility to seamlessly interact with the multitude of emerging network models, as well as the muscle to support the demanding multimedia functionality in a mobile environment, presents some huge challenges to the developer of the wireless implementation platforms. The power budget of the mobile terminal is typically fixed by size considerations and operation time. Cost considerations further constrain the solution space.In response to these challenges, many solutions have been floated and experimented with ranging from multi-processor architectures, advanced DSPs, reconfigurable solutions and hardwired accelerators. While these innovations break new ground in the world of embedded architectures, many questions emerge such as efficiency, flexibility and programming model.This panel will presents a "bake-off" between a number of solutions that have emerged over the recent years.
Francine Bacchini, Jan M. Rabaey, Allan Cox, Frank Lane, Rudy Lauwereins, Ulrich Ramacher, David Witt
DAC2
2005 Receiver initiated rendezvous schemes for sensor networks
abstract
Power efficiency is one of the most critical factors for wireless sensor networks. In this paper, we present a family of receiver-initiated pseudo-asynchronous rendezvous schemes designed specifically to reduce sensor nodes' power consumption. We analyze their performances in terms of power efficiency and latency, under different static and dynamic system parameters. In addition, we verify our analysis by extensive network simulations. Based on the results, our proposed scheme demonstrates superior performance compared to previous proposed schemes
En-Yi A. Lin, Jan M. Rabaey, Sven Wiethölter, Adam Wolisz
GLOBECOM2
2005 On the performance of geographical routing in the presence of localization errors [ad hoc network applications]
abstract
In this paper, a detailed study of the performance of geographic routing protocols in the presence of localization errors is carried out. Both analytical and simulation results illustrate the major impact or localization errors on the protocol goodput and route discovery energy. The performance metrics observed were the packet delivery ratio and the power consumed at a node for routing. It is shown that significant performance deterioration occurs with location errors as low as 20% of a node's radio range with no other obstacles in the network. To counteract this degradation, an enhancement is proposed that increases the error tolerance to about 40% radio range and in addition improves the performance consistently for any location error. Furthermore, the effect of obstacles in conjunction with location errors on the routing performance is also investigated.
Rahul C. Shah, Adam Wolisz, Jan M. Rabaey
ICC3
2005 Low power synchronization for wireless sensor network modems
abstract
A simplification of synchronization requirements is necessary to meet the low power goals of wireless sensor networks (WSNs). Simply scaling existing synchronization systems to WSN data rates consumes more power than the entire allotted node budget. In this work, simplification of synchronization requirements was achieved through system level design of modulation scheme, data rate, packet length, and clock accuracy. Analog and digital implementations of the resulting synchronization functions are explored. A digital implementation was chosen because it requires a smaller header length and therefore reduces system energy consumption. The chosen scheme achieves better than 1e/sup -4/ BER at 13 dB SNR, an implementation loss of 1.6 dB over the ideal detection case, while consuming under 300 /spl mu/W.
M. Josie Ammer, Jan M. Rabaey
WCNC2
2005 Does proper coding make single hop wireless sensor networks reality: the power consumption perspective
abstract
The common belief is that a multi-hop configuration with rather small per-hop distance is the only viable energy-efficient option for wireless sensor networks. We discuss a single hop configuration, utilizing the asymmetry between lightweight sensor nodes and a more powerful "base station" and demonstrate that such a single hop configuration can actually have lower overall power consumption than a multi-hop counterpart.
Lizhi C. Zhong, Jan M. Rabaey, Adam Wolisz
WCNC2
2005 Modeling and Analysis of Opportunistic Routing in Low Traffic Scenarios
abstract
Opportunistic routing protocols have been proposed as efficient methods to exploit the high node densities in sensor networks to mitigate the effect of varying channel conditions and non-availability of nodes that power down periodically. They work by integrating the network and data link layers so that they can take a joint decision as to the next hop forwarding node based on its availability and suitability as a forwarder. This cross-layer integration makes it harder to optimize the protocol due to the dependencies among the different components of the protocol stack. In this paper, we provide a framework to model opportunistic routing that breaks up the functionality into three separate components and simplifies analysis. The framework is used to model two variants of opportunistic routing and is shown to match well with simulation results. In addition, using the model for performance analysis yields important guidelines for the future design of such protocols.
Rahul C. Shah, Sven Wiethölter, Adam Wolisz, Jan M. Rabaey
WiOpt4
2004 Adaptive sleep discipline for energy conservation and robustness in dense sensor networks
abstract
This paper presents an adaptive approach for conserving energy in high-density sensor networks. The proposed method allows sensor nodes to sleep while ensuring that application performance constraints are met. Unlike deterministic algorithms that assume static connectivity, the approach uses a randomized algorithm to provide robustness to the variations in network connectivity. These variations are due to fading channels, depletion or addition of nodes, node mobility, and the sleeping of nodes. The algorithm developed in the paper is extremely lightweight and does not require nodes to keep any state information about their individual neighbors. Based on local observations, each node independently decides when to sleep and wakeup. The algorithm also ensures an evenly distributed workload among nodes, and achieves energy savings proportional to the density of nodes.
Jana van Greunen, Dragan Petrovic, Alvise Bonivento, Jan M. Rabaey, Kannan Ramchandran, Alberto L. Sangiovanni-Vincentelli
ICC4
2004 Power-efficient rendez-vous schemes for dense wireless sensor networks
abstract
We propose two generic rendez-vous schemes for dense wireless sensor networks, including the transmitter- and receiver-initiated cycled receiver schemes. Our studies include the analyses and comparisons of their power efficiencies, especially under a fading channel as a realistic physical layer. We believe our modeling strategies as well as the results are applicable to any rendez-vous scheme of the cycled- receiver nature. The paper further proposes precise guidelines and optimization strategies for synchronization in wireless sensor networks.
En-Yi A. Lin, Jan M. Rabaey, Adam Wolisz
ICC2
2004 An integrated data-link energy model for wireless sensor networks
abstract
In this paper we integrate the energy models of all the components of the data link layer of a wireless sensor network into a single framework, which interacts with the abstracted models of the network layer, the physical layer, and the channel. Such a framework enables intra layer and cross layer optimization. By identifying the important design parameters, it provides the venues and guidelines for energy' reduction and design improvement, as demonstrated in the paper. The validity of the models is verified using OMNET++ network simulations.
Lizhi C. Zhong, Jan M. Rabaey, Adam Wolisz
ICC2
2004 An Integrated, Low Power Localization System for Sensor Networks
abstract
Localization (a.k.a. locationing) is a central concern for ubiquitous self-configuring sensor networks. In this paper the implementation of a distributed, least-squares-based localization algorithm is presented. Low power and energy dissipation are key requirements for sensor networks. As part of the sensor network, the localization system must also conform to these requirements. An ultra-low-power and dedicated hardware implementation of the localization system is therefore presented. The cost of fixed-point implementation is also investigated. The design is implemented in a 0.13 /spl mu/ CMOS process. It dissipates 1.7 mW of active power and 0.122 J/op of active energy with a silicon area of 0.55 mm/sup 2/. The mean calculated location error due to fixed-point implementation is shown to be 6%.
Tufan C. Karalar, Shunzo Yamashita, Michael Sheets, Jan M. Rabaey
MobiQuitous4
2003 A low-energy chip-set for wireless intercom
abstract
A low power wireless intercom system is designed and implemented. Two fully-operational ASICs, integrating custom and commercial IP, implement the entire digital portion of the protocol stack. Combined, the chips consume 13 mW on average when three nodes are connected to the network. A high-level design methodology was used to define the protocol stack and communication algorithms, select architectures, and minimize energy.
M. Josie Ammer, Michael Sheets, Tufan C. Karalar, Mika Kuulusa, Jan M. Rabaey
DAC5
2003 Reshaping EDA for power
abstract
Today's rising power densities have been widely cited as the foremost challenge to continued CMOS scaling. In fact, the current power crisis is reminiscent of the final days previous technologies, such as the once popular bipolar and NMOS technologies and even vacuum tubes. How CMOS technology will respond to the current power challenge to extend CMOS scaling to sub-90nm technology is an important question for designer and CAD tool developers alike. With aggressive scaling a number of new challenges have arisen, such as leakage control, heat removal and power supply distribution, that need to be addressed using new design techniques in conjunction with new CAD solutions.This panel brings together experts in circuit design and CAD tool development to discuss the current status of low-power design and provide opinions on what new EDA capabilities are most important in the power-constrained design era. For instance, how will power be distributed in a robust fashion in sub-90nm ICs, and what are the critical EDA analysis and optimization capabilities? What are the best techniques for leakage reduction, not only in standby modes, but also in the active mode? And how far will voltage scaling take us in attacking the dynamic power consumption issue? What will a power-centric design flow look like and how will it change the way we design ICs? The objective of the panel is to explore these issues and formulate a list of critical issues that need to be addressed by the EDA community to enable successful scaling of CMOS into the sub-90nm era.
Jan M. Rabaey, Dennis Sylvester, David T. Blaauw, Kerry Bernstein, Jerry Frenkil, Mark Horowitz, Wolfgang Nebel, Takayasu Sakurai
DAC1
2003 Distributed algorithms for transmission power control in wireless sensor networks
abstract
In a wireless, multi-hop sensor network, choosing transmission power levels has an important impact on energy efficiency and network lifetime. Two algorithms for dynamically adjusting transmission power level on a per-node basis are proposed here. Network lifetime, convergence speed as well as resulting network connectivity are used as figures of merit for these two algorithms. They have been evaluated in an indoor sensor environment. The network lifetime metrics of these two local algorithms are also benchmarked against power control algorithms using global information. We show that these local algorithms outperform fixed power level assignment and that the lifetime achieved by them is usually within a factor of two of globally computed solution while being scalable.
Martin Kubisch, Holger Karl, Adam Wolisz, Lizhi C. Zhong, Jan M. Rabaey
WCNC5
2003 A study of low level vibrations as a power source for wireless sensor nodes
Shad Roundy, Paul K. Wright, Jan M. Rabaey
Comput. Commun.3
2002 What's the next EDA driver?
abstract
The PC industry was the major consumer of silicon in the 80's and 90's. It defined the requirements for EDA. In a world dominated by PC's, clock frequency was the ultimate measure of performance. Times have changed. Today the internet, wireless communications and consumer applications such as high end gaming stations have replaced the PC as the primary driver. How we measure 'cutting edge' has also changed. Raw computing speed measured in terms of clock speed has given way to alternate forms of measuring performance - MBits/s, MMACS, and Polygons/sec now capture the differing computation requirements for these domains. Depending on the application domain, weight and power are just as important if not more important than raw computational power. How do these changes drive technology development in EDA? Which of these is the dominant EDA driver? How much of this drive is a function of EDA following the money? How much of this drive is a function of unique technical challenges posed by these domains? This diverse set of panelists with representation from various application segments and the EDA industry will attempt to answer these questions.
Jan M. Rabaey, Joachim Kunkel, Dennis Brophy, Raúl Camposano, Davoud Samani, Larry Lerner, Rick Hetherington
DAC1
2002 Robust Positioning Algorithms for Distributed Ad-Hoc Wireless Sensor Networks
Chris Savarese, Jan M. Rabaey, Koen Langendoen
USENIX ATC, General Track2
2002 Energy aware routing for low energy ad hoc sensor networks
abstract
The recent interest in sensor networks has led to a number of routing schemes that use the limited resources available at sensor nodes more efficiently. These schemes typically try to find the minimum energy path to optimize energy usage at a node. In this paper we take the view that always using lowest energy paths may not be optimal from the point of view of network lifetime and long-term connectivity. To optimize these measures, we propose a new scheme called energy aware routing that uses sub-optimal paths occasionally to provide substantial gains. Simulation results are also presented that show increase in network lifetimes of up to 40% over comparable schemes like directed diffusion routing. Nodes also burn energy in a more equitable way across the network ensuring a more graceful degradation of service with time.
Rahul C. Shah, Jan M. Rabaey
WCNC2
2001 Addressing the System-on-a-Chip Interconnect Woes Through Communication-Based Design
abstract
Communication-based design represents a formal method approach to of system-on-a-chip design that considers communication between components as important as the computations they perform. “Our network-on-chip&rdqo ; approach partitions the communication into layers to maximize reuse and provide a programmer with an abstraction of the underlying communication framework. This layered approach is cast in the structure advocated by the OSI Reference network model and is demonstrated with a reconfigurable DSP example. The Metropolis methodology of deriving layers through a sequence of successive adaptation steps between incompatible behaviors refinement of communication is illustrated through the Intercom a design example. In another approach, MESCAL provides a designer with tools for a correct-by-construction protocol stack.
Marco Sgroi, Michael Sheets, Andrew Mihal, Kurt Keutzer, Sharad Malik, Jan M. Rabaey, Alberto L. Sangiovanni-Vincentelli
DAC6
2001 Design methodology for PicoRadio networks
abstract
One of the most compelling challenges of the next decade is the "last-meter" problem, extending the expanding data network into end-user data-collection and monitoring devices. PicoRadio supports the assembly of an ad hoc wireless network of self-contained mesoscale, low-cost, low-energy sensor and monitor nodes. While technology advances have made it conceivable to deploy wireless networks of heterogeneous nodes, the design of a low-power, low-cost, adaptive node in a reduced time to market is still a challenge. We present a design methodology for PicoRadio Networks, from system conception and optimization to silicon platform implementation. For each phase of the design, we demonstrate the applicability of our methodology through promising experimental results.
Julio Leao da Silva Jr., J. Shamberger, M. Josie Ammer, Chunlong Guo, Suet-Fei Li, Rahul C. Shah, Tim Tuan, Michael Sheets, Jan M. Rabaey, Borivoje Nikolic, Alberto L. Sangiovanni-Vincentelli, Paul K. Wright
DATE9
2001 Low power distributed MAC for ad hoc sensor radio networks
abstract
Targeted at multi-hop wireless sensor networks, a set of low power MAC design principles have been proposed, and a novel ultra-low power MAC is designed to be distributed in nature to support scalability, survivability and adaptability requirements. Simple CSMA and spread spectrum techniques are combined to trade off bandwidth and power efficiency. A distributed algorithm is used to do dynamic channel assignment. A novel wake-up radio scheme is incorporated to take advantage of new radio technologies. The notion of mobility awareness is introduced into an adaptive protocol to reduce network maintenance overhead. The resulting protocol shows much higher power efficiency for typical sensor network applications.
Chunlong Guo, Lizhi C. Zhong, Jan M. Rabaey
GLOBECOM3
2001 Location in distributed ad-hoc wireless sensor networks
abstract
Evolving networks of ad-hoc wireless sensing nodes rely heavily on the ability to establish position information. The algorithms presented herein rely on range measurements between pairs of nodes and the a priori coordinates of sparsely located anchor nodes. Clusters of nodes surrounding anchor nodes cooperatively establish confident position estimates through assumptions, checks, and iterative refinements. Once established, these positions are propagated to more distant nodes, allowing the entire network to create an accurate map of itself. Major obstacles include overcoming inaccuracies in range measurements as great as /spl plusmn/50%, as well as the development of initial guesses for node locations in clusters with few or no anchor nodes. Solutions to these problems are presented and discussed, using position error as the primary metric. Algorithms are compared according to position error, scalability, and communication and computational requirements. Early simulations yield average position errors of 5% in the presence of both range and initial position inaccuracies.
Chris Savarese, Jan M. Rabaey, Jan Beutel
ICASSP2
2001 Reconfigurable platform design for wireless protocol processors
abstract
Low-energy protocol processing is a crucial issue in next generation wireless systems. In modern wireless system design, this problem is tightly coupled with the signal processing needs. Fierce market competition and inventive wireless applications are imposing stricter design requirements in energy consumption, cost, size, and flexibility. To deal with these unique constraints, we incorporate the platform-based design methodology to deal with these constraints by advocating reusability. This paper presents this methodology, and its application on PicoRadio, a cutting-edge wireless system. In particular, we describe the design of a reconfigurable architecture optimized for protocol processing.
Tim Tuan, Suet-Fei Li, Jan M. Rabaey
ICASSP3
2001 Wireless beyond the third generation wireless beyond the third generation: facing the energy challenge
abstract
Article Wireless beyond the third generation wireless beyond the third generation: facing the energy challenge Share on Author: Jan M. Rabaey BWRC, EECS Department, University of California at Berkeley BWRC, EECS Department, University of California at BerkeleyView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 1–3https://doi.org/10.1145/383082.383084Online:06 August 2001Publication History 10citation587DownloadsMetricsTotal Citations10Total Downloads587Last 12 Months5Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jan M. Rabaey
ISLPED1
2001 Limitations and challenges of computer-aided design technology for CMOS VLSI
abstract
As manufacturing technology moves toward fundamental limits of silicon CMOS processing, the ability to reap the full potential of available transistors and interconnect is increasingly important. Design technology (DT) is concerned with the automated or semi-automated conception, synthesis, verification, and eventual testing of microelectronic systems. While manufacturing technology faces fundamental limits inherent in physical laws or material properties, design technology faces fundamental limitations inherent in the computational intractability of design optimizations and in the broad and unknown range of potential applications within various design processes. In this paper, we explore limitations to how design technology can enable the implementation of single-chip microelectronic systems that take full advantage of manufacturing technology with respect to such criteria as layout density performance, and power dissipation.
Randal E. Bryant, Kwang-Ting Cheng, Andrew B. Kahng, Kurt Keutzer, Wojciech Maly, A. Richard Newton, Lawrence T. Pileggi, Jan M. Rabaey, Alberto L. Sangiovanni-Vincentelli
Proc. IEEE8
2000 Retargetable estimation scheme for DSP architecture selection
abstract
Abstract — Given the recent wave of innovation and diversification in digital signal processor (DSP) architecture, the need for quickly evaluating the true potential of considered architectural choices for a given application has been rising. We propose a new scheme, called Retargetable Estimation, that involves analysis of a high-level description of a DSP application, with aggressive optimization search, to provide a performance estimate of its optimal implementation on the architectures considered. With this scheme, we present a new parameterized architecture model that allows quick retargeting to a wide range of architectural choices, and that emphasizes capturing an architecture's salient optimizing features. We show that for a set of DSP benchmarks and two full applications, hand-optimized performance can be predicted reliably. We applied this scheme to two different processors. I.
Naji Ghazal, A. Richard Newton, Jan M. Rabaey
ASP-DAC3
2000 Low-power silicon architecture for wireless communications: embedded tutorial
abstract
Wireless communication and networking is experiencing a dramatic growth, and all indicators point to an extension of this growth in the foreseeable future.This paper reflects on the demands and the opportunities offered with respect to the integrated implementation of these applications in the "systems-ona-chip" era.
Jan M. Rabaey
ASP-DAC1
2000 Predicting performance potential of modern DSPs
abstract
High-level development tools for digital signal processors (DSPs) remain unable to extract optimal performance from them without the designer's in-depth knowledge of the architecture. In this paper we describe our approach to Retargetable Estimation and show how and why it can be effective in quickly predicting and guiding toward hand-optimized performance of moderns DSPs for a given application described in a high-level language. We also contrast the advantages of this scheme with those of a full-featured optimizing compiler.
Naji Ghazal, A. Richard Newton, Jan M. Rabaey
DAC3
2000 Designing wireless protocols: methodology and applications
abstract
Communication protocols are essential components of wireless systems. Present methods for protocol design are heuristic in nature and are not suited for next generation wireless systems where time-to-market concerns require correct-the-first-time implementations. In this paper we present a new design methodology for wireless protocols based on the principle of orthogonalization of concerns. In particular, the methodology separates function and architecture design and emphasizes the use of formal models to ensure correctness and reduce design time. Protocols are described using co-design finite state machines (CFSMs), a model of computation that has been introduced to allow the efficient capture of both the control and the data processing parts of the specification. Furthermore, algorithms for automatic hardware and software synthesis from CFSMs are available. This allows a fast exploration of different HW/SW partitions and the analysis of tradeoffs involved. Intercom, a mobile wireless system supporting full-duplex voice communication among different users, is presented and the design of its protocols is described. The design methodology presented here will be used for the design of PicoRadio, a low-power and highly adaptive network of sensors.
Marco Sgroi, Julio Leao da Silva Jr., Fernando De Bernardinis, Fred L. Burghardt, Alberto L. Sangiovanni-Vincentelli, Jan M. Rabaey
ICASSP6
2000 Challenges and Opportunities in Broadband and Wireless Communication Designs
abstract
Communication designs form the fastest growing segment of the semiconductor market. Both network processors and wireless chipsets have been attracting a great deal of research attention, financial resources and design efforts. However, further progress is limited by lack of adequate system methodologies and tools. Our goal in this paper is to provide impetus for development of communication design techniques and tools. The first part addresses network processors (NP) that we study from three viewpoints: application, architecture, and system software and compilation tools. In addition to summary of main issues and representative case studies, we identify main system design issues. The second part of the tutorial focuses on wireless design. The main emphasis is on platform-based design methodology that leverages on functional profiling, architecture exploration, and orthogonalization of concerns to facilitate low-power wireless communication systems. The highlight of the paper, an in-depth study of the state-of-the-art wireless design, PicoRadio, is used as explanatory design example.
Jan M. Rabaey, Miodrag Potkonjak, Farinaz Koushanfar, Suet-Fei Li, Tim Tuan
ICCAD1
2000 Processors for Mobile Applications
abstract
Mobile processors form a large and very fast growing segment of semiconductor market. Although they are used in a great variety of embedded systems such as personal digital organizers (PDAs), smart cards, internet appliances, laptops, smart badges, cellular phones, wearable computers, and sensor networks, they share the common need for low power, code density, security, cost sensitivity and multimedia and communication processing. The goal of this paper is to review the field of processors for mobile applications. We survey a spectrum of processors, their system software, and the accompanying hardware components. The emphasis is on classification and identification of major technology and architecture trends. Companion to this paper is a WWW page [Mob00] which provides comprehensive additional material about mobile processors.
Farinaz Koushanfar, Miodrag Potkonjak, Vandana Prabhu, Jan M. Rabaey
ICCD4
2000 MOS current mode logic for low power, low noise CORDIC computation in mixed-signal environments
abstract
In this work, MOS Current Mode Logic (MCML) is analyzed for application to low power, mixed signal environments. A small MCML cell library is developed and optimized for several different performance requirements. The cells are then applied to the generation of piplelined CORDIC structures and compared with equivalent CMOS circuits. MCML CORDICs are designed which can operate from 125MHz to 310MHz with power consumption varying between 4.3mW and 18.6mW. These power results are up to 1.5 times less than CMOS CORDICs with equivalent propagation delays. Design was done in a 0.25µm standard CMOS process from ST Microelectronics.
Jason M. Musicer, Jan M. Rabaey
ISLPED2
2000 System-level design: orthogonalization of concerns andplatform-based design
abstract
System-level design issues become critical as implementation technology evolves toward increasingly complex integrated circuits and the time-to-market pressure continues relentlessly. To cope with these issues, new methodologies that emphasize re-use at all levels of abstraction are a "must", and this is a major focus of our work in the Gigascale Silicon Research Center. We present some important concepts for system design that are likely to provide at least some of the gains in productivity postulated above. In particular, we focus on a method that separates parts of the design process and makes them nearly independent so that complexity could be mastered. In this domain, architecture-function co-design and communication-based design are introduced and motivated. Platforms are essential elements of this design paradigm. We define system platforms and we argue about their use and relevance. Then we present an application of the design methodology to the design of wireless systems. Finally, we present a new approach to platform-based design called modern embedded systems, compilers, architectures and languages, based on highly concurrent and software programmable architectures and associated design tools.
Kurt Keutzer, A. Richard Newton, Jan M. Rabaey, Alberto L. Sangiovanni-Vincentelli
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2000 Maximally and arbitrarily fast implementation of linear andfeedback linear computations
abstract
By establishing a relationship between the basic properties of linear computations and eight optimizing transformations (distributivity, associativity, commutativity, inverse and zero element law, common subexpression replication and elimination, constant propagation), a computer-aided design platform is developed to optimally speed-up an arbitrary instance from this large class of computations with respect to those transformations. Furthermore, arbitrarily fast implementation of an arbitrary linear computation is obtained by adding loop unrolling to the transformations set. During this process, a novel Horner pipelining scheme is used so that the area-time (AT) product is maintained constant, regardless of achieved speed-up. We also present a generalization of the new approach so that an important subclass of nonlinear computations, named feedback linear computations, is efficiently, maximally, and arbitrarily sped-up.
Miodrag Potkonjak, Jan M. Rabaey
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2000 Low-swing on-chip signaling techniques: effectiveness and robustness
abstract
This paper reviews a number of low-swing on-chip interconnect schemes and presents a thorough analysis of their effectiveness and limitations, especially on energy efficiency and signal integrity. In addition, several new interface circuits presenting even more energy savings and better reliability are proposed. Some of these circuits not only reduce the interconnect swing, but also use very low supply voltages so as to obtain quadratic energy savings. The performance of each of the presented circuits is thoroughly examined using simulation on a benchmark interconnect circuit. Significant energy savings up to a factor of six have been observed.
Hui Zhang 0008, George Varghese, Jan M. Rabaey
IEEE Trans. Very Large Scale Integr. Syst.3
1999 The design of a low energy FPGA
abstract
This work presents the design of an energy efllcient FPGA architecture.Significant reduction in the energy consumption is achieved by tackling both circuit design and architecture optimization issues concurrently.A hybrid interconnect structure incorporating Nearest Neighbor Connections, Symmetric Mesh Architecture, and Hierarchical connectivity is used.The energy of the interconnect is also reduced by employing low-swing circuit techniques.These techniques have been employed to design and fabricate an FPGA.Preliminary analysis show energy improvement of more than an order of magnitude when compared to existing commercial architectures. 1.1
George Varghese, Hui Zhang 0008, Jan M. Rabaey
ISLPED3
1999 Algorithm selection: a quantitative optimization-intensive approach
abstract
Implementation platform selection is an important component of hardware-software codesign process which selects, for a given computation, the most suitable implementation platform. In this paper, we study the complementary component of hardware-software codesign, algorithm selection. Given a set of specifications for the targeted application, algorithm selection refers to choosing the most suitable completely specified computational structure for a given set of design goals and constraints, among several functionally equivalent alternatives. While implementation platform selection has been recently widely and vigorously studied, the algorithm selection problem has not been studied in computer-aided design domain until now. We first introduce the algorithm selection problem, and then analyze and classify its degrees of freedom. Next, we demonstrate an extraordinary impact of algorithm selection for achieving high throughput, and low-cost implementations. We define the algorithm selection problem formally and prove that throughput and area optimization using algorithm selection are computationally intractable problems. We also propose a relaxation-based heuristic for throughput optimization. Finally, we present an algorithm for cost optimization using algorithm selection. The effectiveness of methodology and proposed algorithms is illustrated using real-life examples.
Miodrag Potkonjak, Jan M. Rabaey
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1998 A Methodology for Guided Behavioral-Level Optimization
abstract
Abstract — Optimization at the early stages of design are crucial. However, due to an overwhelming number of design and optimization options, design exploration is often conducted in a qualitative, ad-hoc manner. This paper presents a methodology and interactive environment for guiding the exploration process. A prototype targeting behavioral-level optimization for datapath-intensive ASIC implementations has been developed. The key to the approach is encapsulated knowledge about the various optimizations and a set of techniques to automatically extract the “essence ” of a design description. At each stage in the exploration process, the system suggests and ranks potential optimizations, both in terms of immediate and longer-term impact. It also provides evaluations of the design and of the likely affects each optimization will have on metrics like power and performance. In the new approach, the designer is responsible for making the actual optimization selections. However,
Lisa M. Guerra, Miodrag Potkonjak, Jan M. Rabaey
DAC3
1998 A Multiprocessor DSP System Using PADDI-2
abstract
We have integrated an image processing system built around PADDI-2, a custom 48 node MIMD parallel DSP. The system includes image processing algorithms, a graphical SFG tool, a simulator, routing tools, compilers, hardware configuration and debugging tools, application development libraries, and software implementations for hardware verification. The system board,connected to a SPARCstation via a custom Sbus controller, contains 384 processors in 8 VLSI chips. The software environment supports a multiprocessor system under development (VGI-1). The software tools and libraries are modular, with implementation dependencies isolated in layered encapsulations.
Roy A. Sutton, Vason P. Srini, Jan M. Rabaey
DAC3
1998 An Energy-Conscious Exploration Methodology for Reconfigurable DSPs
Jan M. Rabaey, Marlene Wan
DATE1
1998 Low-energy embedded FPGA structures
abstract
This paper introduces an energy-efficient FPGA module, intended for embedded implementations. The main features of the proposed cell include a rich local-interconnect network, which drastically reduces the energy dissipated in the wiring, and a dual-voltage scheme that allows pass-transistor networks to operate at low-voltages yet maintain decent performance. Simulations on a benchmark set demonstrate that the proposed module succeeds in its goal of reducing energy consumption by an order of magnitude over existing implementations.
Eric Kusse, Jan M. Rabaey
ISLPED2
1998 Low-swing interconnect interface circuits
abstract
This paper introduces an energy-efficient FPGA module, intended for embedded implementations. The main features of the proposed cell include a rich local-interconnect network, which drastically reduces the energy dissipated in the wiring, and a dual-voltage scheme that allows pass-transistor networks to operate at low-voltages yet maintain decent performance. Simulations on a benchmark set demonstrate that the proposed module succeeds in its goal of reducing energy consumption by an order of magnitude over existing implementations.
Hui Zhang 0008, Jan M. Rabaey
ISLPED2
1998 Behavioral-level synthesis of heterogeneous BISR reconfigurable ASIC's
abstract
In this paper, behavioral-level synthesis techniques are presented for the design of reconfigurable hardware. The techniques are applicable for synthesis of several classes of designs, including: (1) design for fault tolerance against permanent faults, (2) design for Improved manufacturability, and (3) design of application specific programmable processors (ASPPs)-processors designed to perform any computation from a specified set on a single implementation platform. This paper focuses on design techniques for efficient built-in self-repair (BISR), and thus directly addresses the former two applications. Previous BISR techniques have been based on replacing a failed module with a backup of the same type. We present new heterogeneous BISR methodologies which remove this constraint and enable replacement of a module with a spare of a different type. The approach is based on the flexibility of behavioral-level synthesis to explore the design space. Two behavioral synthesis techniques are developed; the first method is through assignment and scheduling, and the second utilizes transformations. Experimental results verify the effectiveness of the approaches.
Lisa M. Guerra, Miodrag Potkonjak, Jan M. Rabaey
IEEE Trans. Very Large Scale Integr. Syst.3
1997 A Dynamic Design Estimation and Exploration Environment
abstract
The rapid increase in the complexity of systems demands newapproaches to design exploration and trade-off analysis. Thispaper presents an exploration environment that provides a uniformway to perform exploration at the conceptual (pre-specification)stages of design. The environment encapsulates knowledge abouthow to perform estimations in different application domains, andseparates a designer from the mechanics of the estimation process.
Ole Bentz, Jan M. Rabaey, David Lidsky
DAC2
1997 Reconfigurable processing: the solution to low-power programmable DSP
abstract
One of the most compelling issues in the design of wireless communication components is to keep power dissipation between bounds. While low-power solutions are readily achieved in an application-specific approach, doing so in a programmable environment is a substantially harder problem. This paper presents an approach to low-power programmable DSP that is based on the dynamic reconfiguration of hardware modules. This technique has shown to yield at least an order of magnitude of power reduction compared to traditional instruction-based engines for problems in the area of wireless communication.
Jan M. Rabaey
ICASSP1
1997 System-level power estimation and optimization - challenges and perspectives
abstract
Article System-level power estimation and optimization—challenges and perspectives Share on Author: Jan M. Rabaey Department of EECS, University of California at Berkeley Department of EECS, University of California at BerkeleyView Profile Authors Info & Claims ISLPED '97: Proceedings of the 1997 international symposium on Low power electronics and designAugust 1997 Pages 158–160https://doi.org/10.1145/263272.263314Online:01 August 1997Publication History 3citation158DownloadsMetricsTotal Citations3Total Downloads158Last 12 Months1Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jan M. Rabaey
ISLPED1
1996 Early Power Exploration - A World Wide Web Application
abstract
Article Free Access Share on Early power exploration—a World Wide Web application Authors: David Lidsky University of California, Berkeley University of California, BerkeleyView Profile , Jan M. Rabaey University of California, Berkeley University of California, BerkeleyView Profile Authors Info & Claims DAC '96: Proceedings of the 33rd annual Design Automation ConferenceJune 1996 Pages 27–32https://doi.org/10.1145/240518.240523Published:01 June 1996Publication History 50citation367DownloadsMetricsTotal Citations50Total Downloads367Last 12 Months11Last 6 weeks5 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
David Lidsky, Jan M. Rabaey
DAC2
1996 Exploiting regularity for low-power design
abstract
Current day behavioral-synthesis techniques produce architectures that are power-inefficient in the interconnect. Experiments have demonstrated that in synthesized designs, about 10 to 40% of the total power may be dissipated in buses, multiplexors, and drivers. We present a novel approach targeted at the reduction of power dissipation in interconnect elements-buses, multiplexors, and buffers. The scheduling, assignment, and allocation techniques presented in this paper exploit the regularity and common computational patterns in the algorithm to reduce the fan-outs and fan-ins of the interconnect wires, resulting in reduced bus capacitances and a simplified interconnect structure. Average power savings of 47% and 49% in buses and multiplexors, respectively, are demonstrated on a set of benchmark examples.
Renu Mehra, Jan M. Rabaey
ICCAD2
1996 Which has greater potential power impact: high-level design and algorithms or innovative low power technology? (panel)
James Burr, Laszlo Gal, Ramsey W. Haddad, Jan M. Rabaey, Bruce Wooley
ISLPED4
1996 Performance optimization using template mapping for datapath-intensive high-level synthesis
abstract
This paper introduces a new approach to performance-driven template mapping for high-level synthesis. Template mapping, the process of mapping high-level algorithmic descriptions to specialized hardware libraries or instruction sets, involves template matching, template selection, and clock selection. Efficient algorithms for each are presented, and novel issues such as partial matching are addressed. The paper focuses on datapath-intensive ASIC design, though the concepts are also highly applicable to compiler development. Experimental results on examples from real applications show significant improvements in throughput with limited area overhead.
Miguel R. Corazao, Marwan A. Khalaf, Lisa M. Guerra, Miodrag Potkonjak, Jan M. Rabaey
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
1996 Activity-sensitive architectural power analysis
abstract
Prompted by demands for portability and low-cost packaging, the electronics industry has begun to view power consumption as a critical design criterion. As such there is a growing need for tools that can accurately predict power consumption early in the design process, many high-level power analysis models do not adequately model activity, however, leading to inaccurate results. This paper describes an activity-sensitive power analysis strategy for datapath, memory, control path, and interconnect elements. Since datapath and memory modeling has been described in a previous publication, this paper focuses mainly on a new Activity-Based Control (ABC) model and on a hierarchical interconnect analysis strategy that enables estimates of chip area as well as power consumption. Architecture-level estimates are compared to switch-level measurements based on net lists extracted from the layouts of three chips: a digital filter, a global controller, and a microprocessor. The average power estimation error is about 9% with a standard deviation of 10%, and the area estimates err on average by 14% with a standard deviation of 6%.
Paul E. Landman, Jan M. Rabaey
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1995 Power minimization in DSP application specific systems using algorithm selection
abstract
We introduce the algorithm selection problem for power minimization. After demonstrating the high impact of this synthesis task on the power consumption of the final implementation using a case study, we studied its computational complexity. We present an efficient optimization intensive algorithm for power minimization using algorithm selection. We applied the marginal utility-based algorithm for algorithm selection on three DSP examples. A table illustrates the effectiveness of the power optimization using algorithm selection on one audio (LMS DCT transform domain filter) and two video (NTSC formatter and DPCM coder) applications. On several DSP examples more than an order of magnitude reduction in power is demonstrated.
Miodrag Potkonjak, Jan M. Rabaey
ICASSP2
1995 Design guidance in the power dimension
abstract
This work proposes an approach for high level design guidance for low power using properties of given algorithms and architectures. Several relevant properties (operation count, the ratio of critical path to available time, spatial locality, and regularity) are identified and discussed, with quantitative measures being proposed for the latter two. Significant emphasis is placed on exploiting the regularity and spatial locality algorithm properties for the optimization of interconnect power. Examples illustrate the large savings that can be attained through property-based guidance of algorithm selection and architecture composition. Though demonstrated for ASIC designs, this approach is extensible to different hardware platforms and performance metrics (e.g. speed, area).
Jan M. Rabaey, Lisa M. Guerra, Renu Mehra
ICASSP1
1995 Efficient throughput optimization of feedback linear computations using generalized Horner's scheme
abstract
Presents a generalized Horner's scheme-based approach which enables that a large and important subclass of nonlinear computations, named feedback linear computations, is efficiently, maximally, and arbitrarily sped-up. The new class includes popular nonlinear polynomial Volterra filters and widely used LMS and RLS adaptive filters. The effectiveness and low overhead of the proposed techniques is illustrated on several designs.
Jan M. Rabaey, Miodrag Potkonjak, Kazutoshi Wakabayashi
ICASSP1
1995 Power conscious CAD tools and methodologies: a perspective
abstract
Power consumption is rapidly becoming an area of growing concern in IC and system design houses. Issues such as battery life, thermal limits, packaging constraints and cooling options are becoming key factors in the success of a product. As a consequence, IC and system designers are beginning to see the impact of power on design area, design speed, design complexity and manufacturing cost. While process and voltage scaling can achieve significant power reductions, these are expensive strategies that require industry momentum, that only pay off in the long run. Technology independent gains for power come from the area of design for low power which has a much higher return on investment (ROI). But low power design is not only a new area but is also a complex endeavour requiring a broad range of synergistic capabilities from architecture/microarchitecture design to package design. It changes traditional IC design from a two-dimensional problem (Area/performance) to a three-dimensional one (Area/Performance/Power). This paper describes the CAD tools and methodologies required to effect efficient design for low power. It is targeted to a wide audience and tries to convey an understanding of the breadth of the problem. It explains the state of the art in CAD tools and methodologies. The paper is written in the form of a tutorial, making it easy to read by keeping the technical depth to a minimum while supplying a wealth of technical references. Simultaneously the paper identifies unresolved problems in an attempt to incite research in these areas. Finally an attempt is made to provide commercial CAD tool vendors with an understanding of the needs and time frames for new CAD tools supporting low power design.>
Deo Singh, Jan M. Rabaey, Massoud Pedram, Francky Catthoor, Suresh Rajgopal, Naresh Sehgal, Thomas J. Mozdzen
Proc. IEEE2
1995 Optimizing power using transformations
abstract
The increasing demand for portable computing has elevated power consumption to be one of the most critical design parameters. A high-level synthesis system, HYPER-LP, is presented for minimizing power consumption in application specific datapath intensive CMOS circuits using a variety of architectural and computational transformations. The synthesis environment consists of high-level estimation of power consumption, a library of transformation primitives, and heuristic/probabilistic optimization search mechanisms for fast and efficient scanning of the design space. Examples with varying degree of computational complexity and structures are optimized and synthesized using the HYPER-LP system. The results indicate that more than an order of magnitude reduction in power can be achieved over current-day design methodologies while maintaining the system throughput; in some cases this can be accomplished while preserving or reducing the implementation area.>
Anantha P. Chandrakasan, Miodrag Potkonjak, Renu Mehra, Jan M. Rabaey, Robert W. Brodersen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
1995 Architectural power analysis: The dual bit type method
abstract
This paper describes a novel strategy for generating accurate black-box models of datapath power consumption at the architecture level. This is achieved by recognizing that power consumption in digital circuits is affected by activity, as well as physical capacitance. Since existing strategies characterize modules for purely random inputs, they fail to account for the effect of signal statistics on switching activity. The dual bit type (DBT) model, however, accounts not only for the random activity of the least significant bits (LSB's), but also for the correlated activity of the most significant bits (MSB's), which contain two's-complement sign information. The resulting model is parameterizable in terms of complexity factors such as word length and can be applied to a wide variety of modules ranging from adders, shifters, and multipliers to register files and memories. Since the model operates at the register transfer level (RTL), it is orders of magnitude faster than gate- or circuit-level tools, but while other architecture-level techniques often err by 50-100% or more, the DBT method offers error rates on the order of 10-15%.>
Paul E. Landman, Jan M. Rabaey
IEEE Trans. Very Large Scale Integr. Syst.2
1994 Memory Estimation for High Level Synthesis
abstract
Article Memory estimation for high level synthesis Share on Authors: Ingrid M. Verbauwhede EECS Department, University of California at Berkeley, Cory Hall, Berkeley, CA EECS Department, University of California at Berkeley, Cory Hall, Berkeley, CAView Profile , Chris J. Scheers Zycad Corporation, Fremont, CA Zycad Corporation, Fremont, CAView Profile , Jan M. Rabaey EECS Department, University of California at Berkeley, Cory Hall, Berkeley, CA EECS Department, University of California at Berkeley, Cory Hall, Berkeley, CAView Profile Authors Info & Claims DAC '94: Proceedings of the 31st annual Design Automation ConferenceJune 1994 Pages 143–148https://doi.org/10.1145/196244.196313Online:06 June 1994Publication History 55citation265DownloadsMetricsTotal Citations55Total Downloads265Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Ingrid Verbauwhede, Chris J. Scheers, Jan M. Rabaey
DAC3
1994 Concurrency characteristics in DSP programs
abstract
The exploration of concurrency has emerged as a dominant research problem in the VLSI DSP literature. A great variety of forms of concurrency exploration have been proposed, analyzed and used on a number of hardware platforms. The belief that the DSP domain is amenable to concurrency exploration has been gaining popularity; however, no systematic study has been conducted to confirm or dispel this claim. This paper presents a global view of the concurrency problem, presenting comprehensive statistics on concurrency properties in commonly used DSP programs. Particular emphasis is placed on the potential cost effectiveness of concurrency exploitation.>
Lisa M. Guerra, Miodrag Potkonjak, Jan M. Rabaey
ICASSP (2)3
1994 Specification and support for multidimensional DSP in the SILAGE language
abstract
Data flow languages are a natural way to describe the flow of computations in a DSP application. The SILAGE language has been developed for this purpose. It contains also more-dimensional arrays of signals and a natural extension of it, delayed versions of arrays, e.g. to represent previous frames in video applications. The paper describes new data flow analysis techniques, to support multi-dimensional arrays. It checks single assignment of arrays, checks if for each consumption of an indexed signal, there is a production, and it will create data dependencies between productions and consumptions. These problems are formulated as integer linear programming problems. This formulation is independent of the number of signals in the arrays. Results show very fast running times (>
Ingrid Verbauwhede, Chris J. Scheers, Jan M. Rabaey
ICASSP (2)3
1994 Algorithm selection: a quantitative computation-intensive optimization approach
Miodrag Potkonjak, Jan M. Rabaey
ICCAD2
1994 Is it Possible to achieve a Teraflop/s on a chip? From High Performance Algorithms to Architectures
abstract
The forumnists address the question of high density computations on a single chip. The surface of a chip offers an ideal medium not only to store information or to process data, but also to execute computations. The 1 Giga floating point operations per second per chip mark has been achieved, we are now moving towards the teraflop mark. How is this going to happen, what are the limitations, what are the opportunities-those are central questions.>
Francky Catthoor, Ed F. Deprettere, Yu Hen Hu, Jan M. Rabaey, Heinrich Meyr, Lothar Thiele
ISCAS4
1994 Research challenges in wireless multimedia
abstract
The near future will bring the fusion of four rapidly evolving technologies: high speed networking and associated services, wireless communications, scaled integrated circuit technology, and multimedia-based applications. These new technologies will enable the access of multimedia data from network servers at any time and any place by light weight, low cost wireless terminals.
Robert W. Brodersen, Thomas D. Burd, Fred L. Burghardt, Andrew J. Burstein, Anantha P. Chandrakasan, Roger Doering, Shankar Narayanaswamy, Trevor Pering, Brian C. Richards, Thomas E. Truman, Jan M. Rabaey
PIMRC11
1994 Optimizing resource utilization using transformations
abstract
The goal of the high level synthesis process for real time applications is to minimize the implementation cost, while still satisfying all timing constraints. In this paper, the authors present a combination of four conceptually simple, yet powerful, transformations: namely retiming, associativity, commutativity and inverse element law, which can help to further this goal. Since the minimization problem associated with these transformations is NP complete, a new fast iterative improvement probabilistic algorithm has been developed. The effectiveness of the proposed algorithm and the associated transformations is demonstrated in multiple ways: using standard benchmark examples, with the aid of statistical analysis and through a comparison with estimated minimal bounds.>
Miodrag Potkonjak, Jan M. Rabaey
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1994 Estimating implementation bounds for real time DSP application specific circuits
abstract
This paper discusses techniques for estimating implementation bounds on computational resources and their role in the high-level synthesis process. Accurate estimations can be extremely useful in a multitude of synthesis operations, such as algorithm and architecture selection, design space search, module selection, transformations, allocation, assignment, and scheduling. Several techniques to efficiently estimate sharp minimum and maximum bounds on the resource requirements of a hardware implementation are discussed. The performance of the algorithms as well as their applications is analyzed using an extensive benchmark set. The proposed techniques have been implemented in the HYPER synthesis system.>
Jan M. Rabaey, Miodrag Potkonjak
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1993 Heterogeneous BISR techniques for yield and reliability enhancement using high level synthesis transformations
abstract
Built-In-Self-Repair (BISR) is a fault tolerance technique against permanent faults, where in addition to core operational modules, a set of spare modules is provided. If a faulty core module is detected, it is replaced with a spare module. The BISR methodology has been used only in situations where a failed module of one type can only be replaced by a backup module of the same type. The authors propose a new BISR approach for ASIC design which removes this constraint and enables replacement of modules of different types with the same spare units by exploiting the design space exploration abilities provided by the use of transformations in high level synthesis. Fast and efficient high level synthesis algorithms which take into account peculiarities of transformation-based design for BISR are presented. The potential of the approach is demonstrated on a set of benchmark examples by showing significant yield and relative productivity improvements which are calculated using state-of-the-art yield modeling techniques.>
Miodrag Potkonjak, Lisa M. Guerra, Jan M. Rabaey
ASAP3
1993 On unlimited parallelism of DSP arithmetic computations
Miodrag Potkonjak, Jan M. Rabaey
ICASSP (1)2
1993 Instruction set mapping for performance optimization
abstract
Performance optimization is the primary design goal in most digital signal processing (DSP) and numerically intensive applications. The problem of mapping high-level algorithmic descriptions for these applications to specialized instruction sets has only recently begun to receive attention. In fact, the problem of optimizing performance has yet to be addressed directly. This paper introduces a new approach to instruction set mapping (and template matching in general) targeted toward performance optimization. Several novel issues are addressed including partial matching and automatic clock selection.
Miguel R. Corazao, Marwan A. Khalaf, Lisa M. Guerra, Miodrag Potkonjak, Jan M. Rabaey
ICCAD5
1993 High level synthesis for reconfigurable datapath structures
abstract
High level synthesis techniques for the synthesis of restructurable datapaths are introduced. The techniques can be used in applications such as design for fault tolerance against permanent faults, design for yield improvement, and design of application specific programmable processors. The paper focuses on design techniques for built in self repair (BISR), which addresses the first two of these applications. The new BISR methodology consists of two approaches which exploit the design space exploration abilities of high level synthesis. The first method uses resource allocation, assignment, and scheduling, and the second uses transformations. The effectiveness of the approaches are verified on a set of benchmark examples.
Lisa M. Guerra, Miodrag Potkonjak, Jan M. Rabaey
ICCAD3
1992 An integrated system for rapid prototyping of high performance algorithm specific data paths
abstract
A system has been developed which targets the rapid prototyping of high performance data computation units which are typical to real-time digital signal processing applications. The hardware platform of the system is a family of multiprocessor integrated circuits. The prototype chip of this family contains 8 processors connected via a dynamically controlled crossbar switch. With a maximum clock rate of 25 MHz, it can support a computation rate of 200 MIPs and can sustain a data I/O bandwidth of 400 MByte/sec. An assembler and simulator provide low-level programmability of the hardware. A compiler which takes input described in the high-level data flow language Silage, and performs estimation, transformations, partitioning, assignment, and scheduling before generating assembly code, provides an automated software compilation path.>
D. C. Chen, Lisa M. Guerra, E. H. Ng, Miodrag Potkonjak, D. P. Schultz, Jan M. Rabaey
ASAP6
1992 Hierarchical scheduling of DSP programs onto multiprocessors for maximum throughput
abstract
A multiprocessor scheduling algorithm that simultaneously considers pipelining, retiming, parallel execution and hierarchical node decomposition to maximize performance throughput is presented. The algorithm is able to take into account interprocessor communication delays, and memory and processor availability constraints. The results on a set of benchmarks demonstrate the algorithm's ability to achieve near optimal speedups across a wide range of applications of various types of concurrency, with good scalability with respect to processor count.>
Phu Hoang, Jan M. Rabaey
ASAP2
1992 Pipelining: just another transformation
abstract
A simple formulation of pipelining: 'Pipelining with N stages is equivalent to retiming where the number of delays on all inputs or all outputs, but not both, is increased by N' is used as the basis for a convenient and efficient treatment of pipelining in design of application specific computers. Classification of pipelining according to the optimization goal (throughput and resource utilization) and the latency is introduced. For polynomial complexity pipelining classes, optimal algorithms are presented. For other classes both proof of NP-completeness and efficient probabilistic algorithms are presented. Both theoretical and experimental properties of pipelining are discussed. In particular, a relationship with other transformations is explored. Due to close relationship between software pipelining and pipelining presented, all results can be easily modified for use in compilers for general purpose computers. Also, as a side result, the exact bound (solution) for iteration bound is derived.>
Miodrag Potkonjak, Jan M. Rabaey
ASAP2
1992 A compiler for multiprocessor DSP implementation
abstract
McDAS, a software environment designed to support the real-time implementation of digital signal processing (DSP) applications onto multiple processors, is described. Users program their algorithms as they would on a single processor, and McDAS automatically schedules and compiles the program onto the target multiprocessor. The scheduler maximizes the computational throughput by simultaneously considering pipelining, retiming, and parallelism while accounting for processor and memory constraints, as well as interprocessor communication delays. If the architecture is scalable or configurable, the scheduler can be invoked with different numbers of processors and multiprocessor topologies to explore various implementations. The code generator is similarly retargetable to different memory architectures and core processors. Data buffers and synchronizations are automatically inserted to ensure correct execution. The code generated can also execute the algorithms with either quasi-infinite precision or bit-true precision, allowing the algorithm designer to assess the effects of quantization and truncation. The results on a set of benchmarks are presented.>
Phu Hoang, Jan M. Rabaey
ICASSP2
1992 Fast implementation of recursive programs using transformations
abstract
An automatic transformational approach used to reduce the iteration bound of recursive DSP algorithms is presented. The proposed approach combines delay retiming, algebraic transformations and loop unrolling in a well defined order. The effectiveness of the approach is demonstrated using examples.>
Miodrag Potkonjak, Jan M. Rabaey
ICASSP2
1992 HYPER-LP: a system for power minimization using architectural transformations
abstract
An automated high-level synthesis system, HYPER-LP, for minimizing power consumption in application-specific datapath-intensive CMOS circuits using a variety of architectural and computational transformations is presented. The sources of power consumption are reviewed, and the effects of architectural transformations on the various power components are presented. The synthesis environment consists of high-level estimation of power consumption, a library of transformation primitives (local and global), and heuristic/probabilistic optimization search mechanisms for fast and efficient scanning of the design space. Examples with varying degree of computational complexity and structures are optimized and synthesized. The results indicate that an order of magnitude reduction in power can be achieved over current-day design methodologies while maintaining the system throughput; in some cases, this can be accomplished while preserving or reducing the implementation area.>
Anantha P. Chandrakasan, Miodrag Potkonjak, Jan M. Rabaey, Robert W. Brodersen
ICCAD3
1992 Maximally fast and arbitrarily fast implementation of linear computations
abstract
It is pointed out that by establishing a relationship between the basic properties of linear computations and several optimizing transformations, it is possible to optimally speed up linear computations with respect to those transformations while keeping the latency fixed. Furthermore, arbitrarily fast, asymptotically optimal implementations can be obtained by adding retiming and loop unrolling to the transformations set and trading latency for throughput. The proposed techniques have yielded results superior to the best published previously on all benchmark examples. The presented approach is also applicable to general (nonlinear) computations.>
Miodrag Potkonjak, Jan M. Rabaey
ICCAD2
1991 Optimizing Resource Utilization Using Transformations
abstract
A transformational approach aimed at improving the resource utilization in high level synthesis is introduced. The current implementation combines retiming and associativity in a single framework. This combination of transformations results in considerable area improvements, as is amply demonstrated by benchmark examples. A novel learning while searching iterative improvement probabilistic algorithm has been developed and is used to resolve the associated NP-complete combinatorial optimization problem. The effectiveness of the proposed algorithms and the transformations is demonstrated using standard benchmark examples, with the aid of statistical analysis, and through a comparison with estimated minimal bounds. The proposed algorithm has proven to be very effective in reaching the optimal solution as well as in runtime.>
Miodrag Potkonjak, Jan M. Rabaey
ICCAD2
1991 An integrated CAD system for algorithm-specific IC design
abstract
LAGER is an integrated computer-aided design system for algorithm-specific integrated circuit design, targeted at applications such as speech processing, image processing, telecommunications, and robot control. LAGER provides user interfaces at behavioral, structural, and physical levels and allows easy integration of novel CAD tools. LAGER consists of a behavioral mapper and a silicon assembler. The behavioral mapper maps the behavior onto a parameterized structure to produce microcode and parameter values. The silicon assembler then translates the filled-out structural description into a physical layout, and, with the aid of simulation tools, the user can fine tune the data path by iterating this process. The silicon assembler can also be used without the behavioral mapper for high-sample-rate applications. A number of algorithm-specific ICs designed with LAGER have been fabricated and tested, and as examples, a robot arm controller chip and a real-time image segmentation chip are described.>
C. Bernard Shung, Rajeev Jain, Ken Rimey, Mani Srivastava 0001, Brian C. Richards, Erik Lettang, Syed Khalid Azim, Lars E. Thon, Paul N. Hilfinger, Jan M. Rabaey, Robert W. Brodersen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.11
1990 DSP specification using the Silage language
abstract
The Silage language is presented, together with a number of extensions to the original design which result in a powerful specification language and design environment for complex digital signal processing (DSP) systems. Aspects of Silage discussed include the applicative language, timing information and time-domain operations, data typing, and functional language, pragmatic directives, and loops and conditionals.>
Dominique Genin, Paul N. Hilfinger, Jan M. Rabaey, Chris J. Scheers, Hugo De Man
ICASSP3
1990 An efficient microcode compiler for application specific DSP processors
abstract
A computer program for microcode compilation for custom digital signal processors is presented. This tool is part of the CATHEDRAL II silicon compiler. The following optimization problems are highlighted: scheduling, hardware assignment, and loop folding. Efficient techniques to solve these problems are developed. This allows for the automatic synthesis of processor architectures which simultaneously exploit pipelining and parallelism. A demonstrator design is presented.>
Gert Goossens, Jan M. Rabaey, Joos Vandewalle, Hugo De Man
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1989 A Scheduling and Resource Allocation Algorithm for Hierarchical Signal Flow Graphs
abstract
The paper describes a new algorithm for the scheduling and resource allocation problem in high-level synthesis. The algorithm can not only efficiently treat flattened signal flow graphs, but also handles graphs with embedded control constructs such as conditional branches and loops. Based on simple and clear, but powerful principles, the algorithm simultaneously minimizes the number of execution units, the number of registers and the amount of interconnections. The algorithm has been implemented and we present the first results, which are very promising.
Miodrag Potkonjak, Jan M. Rabaey
DAC2
1989 A large-vocabulary real-time continuous-speech recognition system
abstract
A system architecture has been developed to implement real-time large-vocabulary continuous-speech recognition using HMM (hidden Markov model) algorithms and bigram language models. It is shown that the largest bottleneck in such a system is located in the memory access. The architecture exploits a variety of techniques, such as partitioning and replication, to cope with this memory bottleneck. The required throughput is achieved with the aid of extensive pipelining (up to thirteen levels deep) and concurrency. The architecture allows extension to larger vocabularies by the addition of more parallel units. Pin count considerations have resulted in the definition of five custom integrated circuits which are currently being tested. Using the proposed approach, the authors are currently designing and debugging a real-time 3000-word continuous-speech recognition system that uses bigram language models.>
Hy Murveit, J. Mankoski, Jan M. Rabaey, Robert W. Brodersen, T. Stoelzle, Dev C. Chen, Shankar Narayanaswamy, P. Schrupp, R. Schwartz
ICASSP3
1989 HYPER: an interactive synthesis environment for high performance real time applications
abstract
A synthesis system called HYPER is proposed for real-time applications. HYPER takes a flow graph description of an algorithm as the input and performs scheduling, resource allocation, optimizations, and transformations. A dedicated bit-sliced data path cluster is generated by the system and the layouts can be further generated through the LAGER IV system.>
Chi-Min Chu, Miodrag Potkonjak, Markus Thaler, Jan M. Rabaey
ICCD4
1986 An intelligent module generator environment
abstract
An environment for the generation of modules is described. It includes tools for interactive design of parameterised procedures describing the structure as well as the topology. For the layout symbolic cells are used which are automatically fitted together as defined by the topology.
Paul Six, Luc Claesen, Jan M. Rabaey, Hugo De Man
DAC3
1986 Experiences with automatic generation of audio band digital signal processing circuits
abstract
A set of digital signal processing circuits, generated by the Lager design synthesis system ([PO84], [Ra85], is described. This system maps a behavioral, assembler level description of an algorithm into a set of concurrent operating, customized processors and dedicated i/o circuitry. The nature of this restricted, but flexible target architecture makes it possible to span a large range of applications in the field of speech processing, audio and telecommunications.
Jan M. Rabaey, Robert W. Brodersen
ICASSP1
1985 An Integrated Automated Layout Generation System for DSP Circuits
abstract
An integrated CAD system for the automated design of digital signal-processing (DSP) circuits for audio and telecommunication applications is described. The system uses as unique input a symbolic description of algorithm. This representation is translated into an actual layout using a two-step process. First, the symbolic input is mapped into the target architecture, which consists basically of a set of concurrent processors and dedicated I/O circuitry. The resulting hardware configuration is compiled into a layout description through a full exploitation of the hierarchy and the modularity of the architecture, calling consecutively a tiler, a floorplanner, and a global placement and routing tool. All these layout generation tools are able to support a wide range of technologies. The provision of a dedicated register transfer level simulator allows for the efficient debugging and algorithmic checking of the real-time operating signal-processing algorithms. The efficiency and the usefulness of this design methodology has been demonstrated by multiple examples. Experiments have shown that the use of these techniques can reduce the complete design process to a few months.
Jan M. Rabaey, Stephen P. Pope, Robert W. Brodersen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1