Nikil Dutt

dblp:d/NikilDDutt · also Nikil D. Dutt · DBLP profile ↗
← Back
334ranked-venue papers
26as first author
46since 2021 · last 2026
0000-0002-3060-8119ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 280 · 24 first-author · 33 since 2021Software engineering, systems software and programming languages · 57 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 14 · 3 since 2021Computer networks · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization
abstract
Recent advances in LLMs have outpaced the computational and memory capacities of edge platforms that primarily employ CPUs, thereby challenging efficient and scalable deployment. While ternary quantization enables significant resource savings, existing CPU solutions rely heavily on memory-based lookup tables (LUTs) which limit scalability, and FPGA or GPU accelerators remain impractical for edge use. This paper presents T-SAR, the first framework to achieve scalable ternary LLM inference on CPUs by repurposing the SIMD register file for dynamic, in-register LUT generation with minimal hardware modifications. T-SAR eliminates memory bottlenecks and maximizes data-level parallelism, delivering 5.6–24.5× and 1.1–86.2× improvements in GEMM latency and GEMV throughput, respectively, with only 3.2% power and 1.4% area overheads in SIMD units. T-SAR achieves up to 2.5–4.9× the energy efficiency of an NVIDIA Jetson AGX Orin, establishing a practical approach for efficient LLM inference on edge platforms.
Hyunwoo Oh, KyungIn Nam, Rajat Bhattacharjya, Hanning Chen, Tamoghno Das, Sanggeon Yun, Suyeon Jang, Andrew Ding, Nikil Dutt, Mohsen Imani
DATE9
2026 DPAS: A Prompt, Accurate and Safe I/O Completion Method for SSDs
Dongjoo Seo, Jihyeon Jung, Yeohwan Yoon, Ping-Xiang Chen, Yongsoo Joo, Sung-Soo Lim, Nikil Dutt
FAST7
2026 ACCESS-AV: Adaptive Communication-Computation Codesign for Sustainable Autonomous Vehicle Localization in Smart Factories
abstract
Autonomous Delivery Vehicles (ADVs) are increasingly used for transporting goods in 5G network-enabled smart factories, with the compute-intensive localization module presenting a significant opportunity for optimization. We propose ACCESS-AV , an energy-efficient Vehicle-to-Infrastructure (V2I) localization framework that leverages existing 5G infrastructure in smart factory environments. By opportunistically accessing the periodically broadcast 5G Synchronization Signal Blocks (SSBs) for localization, ACCESS-AV obviates the need for dedicated Roadside Units (RSUs) or additional onboard sensors to achieve energy efficiency as well as cost reduction. We implement an Angle-of-Arrival (AoA)-based estimation method using the Multiple Signal Classification (MUSIC) algorithm, optimized for resource-constrained ADV platforms through an adaptive communication-computation strategy that dynamically balances energy consumption with localization accuracy based on environmental conditions such as Signal-to-Noise Ratio (SNR) and vehicle velocity. Experimental results demonstrate that ACCESS-AV achieves an average energy reduction of 43.09% compared to non-adaptive systems employing AoA algorithms such as vanilla MUSIC, ESPRIT, and Root-MUSIC. It maintains sub-30 cm localization accuracy while also delivering substantial reductions in infrastructure and operational costs, establishing its viability for sustainable smart factory environments.
Rajat Bhattacharjya, Arnab Sarkar 0001, Ish Kool, Sabur Baidya, Nikil Dutt
ACM Trans. Embed. Comput. Syst.5
2025 Invited Paper: Mindful AI for Pervasive Health and Wellbeing (PHW)
abstract
Emerging AI-driven pervasive health and wellbeing (PHW) services (e.g., personalized health assistants and mobile health applications) face critical challenges in handling noisy/intermittent sensory data, integrating cross-modal insights, and stringent energy and compute constraints. We present Mindful AI, a cognitive-inspired framework designed to enable adaptive, resilient, and efficient PHW services in real-world conditions. Our dual-mode intelligence—Automatic (System 1) and Reflective (System 2)—selectively directs system attention toward the most relevant sensing and compute contexts, unifying bottom-up stimuli (driven by input quality, inference demands and model confidence, and resource availability) with top-down insights (reflecting user demands, system goals/constraints, and contextual information). Our framework distills and orchestrates insights across sensing, communication, and computation through hybrid attention toward bottom-up and top-down insights that support cross-layer sense-compute co-optimization to achieve resilient, low-latency, and energy-efficient PHW services. We evaluate our approach on multi-tier device-edge-cloud platforms, using real-world case studies in pain assessment, stress monitoring, and human activity recognition to demonstrate adaptation to real-world uncertainties (e.g., sensor degradation, context drift, network variability), while maintaining strict QoS, accuracy, and latency guarantees.
Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt
ICCAD5
2025 Exploiting Approximate SRAM for Energy-Efficient Integer Motion Estimation on VVC Encoders
abstract
Video coding is a critical technology for enabling many modern applications. However, the high complexity of state-of-the-art encoders leads to energy consumption challenges, particularly in memory systems. This paper proposes an integer motion estimation (IME) system using approximate SRAM memories to enhance the energy efficiency of VVC encoders. The approximate SRAM memories are employed for both the current block and the search area buffers by reducing the supply voltage. The proposed IME system is evaluated using a customized tool that simulates the approximation effects on both buffers based on real energy measurements from a 28nm SRAM. This tool is integrated with VVenC, a fast VVC implementation, to assess the impact on coding efficiency and energy consumption. Experimental results show that the IME system using approximate SRAM can reduce energy consumption in reading operations by up to 55%, with a low impact on coding efficiency.
Matheus Isquierdo, Felipe Sampaio, Bruno Zatt, Nikil Dutt, Daniel Palomino 0001
ISCAS4
2025 SCHED: Safe CPU Scheduling Framework with Reinforcement Learning and Decision Trees for Autonomous Vehicles
abstract
Autonomous vehicles (AVs) require consistently low-latency computations; however, operating system (OS) CPU scheduler can lead to high tail latency that threatens timely decision-making and safety. General-purpose OS schedulers prioritize fairness and throughput over individual task deadlines, posing challenges for complex AV workloads with mixed-criticality tasks. Despite extensive research on CPU scheduling, reinforcement learning (RL) has not been explored at the OS level due to feasibility concerns. This paper presents SCHED, the first RL-based CPU scheduling framework optimized for OS integration. SCHED learns a scheduling policy via reinforcement learning and deploys it as a lightweight, decision-tree-based scheduler. We demonstrate that SCHED remains sufficiently fast for OS-level integration compared to existing OS schedulers. Furthermore, we validate SCHED on a realistic autonomous vehicle pipeline, demonstrating its practical potential in maintaining low latency and ensuring safer, more responsive AV operations.
Dongjoo Seo, Changhoon Sung 0001, Ping-Xiang Chen, Bryan Donyanavard, Nikil Dutt
VTC2025-Spring5
2025 Runtime Adaptivity for Efficient Neural Network Inference on Autonomous Systems
abstract
Neural network pruning and dynamic training have emerged as key techniques for optimizing deep learning models to meet the constraints of resource-limited systems. However, achieving both efficiency and adaptability without compromising safety or performance remains a significant challenge in real-time autonomous applications. We present Back to the Future and USA-Nets , two complementary approaches that address this challenge. Back to the Future combines pruning with dynamic routing to enable latency gains and dynamic reconfiguration at runtime, allowing a pruned model to seamlessly revert to the full model when unsafe or anomalous behavior is detected. USA-Nets extend this concept by enabling runtime adaptability through dynamically trained networks that can adjust their width without requiring additional annotated data or excessive storage overhead. Together, these methods deliver significant performance improvements while maintaining safety and flexibility, as evidenced by experimental results demonstrating that Back to the Future achieves a 32× faster reversion time compared to loading the full model, and USA-Nets achieve up to 85% latency reduction with minimal accuracy degradation. These innovations pave the way for efficient, adaptable, and safe deployment of deep learning models in diverse real-time and resource-constrained environments, with future work focusing on advanced pruning techniques and runtime optimizations.
Danny Abraham, Biswadip Maity, Bryan Donyanavard, Nikil Dutt
ACM Trans. Embed. Comput. Syst.4
2025 OASIS: Optimized Adaptive System for Intelligent SLAM
abstract
Visual Simultaneous Localization and Mapping (VSLAM) is essential for mobile autonomous systems operating in complex dynamic environments. VSLAM algorithms are computationally intensive and must execute in real-time on resource-constrained embedded devices. Variations in environmental complexity can lead to longer frame processing times, causing dropped frames, lost localization information, and degraded accuracy. To address these challenges, we introduce OASIS, a novel adaptive approximation method that dynamically reduces input frame areas based on realtime visual importance. Unlike traditional optimizations that require adjusting internal SLAM parameters, OASIS selectively minimizes computation by adaptively filtering less critical image regions, significantly reducing computational load. Evaluations on the EuRoC MAV dataset demonstrate that our approach balances accuracy and system predictability, achieving up to a 71.8% reduction in worst-case pose estimation errors. OASIS offers a significant advancement in reliable, predictable, and energy-efficient SLAM tailored for mobile autonomous robotic applications.
Alles Rebel, Nikil Dutt, Bryan Donyanavard
ACM Trans. Embed. Comput. Syst.2
2025 Exploiting Approximation for Run-time Resource Management of Embedded HMPs
abstract
Run-time resource management (RTM) of multi-programmed workloads on heterogeneous multi-core platforms is challenging due to (i) fixed power budget of the device, (ii) variable performance requirements of the workloads, and (iii) unknown arrival of the applications. Existing RTM solutions lack power-performance coordination, resulting in performance degradation during power actuation or power violations during performance provisioning. Exploiting inherent error-resilience of the applications can address the performance loss incurred in power actuation, by combining run-time approximation with traditional power knobs (including Dynamic Voltage/Frequency Scaling, Task Migration, Degree of Parallelism, and CPU Quota ). In this work, we present an accuracy-aware resource management framework that jointly actuates run-time approximation and traditional power knobs for efficient power-performance management of multi-programmed and multi-threaded workloads running on heterogeneous mobile platforms. Our strategy configures the accuracy of the applications at run-time to exploit accuracy-performance trade-offs, by considering system-wide power-performance dynamics. We use heuristic estimation models to jointly enforce accuracy configuration and traditional power knobs settings at run-time. We evaluated our framework on real-world embedded mobile platforms, including Odroid XU3 and Asus Tinker Edge R boards to demonstrate the efficiency of our proposed approach across multiple workload scenarios. Our approach achieved 25% lower performance violations against the state-of-the-art run-time resource management policies at the cost of 2.2% accuracy loss across six applications.
Zain Taufique, Anil Kanduri, Antonio Miele, Amir-Mohammad Rahmani, Cristiana Bolchini, Nikil Dutt, Pasi Liljeberg
ACM Trans. Embed. Comput. Syst.6
2024 BoostIID: Fault-agnostic Online Detection of WCET Changes in Autonomous Driving
abstract
The lifespan of autonomous vehicles is increasing, exposing them to aging and permanent faults that can impact timing safety based on design-time analyses such as worst-case execution time (WCET). Conventional fault-aware WCET methods incorporate potential faults into the analysis, which can result in severely pessimistic WCET estimates. The resulting underutilized system can exhibit significant energy inefficiency during normal operation. We propose BoostIID, a dynamic fault detection mechanism that achieves improved utilization and energy efficiency by monitoring changes in the statistical distribution of WCET at runtime. Unlike prior monitoring methods, BoostIID proactively detects faults before actual timing violations occur, allowing time for recovery measures. As a result, BoostIID eliminates the overhead of fault recovery from classical design-time WCET estimates, resulting in improved energy efficiency. We further improve the detection accuracy using a collection of independent and identically distributed (i.i.d) tests with a boosting technique. Our experimental results with an autonomous driving benchmark suite show 62.6% energy reduction over pessimistic WCET methods, demonstrating BoostIID’s utility for energy-efficient safety-critical system design.
Saehanseul Yi, Nikil Dutt
ASPDAC2
2024 Work-in-Progress: Context and Noise Aware Resilience for Autonomous Driving Applications
abstract
Autonomous Vehicles (AVs) often use noise prone sensory data from cameras and LiDAR for perception. In specific noisy scenarios, different object detection models exhibit non-intuitive and varying degrees of resilience, necessitating adaptive model selection. In this work, we develop a context and noise aware framework for run-time adaptive configuration of objection models for high accuracy and low latency inference. We combine driving scene context and input data noise to prioritize among input modalities, followed by selection and configuration of most resilient object detection model appropriate for the context. Our evaluation for 2D object detection on nuScenes dataset provided average 1.83x speedup in latency compared to baseline while preserving average prediction confidence.
Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt
CODES+ISSS5
2024 Back to the Future: Reversible Runtime Neural Network Pruning for Safe Autonomous Systems
abstract
Neural network pruning has emerged as a technique to reduce the size of networks at the cost of accuracy to enable deployment in resource-constrained systems. However, low-accuracy pruned models may compromise the safety of realtime autonomous systems when encountering unpredictable scenarios, e.g., due to anomalous or emergent behavior. We propose Back to the Future: a novel approach that combines pruning with dynamic routing to achieve both latency gains and dynamic reconfiguration to meet desired accuracy at runtime. Our approach enables the pruned model to quickly revert to the full model when unsafe behavior is detected, enhancing safety and reliability. Experimental results demonstrate that our swapping approach is 32× faster than loading the original model from disk, providing seamless reversion to the accurate version of the model, demonstrating its applicability for safe autonomous systems design.
Danny Abraham, Biswadip Maity, Bryan Donyanavard, Nikil Dutt
DATE4
2024 SEAL: Sensing Efficient Active Learning on Wearables through Context-awareness
abstract
In this paper, we introduce SEAL, a co-optimization framework designed to enhance both sensing and querying strategies in wearable devices for mHealth applications. Employing Reinforcement Learning (RL), SEAL strategically utilizes user contextual information and the machine learning model's confidence levels to make efficient decisions. This innovative approach is particularly significant in addressing the challenge of battery drain due to continuous physiological signal sensing, such as Photoplethysmography (PPG). Our framework demonstrates its effectiveness in a stress monitoring application, achieving a substantial reduction of 76% in the volume of PPG signals collected, while only experiencing a minor 6% decrease in user-labeled data quality. This balance showcases SEAL's potential in optimizing data collection in a way that is considerate of both device constraints and data integrity.
Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt
DATE6
2024 KDTree-SOM: Self-organizing Map based Anomaly Detection for Lightweight Autonomous Embedded Systems
abstract
Self-Organizing Maps (SOM) promise a lightweight approach for multivariate time series anomaly detection in lightweight autonomous embedded systems. However, the enormous volume of time series data from autonomous systems testing requires huge SOMs with impractical search overhead. We present KDTree-SOM that effectively optimizes the winner node search for huge SOMs by reconstructing the SOM as a k-dimensional tree (kd-tree). KDTree-SOM achieves on average a 4 × inference time reduction for huge SOMs while achieving up to 95% anomaly detection accuracy with only KB-level memory overhead, demonstrating its potential for anomaly detection in lightweight autonomous embedded platforms.
Ping-Xiang Chen, Dongjoo Seo, Biswadip Maity, Nikil Dutt
ACM Great Lakes Symposium on VLSI4
2024 Improving Virtualized I/O Performance by Expanding the Polled I/O Path of Linux
abstract
The continuing advancement of storage technology has introduced ultra-low latency (ULL) SSDs that feature 20 μs or less access latency. Therefore, the context switching overhead of interrupts has become more pronounced on these SSDs, prompting consideration of polling as an alternative to mitigate this overhead. At the same time, the high price of ULL SSDs is a major issue preventing the wide adoption of polling.
Dongjoo Seo, Yongsoo Joo, Nikil Dutt
HotStorage3
2024 EA^2: Energy Efficient Adaptive Active Learning for Smart Wearables
abstract
Mobile Health (mHealth) applications rely on supervised Machine Learning (ML) algorithms, requiring end-user-labeled data for the training phase. The gold standard for obtaining such labeled data is by sending queries to users and gathering responses for the corresponding label, which was conventionally done through triggering questions sent at random. Active Learning (AL) methods use intelligent query-sending policies by incorporating users' contextual information to maximize the response rate and informativeness of the collected labeled data. However, wearable devices' substantial battery drainage associated with the sensing of physiological signals underscores the need for developing an efficient sensing policy in addition to a query-sending policy. In this work, we present a co-optimization framework for both sensing and querying strategies within wearable devices, leveraging contextual information and ML model's prediction confidence. We designed a Reinforcement Learning (RL) agent to quantify different contextual parameters combined with model confidence to determine sensing and querying decisions. Our evaluation of an exemplar stress monitoring application showed a 76% reduction in sensing and data transmission energy consumption, with only a 6% drop in user-labeled data.
Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt
ISLPED6
2024 Message from the DM-SmartHealth 2024 Co-Chairs; SMARTCOMP 2024
abstract
It is our great pleasure to welcome you to the 1st IEEE International Workshop on Digital and Mobile Smart Health Systems (DM-SmartHealth 2024) co-located with the 10th IEEE International Conference on Smart Computing (SMARTCOMP 2024), to be held in person on June 29th, 2024, in Osaka, Japan.
Mattia Giovanni Campana, Nikil Dutt
SMARTCOMP2
2024 PERFECT: Personalized Exercise Recommendation Framework and architECTure
abstract
Background : The health benefits of regular physical activity (PA) are well-established and widely acknowledged. Through the integration of wearable trackers, the Internet of Things (IoT)—a network of interconnected devices capable of collecting and exchanging data—coupled with mobile health (mHealth), which refers to the use of mobile devices to support medical and public health practices, it is now feasible to systematically gather and present individual exercise behaviors. This advanced approach enables the precise correlation of users’ physiological data and daily activities with their specific fitness needs, offering a personalized pathway to improving health outcomes. Objective : This study aims to enhance PA levels among individuals by developing a personalized exercise recommendation system. Utilizing reinforcement learning, the system proposes tailored exercise plans based on biomarkers and the user’s specific context. Methods : In this study, we developed applications for smartphones and smartwatches designed to gather, monitor, and recommend exercise routines through the application of a contextual multi-arm bandit algorithm. To evaluate the efficacy of this mHealth exercise regimen, we enlisted the participation of twenty female college students. Results : The outcomes of our investigation revealed a significant enhancement in the average daily duration of exercise (P \({\lt}\) . 001). Participants expressed high levels of satisfaction with both the walking program and the recommendation system, achieving average ratings of 4.31 (SD \(=\) 0.60) and 3.69 (SD \(=\) 0.95), respectively, on a 5-point scale. Furthermore, the average scores for participants’ confidence in safely performing the recommended walking exercises, as well as their perception of the study’s effectiveness in meeting their PA needs, were both above 4, indicating a positive reception and confidence in the program’s design and implementation. Conclusions : The evolution of the IoT and wearable technology has marked the beginning of a new era for mHealth systems, particularly in the personalization of health interventions. Such advancements enable the precise personalization of PA recommendations, potentially enhancing user engagement and performance outcomes. This paper introduces a novel exercise recommendation system that utilizes reinforcement learning to personalize walking exercises based on the user’s biomarkers and context, aiming to improve the user’s aerobic capacity significantly.
Milad Asgari Mehrabadi, Elahe Khatibi, Tamara Jimah, Sina Labbaf, Holly Borg, Laura Narvaez, Pamela Pimentel, Arlene Turner, Nikil Dutt, Amir-Mohammad Rahmani
ACM Trans. Comput. Heal.9
2024 ZoneTrace: Zone Monitoring Tool for F2FS on ZNS SSDs
abstract
We present ZoneTrace , a runtime monitoring tool for the Flash-friendly File System (F2FS) on Zoned Namespace (ZNS) Solid-state Drives (SSDs). ZNS SSD organizes its storage into zones of sequential write access. Due to ZNS SSD’s sequential write nature, F2FS is a log-structured file system that has recently been adopted to support ZNS SSDs. To present the space management with the zone concept between F2FS and the underlying ZNS SSD, we developed ZoneTrace , a tool that enables users to visualize and analyze the space management of F2FS on ZNS SSDs. ZoneTrace utilizes the extended Berkeley Packet Filter (eBPF) to trace the updated segment bitmap in F2FS and visualize each zone space usage accordingly. Furthermore, ZoneTrace is able to analyze on file fragmentation in F2FS and provides users with informative fragmentation histogram to serve as an indicator of file fragmentation. Using ZoneTrace ’s visualization, we are able to identify the current F2FS space management scheme’s inability to fully optimize space for streaming data recording in autonomous systems, which leads to serious file fragmentation on ZNS SSDs. Our evaluations show that ZoneTrace is lightweight and assists users in getting useful insights for effortless monitoring on F2FS with ZNS SSD with both synthetic and realistic workloads. We believe ZoneTrace can help users analyze F2FS with ease and open up space management research topics with F2FS on ZNS SSDs.
Ping-Xiang Chen, Dongjoo Seo, Changhoon Sung 0001, Jongheum Park, Minchul Lee, Huaicheng Li, Matias Bjørling, Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.8
2023 Impact of COVID-19 Pandemic on Sleep Including HRV and Physical Activity as Mediators: A Causal ML Approach
abstract
Sleep quality is crucial to both mental and physical well-being. The COVID-19 pandemic, which has notably affected the population’s health worldwide, has been shown to deteriorate people’s sleep quality. Numerous studies have been conducted to evaluate the impact of the COVID-19 pandemic on sleep efficiency, investigating their relationships using correlation-based methods. These methods merely rely on learning spurious correlation rather than the causal relations among variables. Furthermore, they fail to pinpoint potential sources of bias and mediators and envision counterfactual scenarios, leading to a poor estimation. In this paper, we develop a Causal Machine Learning method, which encompasses causal discovery and causal inference components, to extract the causal relations between the COVID-19 pandemic (treatment variable) and sleep quality (outcome) and estimate the causal treatment effect, respectively. We conducted a wearable-based health monitoring study to collect data, including sleep quality, physical activity, and Heart Rate Variability (HRV) from college students before and after the COVID-19 lockdown in March 2020. Our causal discovery component generates a causal graph and pinpoints mediators in the causal model. We incorporate the strongly contributing mediators (i.e., HRV and physical activity) into our causal inference component to estimate the robust, accurate, and explainable causal effect of the pandemic on sleep quality. Finally, we validate our estimation via three refutation analysis techniques. Our experimental results indicate that the pandemic exacerbates college students’ sleep scores by 8%. Our validation results show significant p-values confirming our estimation.
Elahe Khatibi, Mahyar Abbasian, Iman Azimi, Sina Labbaf, Mohammad Feli, Jessica L. Borelli, Nikil Dutt, Amir-Mohammad Rahmani
BSN7
2023 Loneliness Forecasting Using Multi-modal Wearable and Mobile Sensing in Everyday Settings
abstract
The adverse effects of loneliness on both physical and mental well-being are profound. Although previous research has utilized mobile sensing techniques to detect mental health issues, few studies have utilized state-of-the-art wearable devices to forecast loneliness and comprehend the physiological manifestations of loneliness and its predictive nature. The primary objective of this study is to examine the feasibility of forecasting loneliness by employing wearable devices, such as smart rings and watches, to monitor early physiological indicators of loneliness. Furthermore, smartphones are employed to capture initial behavioral signs of loneliness. To accomplish this, we employed personalized machine learning techniques, leveraging a comprehensive dataset comprising physiological and behavioral information obtained during our study involving the monitoring of college students. Through the development of personalized models, we achieved a notable accuracy of 0.82 and an F-1 score of 0.82 in forecasting loneliness levels seven days in advance. Additionally, the application of Shapley values facilitated model explainability. The wealth of data provided by this study, coupled with the forecasting methodology employed, possesses the potential to augment interventions and facilitate the early identification of loneliness within populations at risk.
Zhongqi Yang, Iman Azimi, Salar Jafarlou, Sina Labbaf, Jessica L. Borelli, Nikil Dutt, Amir-Mohammad Rahmani
BSN6
2023 Tutorial: MARS: A Framework for Runtime Monitoring, Modeling, and Management of Realtime Systems
Bryan Donyanavard, Nikil Dutt, Biswadip Maity, Parth Malani, Tiago Rogério Mück
CODES+ISSS2
2023 Lightning Talk: The New Era of Computational Cognitive Intelligence
abstract
The triple whammy of variability in platforms (e.g., process variability), applications (e.g., dynamic use cases in autonomous systems), and the environment (e.g., context) renders ineffective the classical computational/algorithmic/numerical computing paradigm in dealing with the inherent runtime dynamism and uncertainty faced by emerging systems. We posit that this requires a fundamental change from classical "static" computing to a new era that deploys a computational cognitive intelligence (CCI) paradigm that is able to learn and evolve at runtime. The CCI paradigm empowers systems to be adaptable and evolvable by exploiting biologically-inspired cognitive intelligence principles.
Nikil Dutt, Bryan Donyanavard
DAC1
2023 Self-awareness in Cyber-Physical Systems: Recent Developments and Open Challenges
abstract
Self-aware computing systems enable computing systems to reflect on their actions and behavior. This becomes even more relevant in Cyber-Physical Systems where computing systems have to control and interact with elements in the real world. This paper reports on recent advances made in computational self-awareness for cyber-physical systems.
Lukas Esterle, Nikil Dutt, Christian Gruhl, Peter R. Lewis 0001, Lucio Marcenaro, Carlo S. Regazzoni, Axel Jantsch
DATE2
2023 Information Processing Factory 2.0 - Self-awareness for Autonomous Collaborative Systems
abstract
This paper summarizes the talks of a special session on the IPF 2.0 project, a collaborative German-US research project that leverages self-awareness principles for the self-management of distributed systems of autonomous multiprocessor systems-on-chip (MPSoCs).
Nora Sperling, Alex Bendrick, Dominik Stöhrmann, Rolf Ernst, Bryan Donyanavard, Florian Maurer 0003, Oliver Lenke, Anmol Surhonne, Andreas Herkersdorf, Walaa Amer, Caio Batista de Melo, Ping-Xiang Chen, Quang Anh Hoang, Rachid Karami, Biswadip Maity, Paul Nikolian, Mariam Rakka, Dongjoo Seo, Saehanseul Yi, Minjun Seo, Nikil Dutt, Fadi J. Kurdahi
DATE21
2023 Locate: Low-Power Viterbi Decoder Exploration using Approximate Adders
abstract
Viterbi decoders are widely used in communication systems, natural language processing (NLP), and other domains. While Viterbi decoders are compute-intensive and power-hungry, we can exploit approximations for early design space exploration (DSE) of trade-offs between accuracy, power, and area. We present Locate, a DSE framework that uses approximate adders in the critically compute and power-intensive Add-Compare-Select Unit (ACSU) of the Viterbi decoder. We demonstrate the utility of Locate for early DSE of accuracy-power-area trade-offs for two applications: communication systems and NLP, showing a range of pareto-optimal design configurations. For instance, in the communication system, using an approximate adder, we observe savings of 21.5% area and 31.02% power with only 0.142% loss in accuracy averaged across three modulation schemes. Similarly, for a Parts-of-Speech Tagger in an NLP setting, out of 15 approximate adders, 7 report 100% accuracy while saving 22.75% area and 28.79% power on average when compared to using a Carry-Lookahead Adder in the ACSU. These results show that Locate can be used synergistically with other optimization techniques to improve the end-to-end efficiency of Viterbi decoders for various application domains.
Rajat Bhattacharjya, Biswadip Maity, Nikil Dutt
ACM Great Lakes Symposium on VLSI3
2023 Is Garbage Collection Overhead Gone? Case study of F2FS on ZNS SSDs
abstract
The sequential write nature of ZNS SSDs makes them very well-suited for log-structured file systems. The Flash-Friendly File System (F2FS), is one such log-structured file system and has recently gained support for use with ZNS SSDs. The large F2FS over-provisioning space for ZNS SSDs greatly reduces the garbage collection (GC) overhead in the log-structured file systems. Motivated by this observation, we explore the trade-off between disk utilization and over-provisioning space, which affects the garbage collection process, as well as the user application performance. To address the performance degradation in write-intensive workloads caused by GC overhead, we propose a modified free segment-finding policy and a Parallel Garbage Collection (P-GC) scheme for F2FS that efficiently reduces GC overhead. Our evaluation results demonstrate that our P-GC scheme can achieve up to 42% performance enhancement with various workloads.
Dongjoo Seo, Ping-Xiang Chen, Huaicheng Li, Matias Bjørling, Nikil Dutt
HotStorage5
2023 Robust Detection of Social Isolation in Older Adults by Combining Biometrics with Social Interaction Data
abstract
Several recent studies, in the aftermath of Covid-19, point to dangers of social isolation that negatively impacts both mental and physical health especially amongst older adults. Isolation often leads to self-destructive behaviour such as drug and alcohol misuse deteriorating the quality of life and compounding health complications. Recent research has explored several technologies to detect the onset of isolation and to trigger interventions (e.g., nudge caregivers, etc.) to mitigate its impact. Such mechanisms span a range from using survey instruments, using biometrics to detect stress (an effect of isolation), and those using monitoring social interactions to detect loneliness amongst individuals. This paper studies biometric based and social interaction-based methods with the objective to understand their relative benefits/disadvantages and explores their combined usage to create a robust isolation detection mechanism. In particular, we explore the design of an integrated system based on both biometric (via wearables) and social interaction data (using call log analysis) to study their efficacy both individually and in combination.
Raghav Mehrotra-Venkat, Nikil Dutt, Julie Rousseau
SMARTCOMP2
2023 A Deep Learning-based PPG Quality Assessment Approach for Heart Rate and Heart Rate Variability
abstract
Photoplethysmography (PPG) is a non-invasive optical method to acquire various vital signs, including heart rate (HR) and heart rate variability (HRV). The PPG method is highly susceptible to motion artifacts and environmental noise. Unfortunately, such artifacts are inevitable in ubiquitous health monitoring, as the users are involved in various activities in their daily routines. Such low-quality PPG signals negatively impact the accuracy of the extracted health parameters, leading to inaccurate decision-making. PPG-based health monitoring necessitates a quality assessment approach to determine the signal quality according to the accuracy of the health parameters. Different studies have thus far introduced PPG signal quality assessment methods, exploiting various indicators and machine learning algorithms. These methods differentiate reliable and unreliable signals, considering morphological features of the PPG signal and focusing on the cardiac cycles. Therefore, they can be utilized for HR detection applications. However, they do not apply to HRV, as only having an acceptable shape is insufficient, and other signal factors may also affect the accuracy. In this article, we propose a deep learning–based PPG quality assessment method for HR and various HRV parameters. We employ one customized one-dimensional (1D) and three 2D Convolutional Neural Networks (CNN) to train models for each parameter. Reliability of each of these parameters will be evaluated against the corresponding electrocardiogram signal, using 210 hours of data collected from a home-based health monitoring application. Our results show that the proposed 1D CNN method outperforms the other 2D CNN approaches. Our 1D CNN model obtains the accuracy of 95.63%, 96.71%, 91.42%, 94.01%, and 94.81% for the HR, average of normal to normal interbeat (NN) intervals, root mean square of successive NN interval differences, standard deviation of NN intervals, and ratio of absolute power in low frequency to absolute power in high frequency ratios, respectively. Moreover, we compare the performance of our proposed method with state-of-the-art algorithms. We compare our best models for HR-HRV health parameters with six different state-of-the-art PPG signal quality assessment methods. Our results indicate that the proposed method performs better than the other methods. We also provide the open source model implemented in Python for the community to be integrated into their solutions.
Emad Kasaeyan Naeini, Fatemeh Sarhaddi, Iman Azimi, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani
ACM Trans. Comput. Heal.5
2023 EASYR: Energy-Efficient Adaptive System Reconfiguration for Dynamic Deadlines in Autonomous Driving on Multicore Processors
abstract
The increasing computing demands of autonomous driving applications have driven the adoption of multicore processors in real-time systems, which in turn renders energy optimizations critical for reducing battery capacity and vehicle weight. A typical energy optimization method targeting traditional real-time systems finds a critical speed under a static deadline, resulting in conservative energy savings that are unable to exploit dynamic changes in the system and environment. We capture emerging dynamic deadlines arising from the vehicle’s change in velocity and driving context for an additional energy optimization opportunity. In this article, we extend the preliminary work for uniprocessors [ 66 ] to multicore processors, which introduces several challenges. We use the state-of-the-art real-time gang scheduling [ 5 ] to mitigate some of the challenges. However, it entails an NP-hard combinatorial problem in that tasks need to be grouped into gangs of tasks, gang formation, which could significantly affect the energy saving result. As such, we present EASYR, an adaptive system optimization and reconfiguration approach that generates gangs of tasks from a given directed acyclic graph for multicore processors and dynamically adapts the scheduling parameters and processor speeds to satisfy dynamic deadlines while consuming as little energy as possible. The timing constraints are also satisfied between system reconfigurations through our proposed safe mode change protocol. Our extensive experiments with randomly generated task graphs show that our gang formation heuristic performs 32% better than the state-of-the-art one. Using an autonomous driving task set from Bosch and real-world driving data, our experiments show that EASYR achieves energy reductions of up to 30.3% on average in typical driving scenarios compared with a conventional energy optimization method with the current state-of-the-art gang formation heuristic in real-time systems, demonstrating great potential for dynamic energy optimization gains by exploiting dynamic deadlines.
Saehanseul Yi, Jongchan Kim 0001, Nikil Dutt
ACM Trans. Embed. Comput. Syst.4
2022 AMSER: Adaptive Multimodal Sensing for Energy Efficient and Resilient eHealth Systems
abstract
eHealth systems deliver critical digital healthcare and wellness services for users by continuously monitoring physiological and contextual data. eHealth applications use multi-modal machine learning kernels to analyze data from different sensor modalities and automate decision-making. Noisy inputs and motion artifacts during sensory data acquisition affect the i) prediction accuracy and resilience of eHealth services and ii) energy efficiency in processing garbage data. Monitoring raw sensory inputs to identify and drop data and features from noisy modalities can improve prediction accuracy and energy efficiency. We propose a closed-loop monitoring and control framework for multi-modal eHealth applications, AMSER, that can mitigate garbage-in garbage-out by i) monitoring input modalities, ii) analyzing raw input to selectively drop noisy data and features, and iii) choosing appropriate machine learning models that fit the configured data and feature vector - to improve prediction accuracy and energy efficiency. We evaluate our AMSER approach using multi-modal eHealth applications of pain assessment and stress monitoring over different levels and types of noisy components incurred via different sensor modalities. Our approach achieves up to 22% improvement in prediction accuracy and 5.6× energy consumption reduction in the sensing phase against the state-of-the-art multi-modal monitoring application.
Emad Kasaeyan Naeini, Sina Shahhosseini, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt
DATE6
2022 Flexible and Personalized Learning for Wearable Health Applications using HyperDimensional Computing
abstract
Health and wellness applications increasingly rely on machine learning techniques to learn end-user physiological and behavioral patterns in everyday settings, posing two key challenges: inability to perform on-device online learning for resource-constrained wearables, and learning algorithms that support privacy-preserving personalization. We exploit a Hyperdimensional computing (HDC) solution for wearable devices that offers flexibility, high efficiency, and performance while enabling on-device personalization and privacy protection. We evaluate the efficacy of our approach using three case studies and show that our system improves performance of training by up to 35.8x compared with the state-of-the-art while offering a comparable accuracy.
Sina Shahhosseini, Yang Ni 0001, Emad Kasaeyan Naeini, Mohsen Imani, Amir-Mohammad Rahmani, Nikil Dutt
ACM Great Lakes Symposium on VLSI6
2022 Composing Graphical Models with Generative Adversarial Networks for EEG Signal Modeling
abstract
Neural oscillations in the form of electroencephalogram (EEG) can reveal underlying brain functions, such as cognition, memory, perception, and consciousness. A comprehensive EEG computational model provides not only a stochastic procedure that directly generates data but also insights to further understand the neurological mechanisms. Here, we propose a generative and inference approach that combines the complementary benefits of probabilistic graphical models and generative adversarial networks (GANs) for EEG signal modeling. We investigate the method’s ability to jointly learn coherent generation and inverse inference models on the CHI-MIT epilepsy multi-channel EEG dataset. We further study the efficacy of the learned representations in epilepsy seizure detection formulated as an unsupervised learning problem. Quantitative and qualitative experimental results demonstrate the effectiveness and efficiency of our approach.
Khuong Vo, Manoj Vishwanath, Ramesh Srinivasan, Nikil Dutt, Hung Cao
ICASSP4
2022 Hyperdimensional Hybrid Learning on End-Edge-Cloud Networks
abstract
In this paper, we present Hyperdimensional Hybrid Learning (HDHL), which combines model-free and model-based Reinforcement Learning, to effectively reduce the computational cost and environment interaction for optimizing an intelligent cloud service. We first show that Hyperdimensional Q-Learning (QHD), the state-of-the-art Hyperdimensional Computing value-based Reinforcement Learning algorithm, is computationally faster than the Deep Q-Network (DQN) for this task. In addition, we demonstrate how HDHL reduces the number of environment interactions by 4.8× to learn the near optimal configuration. Our evaluation shows that HDHL is computationally more efficient than both Q-Learning algorithms, with the total time being reduced by 21.0× compared to DQN and 16.5× compared to QHD.
Mariam Issa, Sina Shahhosseini, Yang Ni 0001, Danny Abraham, Amir-Mohammad Rahmani, Nikil Dutt, Mohsen Imani
ICCD7
2022 CARLsim 6: An Open Source Library for Large-Scale, Biologically Detailed Spiking Neural Network Simulation
abstract
Mature simulation systems for Spiking Neural Networks (SNNs) become more relevant than ever for understanding the brain and supporting neuromorphic computing. The CARL-sim SNN platform is one of the first Open Source simulation systems that utilized CUDA GPUs to address the tremendous parallel processing demands of natural brains. It has evolved over almost a decade in numerous scientific research projects requiring efficient biologically plausible modeling at scale. With its sixth major release, CARLsim 6 respects this legacy by supporting the latest versions of operating systems, development tool chains, multi-core computers, and of course GPUs. It runs on a range of platforms; from Notebooks up to the NVIDIA DGX-A100 supercomputer, and is used in biologically plausible simulations of the hippocampus and neocortex. The latest version has added flexibility for incorporating long-term and short-term synaptic plasticity. Neuromodulation is an important property of neurobiology that can lead to rapid few shot learning, network rewiring, and neural activity modulation. Because of this, CARLsim 6 now supports four multiple neuromodulators for simulating neural excitability and synaptic plasticity.
Lars Niedermeier, Kexin Chen 0002, Jinwei Xing, Anup Das 0001, Jeffrey Kopsick, Eric Scott, Nate Sutton, Killian Weber, Nikil Dutt, Jeffrey L. Krichmar
IJCNN9
2022 ProSwap: Period-aware Proactive Swapping to Maximize Embedded Application Performance
abstract
Linux prevents errors due to physical memory limits by swapping out active application memory from main memory to secondary storage. Swapping degrades application performance due to swap-in/out latency overhead. To mitigate the swapping overhead in periodic applications, we present ProSwap: a period-aware proactive and adaptive swapping policy for em-bedded systems. ProSwap exploits application periodic behavior to proactively swap-out rarely-used physical memory pages, creating more space for active processes. A flexible memory reclamation time-window enables adaptation to memory limitations that vary between applications. We demonstrate ProSwap's efficacy for an autonomous vehicle application scenario executing multi-application pipelines, and show that our policy achieves up to 1.26×performance gain via proactive swapping.
Dongjoo Seo, Biswadip Maity, Ping-Xiang Chen, Dukyoung Yun, Bryan Donyanavard, Nikil Dutt
NAS6
2022 Demand Layering for Real-Time DNN Inference with Minimized Memory Usage
abstract
When executing a deep neural network (DNN), its model parameters are loaded into GPU memory before execution, incurring a significant GPU memory burden. There are studies that reduce GPU memory usage by exploiting CPU memory as a swap device. However, this approach is not applicable in most embedded systems with integrated GPUs where CPU and GPU share a common memory. In this regard, we present Demand Layering, which employs a fast solid-state drive (SSD) as a co-running partner of a GPU and exploits the layer-by-layer execution of DNNs. In our approach, a DNN is loaded and executed in a layer-by-layer manner, minimizing the memory usage to the order of a single layer. Also, we developed a pipeline architecture that hides most additional delays caused by the interleaved parameter loadings alongside layer executions. Our implementation shows a 96.5% memory reduction with just 14.8% delay overhead on average for representative DNNs. Furthermore, by exploiting the memory-delay tradeoff, near-zero delay overhead (under 1 ms) can be achieved with a slightly increased memory usage (still an 88.4% reduction), showing the great potential of Demand Layering.
Mingoo Ji, Saehanseul Yi, Changjin Koo, Sol Ahn, Dongjoo Seo, Nikil Dutt, Jongchan Kim 0001
RTSS6
2022 SIC-EDGE: Semantic Iterative ECG Compression for Edge-Assisted Wearable Systems
abstract
Wearable sensors and Internet of Things technologies are enabling automated health monitoring applications, where signals captured by sensors are analyzed in real-time by algorithms detecting health issues and conditions. However, continuous clinical-level monitoring of patients in everyday settings often requires computation, storage and connectivity capabilities beyond those possessed by wearable sensors. While edge computing partially resolves this issue by connecting the sensors to compute-capable devices positioned at the network edge, the wireless links connecting the sensors to the edge servers may not have sufficient capacity to transfer the information-rich data that characterize these applications. A possible solution is to compress the signal to be transferred, accepting the tradeoff between compression gain and detection accuracy. In this paper, we propose SIC-EDGE: a "semantic compression" framework whose goal is to dynamically optimize the resolution of an electrocardiogram (ECG) signal transferred from a wearable sensor to an edge server to perform real-time detection of heart diseases. The core idea is to establish a collaborative control loop between the sensor and the edge server to iteratively build a semantic representation that is: (i) ECG-cycle specific; (ii) personalized, and (iii) targeted to support the classification task rather than signal reconstruction. The core of SIC-EDGE is a Sequential Hypothesis Testing (SHT) algorithm that analyzes partial representations along the iterations to determine which and how many representation layers (wavelet coefficients in our implementation) are requested. Our results on established datasets demonstrates the need for adaptive "semantic" compression, and illustrate the dynamic compression strategy realized by SIC-EDGE. We show that SIC-EDGE leads to an increase in terms of recall and F1 score of up to 35% and 26% respectively compared to an optimized but static wavelet compression for a given maximum channel usage.
Delaram Amiri, Janne Takalo-Mattila, Luca Bedogni, Marco Levorato, Nikil Dutt
WoWMoM5
2022 Exploring computation offloading in IoT systems
abstract
Internet of Things (IoT) paradigm raises challenges for devising efficient strategies that offload applications to the fog or the cloud layer while ensuring the optimal response time for a service. Traditional computation offloading policies assume the response time is only dominated by the execution time. However, the response time is a function of many factors including contextual parameters and application characteristics that can change over time. For the computation offloading problem, the majority of existing literature presents efficient solutions considering a limited number of parameters (e.g., computation capacity and network bandwidth) neglecting the effect of the application characteristics and dataflow configuration. In this paper, we explore the impact of the computation offloading on total application response time in three-layer IoT systems considering more realistic parameters, e.g., application characteristics, system complexity, communication cost, and dataflow configuration. This paper also highlights the impact of a new application characteristic parameter defined as Output–Input Data Generation (OIDG) ratio and dataflow configuration on the system behavior. In addition, we present a proof-of-concept end-to-end dynamic computation offloading technique, implemented in a real hardware setup, that observes the aforementioned parameters to perform real-time decision-making.
Sina Shahhosseini, Arman Anzanpour, Iman Azimi, Sina Labbaf, Dongjoo Seo, Sung-Soo Lim, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani
Inf. Syst.8
2022 Online Learning for Orchestration of Inference in Multi-user End-edge-cloud Networks
abstract
Deep-learning-based intelligent services have become prevalent in cyber-physical applications, including smart cities and health-care. Deploying deep-learning-based intelligence near the end-user enhances privacy protection, responsiveness, and reliability. Resource-constrained end-devices must be carefully managed to meet the latency and energy requirements of computationally intensive deep learning services. Collaborative end-edge-cloud computing for deep learning provides a range of performance and efficiency that can address application requirements through computation offloading. The decision to offload computation is a communication-computation co-optimization problem that varies with both system parameters (e.g., network condition) and workload characteristics (e.g., inputs). However, deep learning model optimization provides another source of tradeoff between latency and model accuracy. An end-to-end decision-making solution that considers such computation-communication problem is required to synergistically find the optimal offloading policy and model for deep learning services. To this end, we propose a reinforcement-learning-based computation offloading solution that learns optimal offloading policy considering deep learning model selection techniques to minimize response time while providing sufficient accuracy. We demonstrate the effectiveness of our solution for edge devices in an end-edge-cloud system and evaluate with a real-setup implementation using multiple AWS and ARM core configurations. Our solution provides 35% speedup in the average response time compared to the state-of-the-art with less than 0.9% accuracy reduction, demonstrating the promise of our online learning framework for orchestrating DL inference in end-edge-cloud systems.
Sina Shahhosseini, Dongjoo Seo, Anil Kanduri, Sung-Soo Lim, Bryan Donyanavard, Amir-Mohammad Rahmani, Nikil Dutt
ACM Trans. Embed. Comput. Syst.8
2022 Endurance-Aware Mapping of Spiking Neural Networks to Neuromorphic Hardware
abstract
Neuromorphic computing systems are embracing memristors to implement high density and low power synaptic storage as crossbar arrays in hardware. These systems are energy efficient in executing Spiking Neural Networks (SNNs). We observe that long bitlines and wordlines in a memristive crossbar are a major source of parasitic voltage drops, which create current asymmetry. Through circuit simulations, we show the significant endurance variation that results from this asymmetry. Therefore, if the critical memristors (ones with lower endurance) are overutilized, they may lead to a reduction of the crossbar's lifetime. We propose eSpine, a novel technique to improve lifetime by incorporating the endurance variation within each crossbar in mapping machine learning workloads, ensuring that synapses with higher activation are always implemented on memristors with higher endurance, and vice versa. eSpine works in two steps. First, it uses the Kernighan-Lin Graph Partitioning algorithm to partition a workload into clusters of neurons and synapses, where each cluster can fit in a crossbar. Second, it uses an instance of Particle Swarm Optimization (PSO) to map clusters to tiles, where the placement of synapses of a cluster to memristors of a crossbar is performed by analyzing their activation within the workload. We evaluate eSpine for a state-of-the-art neuromorphic hardware model with phase-change memory (PCM)-based memristors. Using 10 SNN workloads, we demonstrate a significant improvement in the effective lifetime.
Twisha Titirsha, Shihao Song, Anup Das 0001, Jeffrey L. Krichmar, Nikil Dutt, Nagarajan Kandasamy, Francky Catthoor
IEEE Trans. Parallel Distributed Syst.5
2021 Energy-Efficient Adaptive System Reconfiguration for Dynamic Deadlines in Autonomous Driving
abstract
The increasing computing demands of autonomous driving applications make energy optimizations critical for reducing battery capacity and vehicle weight. Current energy optimization methods typically target traditional real-time systems with static deadlines, resulting in conservative energy savings that are unable to exploit additional energy optimizations due to dynamic deadlines arising from the vehicle's change in velocity and driving context. We present an adaptive system optimization and reconfiguration approach that dynamically adapts the scheduling parameters and processor speeds to satisfy dynamic deadlines while consuming as little energy as possible. Our experimental results with an autonomous driving task set from Bosch and realworld driving data show energy reductions up to 46.4% on average in typical dynamic driving scenarios compared with traditional static energy optimization methods, demonstrating great potential for dynamic energy optimization gains by exploiting dynamic deadlines.
Saehanseul Yi, Jongchan Kim 0001, Nikil Dutt
ISORC4
2021 Dynamic Reliability Management in Neuromorphic Computing
abstract
Neuromorphic computing systems execute machine learning tasks designed with spiking neural networks. These systems are embracing non-volatile memory to implement high-density and low-energy synaptic storage. Elevated voltages and currents needed to operate non-volatile memories cause aging of CMOS-based transistors in each neuron and synapse circuit in the hardware, drifting the transistor’s parameters from their nominal values. If these circuits are used continuously for too long, the parameter drifts cannot be reversed, resulting in permanent degradation of circuit performance over time, eventually leading to hardware faults. Aggressive device scaling increases power density and temperature, which further accelerates the aging, challenging the reliable operation of neuromorphic systems. Existing reliability-oriented techniques periodically de-stress all neuron and synapse circuits in the hardware at fixed intervals, assuming worst-case operating conditions, without actually tracking their aging at run-time. To de-stress these circuits, normal operation must be interrupted, which introduces latency in spike generation and propagation, impacting the inter-spike interval and hence, performance (e.g., accuracy). We observe that in contrast to long-term aging, which permanently damages the hardware, short-term aging in scaled CMOS transistors is mostly due to bias temperature instability. The latter is heavily workload-dependent and, more importantly, partially reversible. We propose a new architectural technique to mitigate the aging-related reliability problems in neuromorphic systems by designing an intelligent run-time manager (NCRTM), which dynamically de-stresses neuron and synapse circuits in response to the short-term aging in their CMOS transistors during the execution of machine learning workloads, with the objective of meeting a reliability target. NCRTM de-stresses these circuits only when it is absolutely necessary to do so, otherwise reducing the performance impact by scheduling de-stress operations off the critical path. We evaluate NCRTM with state-of-the-art machine learning workloads on a neuromorphic hardware. Our results demonstrate that NCRTM significantly improves the reliability of neuromorphic hardware, with marginal impact on performance.
Shihao Song, Jui Hanamshet, Adarsha Balaji, Anup Das 0001, Jeffrey L. Krichmar, Nikil Dutt, Nagarajan Kandasamy, Francky Catthoor
ACM J. Emerg. Technol. Comput. Syst.6
2021 SEAMS: Self-Optimizing Runtime Manager for Approximate Memory Hierarchies
abstract
Memory approximation techniques are commonly limited in scope, targeting individual levels of the memory hierarchy. Existing approximation techniques for a full memory hierarchy determine optimal configurations at design-time provided a goal and application. Such policies are rigid: they cannot adapt to unknown workloads and must be redesigned for different memory configurations and technologies. We propose SEAMS: the first self-optimizing runtime manager for coordinating configurable approximation knobs across all levels of the memory hierarchy. SEAMS continuously updates and optimizes its approximation management policy throughout runtime for diverse workloads. SEAMS optimizes the approximate memory configuration to minimize energy consumption without compromising the quality threshold specified by application developers. SEAMS can (1) learn a policy at runtime to manage variable application quality of service ( QoS ) constraints, (2) automatically optimize for a target metric within those constraints, and (3) coordinate runtime decisions for interdependent knobs and subsystems. We demonstrate SEAMS’ ability to efficiently provide functions (1)–(3) on a RISC-V Linux platform with approximate memory segments in the on-chip cache and main memory. We demonstrate SEAMS’ ability to save up to 37% energy in the memory subsystem without any design-time overhead. We show SEAMS’ ability to reduce QoS violations by 75% with < 5% additional energy.
Biswadip Maity, Bryan Donyanavard, Anmol Surhonne, Amir-Mohammad Rahmani, Andreas Herkersdorf, Nikil Dutt
ACM Trans. Embed. Comput. Syst.6
2021 Chauffeur: Benchmark Suite for Design and End-to-End Analysis of Self-Driving Vehicles on Embedded Systems
abstract
Self-driving systems execute an ensemble of different self-driving workloads on embedded systems in an end-to-end manner, subject to functional and performance requirements. To enable exploration, optimization, and end-to-end evaluation on different embedded platforms, system designers critically need a benchmark suite that enables flexible and seamless configuration of self-driving scenarios, which realistically reflects real-world self-driving workloads’ unique characteristics. Existing CPU and GPU embedded benchmark suites typically (1) consider isolated applications, (2) are not sensor-driven, and (3) are unable to support emerging self-driving applications that simultaneously utilize CPUs and GPUs with stringent timing requirements. On the other hand, full-system self-driving simulators (e.g., AUTOWARE, APOLLO) focus on functional simulation, but lack the ability to evaluate the self-driving software stack on various embedded platforms. To address design needs, we present Chauffeur, the first open-source end-to-end benchmark suite for self-driving vehicles with configurable representative workloads. Chauffeur is easy to configure and run, enabling researchers to evaluate different platform configurations and explore alternative instantiations of the self-driving software pipeline. Chauffeur runs on diverse emerging platforms and exploits heterogeneous onboard resources. Our initial characterization of Chauffeur on different embedded platforms – NVIDIA Jetson TX2 and Drive PX2 – enables comparative evaluation of these GPU platforms in executing an end-to-end self-driving computational pipeline to assess the end-to-end response times on these emerging embedded platforms while also creating opportunities to create application gangs for better response times. Chauffeur enables researchers to benchmark representative self-driving workloads and flexibly compose them for different self-driving scenarios to explore end-to-end tradeoffs between design constraints, power budget, real-time performance requirements, and accuracy of applications.
Biswadip Maity, Saehanseul Yi, Dongjoo Seo, Leming Cheng, Sung-Soo Lim, Jongchan Kim 0001, Bryan Donyanavard, Nikil Dutt
ACM Trans. Embed. Comput. Syst.8
2021 An Interpretable Machine Learning Model Enhanced Integrated CPU-GPU DVFS Governor
abstract
Modern heterogeneous CPU-GPU-based mobile architectures, which execute intensive mobile gaming/graphics applications, use software governors to achieve high performance with energy-efficiency. However, existing governors typically utilize simple statistical or heuristic models, assuming linear relationships using a small unbalanced dataset of mobile games; and the limitations result in high prediction errors for dynamic and diverse gaming workloads on heterogeneous platforms. To overcome these limitations, we propose an interpretable machine learning (ML) model enhanced integrated CPU-GPU governor: (1) It builds tree-based piecewise linear models (i.e., model trees) offline considering both high accuracy (low error) and interpretable ML models based on mathematical formulas using a simulatability operation counts quantitative metric. And then (2) it deploys the selected models for online estimation into an integrated CPU-GPU Dynamic Voltage Frequency Scaling governor. Our experiments on a test set of 20 mobile games exhibiting diverse characteristics show that our governor achieved significant energy efficiency gains of over 10% (up to 38%) improvements on average in energy-per-frame with a surprising-but-modest 3% improvement in Frames-per-Second performance, compared to a typical state-of-the-art governor that employs simple linear regression models.
Jurn-Gyu Park, Nikil Dutt, Sung-Soo Lim
ACM Trans. Embed. Comput. Syst.2
2020 CryptoPIM: In-memory Acceleration for Lattice-based Cryptographic Hardware
abstract
Quantum computers promise to solve hard mathematical problems such as integer factorization and discrete logarithms in polynomial time, making standardized public-key cryptosystems insecure. Lattice-Based Cryptography (LBC) is a promising post-quantum public key cryptographic protocol that could replace standardized public key cryptography, thanks to the inherent post-quantum resistant properties, efficiency, and versatility. A key mathematical tool in LBC is the Number Theoretic Transform (NTT), a common method to compute polynomial multiplication. It is the most compute-intensive routine and requires acceleration for practical deployment of LBC protocols. In this paper, we propose CryptoPIM, a high-throughput Processing In-Memory (PIM) accelerator for NTT-based polynomial multiplier with the support of polynomials with degrees up to 32k. Compared to the fastest FPGA implementation of an NTT-based multiplier, CryptoPIM achieves on average 31x throughput improvement with the same energy and only 28% performance reduction, thereby showing promise for practical deployment of LBC.
Hamid Nejatollahi, Saransh Gupta, Mohsen Imani, Tajana Rosing, Rosario Cammarota, Nikil Dutt
DAC6
2020 Emergent Control of MPSoC Operation by a Hierarchical Supervisor / Reinforcement Learning Approach
abstract
MPSoCs increasingly depend on adaptive resource management strategies at runtime for efficient utilization of resources when executing complex application workloads. In particular, conflicting demands for adequate computation performance and power-/energy-efficiency constraints make desired application goals hard to achieve. We present a hierarchical, cross-layer hardware/software resource manager capable of adapting to changing workloads and system dynamics with zero initial knowledge. The manager uses rule-based reinforcement learning classifier tables (LCTs) with an archive-based backup policy as leaf controllers. The LCTs directly manipulate and enforce MPSoC building block operation parameters in order to explore and optimize potentially conflicting system requirements (e.g., meeting a performance target while staying within the power constraint). A supervisor translates system requirements and application goals into per-LCT objective functions (e.g., core instructions-per-second (IPS). Thus, the supervisor manages the possibly emergent behavior of the low-level LCT controllers in response to 1) switching between operation strategies (e.g., maximize performance vs. minimize power; and 2) changing application requirements. This hierarchical manager leverages the dual benefits of a software supervisor (enabling flexibility), together with hardware learners (allowing quick and efficient optimization). Experiments on an FPGA prototype confirmed the ability of our approach to identify optimized MPSoC operation parameters at runtime while strictly obeying given power constraints.
Florian Maurer 0003, Bryan Donyanavard, Amir-Mohammad Rahmani, Nikil Dutt, Andreas Herkersdorf
DATE4
2020 Exploring Energy Efficient Quantum-resistant Signal Processing Using Array Processors
abstract
Quantum computers threaten to break public-key cryptography schemes such as DSA and ECDSA in polynomial time, which poses an imminent threat to secure signal processing. Ring learning with error (RLWE) lattice-based cryptography (LBC) is one of the most promising families of post-quantum cryptography (PQC) schemes in terms of efficiency and versatility. Two conventional methods to compute polynomial multiplication, the most compute-intensive routine in the RLWE schemes, are convolutions and Number Theoretic Transform (NTT).In this work, we explore the energy efficiency of polynomial multiplier using systolic architecture for the first time. As an early exploration, we design two high-throughput systolic array polynomial multipliers, including NTT-based and convolution-based, and compare them to our low-cost sequential (non-systolic) NTT-based multiplier. Our sequential NTT-based multiplier achieves 3x speedup over the state-of-the-art FGPA implementation of the polynomial multiplier in the NewHope-Simple key exchange mechanism on a low-cost Artix7 FPGA. When synthesized on a Zynq UltraScale+ FPGA, the NTT-based systolic and convolution-based systolic designs achieve on average 1.7x and 7.5x speedup over our sequential NTT-based multiplier respectively, which can lead to generating over 2x more signatures per second by CRYSTALS-Dilithium, a PQC digital signature scheme. These explorations help designers select the right PQC implementations for making future signal processing applications quantum-resistant.
Hamid Nejatollahi, Sina Shahhosseini, Rosario Cammarota, Nikil Dutt
ICASSP4
2020 PyCARL: A PyNN Interface for Hardware-Software Co-Simulation of Spiking Neural Network
abstract
We present PyCARL, a PyNN-based common Python programming interface for hardware-software cosimulation of spiking neural network (SNN). Through PyCARL, we make the following two key contributions. First, we provide an interface of PyNN to CARLsim, a computationally- efficient, GPU-accelerated and biophysically-detailed SNN simulator. PyCARL facilitates joint development of machine learning models and code sharing between CARLsim and PyNN users, promoting an integrated and larger neuromorphic community. Second, we integrate cycle-accurate models of state-of-the-art neuromorphic hardware such as TrueNorth, Loihi, and DynapSE in PyCARL, to accurately model hardware latencies, which delay spikes between communicating neurons, degrading performance of machine learning models. PyCARL allows users to analyze and optimize the performance difference between software-based simulation and hardware-oriented simulation. We show that system designers can also use PyCARL to perform design-space exploration early in the product development stage, facilitating faster time-to-market of neuromorphic products.
Adarsha Balaji, Prathyusha Adiraju, Hirak J. Kashyap, Anup Das 0001, Jeffrey L. Krichmar, Nikil Dutt, Francky Catthoor
IJCNN6
2020 STINT: selective transmission for low-energy physiological monitoring
abstract
Noninvasive, and continuous physiological sensing enabled by novel wearable sensors is generating unprecedented diagnostic insights in many medical practices. However, the limited battery capacity of these wearable sensors poses a critical challenge in extending device lifetime in order to prevent omission of informative events. In this work, we exploit the inherent sparsity of physiological signals to intelligently enable selective transmission of these signals and thereby improve the energy efficiency of wearable sensors. We propose STINT, a selective transmission framework that generates a sparse representation of the raw signal based on domain-specific knowledge, and which can be integrated into a wide range of resource-constrained embedded sensing IoT platforms. STINT employs a neural network (NN) for selective transmission: the NN identifies, and transmits only the informative parts of the raw signal, thereby achieving low power operation. We validate STINT and establish its efficacy in the domain of IoT for energy-efficient physiological monitoring, by testing our framework on EcoBP - a novel miniaturized, and wireless continuous blood pressure sensor. Early experimental results on the EcoBP device demonstrate that the STINT-enabled EcoBP sensor outperforms the native platform by 14% of sensor energy consumption, with room for additional energy savings via complementary bluetooth and wireless optimizations.
Tao-Yi Lee, Khuong Vo, Wongi Baek, Michelle Khine, Nikil Dutt
ISLPED5
2020 R-TOD: Real-Time Object Detector with Minimized End-to-End Delay for Autonomous Driving
abstract
For realizing safe autonomous driving, the end-to-end delays of real-time object detection systems should be thoroughly analyzed and minimized. However, despite recent development of neural networks with minimized inference delays, surprisingly little attention has been paid to their end-to-end delays from an object's appearance until its detection is reported. With this motivation, this paper aims to provide more comprehensive understanding of the end-to-end delay, through which precise best- and worst-case delay predictions are formulated, and three optimization methods are implemented: (i) on-demand capture, (ii) zero-slack pipeline, and (iii) contention-free pipeline. Our experimental results show a 76% reduction in the end-to-end delay of Darknet YOLO (You Only Look Once) v3 (from 1070 ms to 261 ms), thereby demonstrating the great potential of exploiting the end-to-end delay analysis for autonomous driving. Furthermore, as we only modify the system architecture and do not change the neural network architecture itself, our approach incurs no penalty on the detection accuracy.
Wonseok Jang, Hansaem Jeong, Kyungtae Kang, Nikil Dutt, Jongchan Kim 0001
RTSS4
2020 Context-Aware Sensing via Dynamic Programming for Edge-Assisted Wearable Systems
abstract
Healthcare applications supported by the Internet of Things enable personalized monitoring of a patient in everyday settings. Such applications often consist of battery-powered sensors coupled to smart gateways at the edge layer. Smart gateways offer several local computing and storage services (e.g., data aggregation, compression, local decision making), and also provide an opportunity for implementing local closed-loop optimization of different parameters of the sensor layer, particularly energy consumption. To implement efficient optimization methods, information regarding the context and state of patients need to be considered to find opportunities to adjust energy to demanded accuracy. Edge-assisted optimization can manage energy consumption of the sensor layer but may also adversely affect the quality of sensed data, which could compromise the reliable detection of health deterioration risk factors. In this article, we propose two approaches: myopic and Markov decision processes (MDPs)—to consider both energy constraints and risk factor requirements for achieving a twofold goal: energy savings while satisfying accuracy requirements of abnormality detection in a patient’s vital signs. Vital signs, including heart rate, respiration rate, and oxygen saturation, are extracted from a photoplethysmogram signal and errors of extracted features are compared to a ground truth that is modeled as a Gaussian distribution. We control the sensor’s sensing energy to minimize the power consumption while meeting a desired level of satisfactory detection performance. We present experimental results on realistic case studies using a reconfigurable photoplethysmogram sensor in an IoT system, and show that compared to nonadaptive methods, myopic reduces an average of 16.9% in sensing energy consumption with the maximum probability of abnormality misdetection on the order of 0.17 in a 24-hour health monitoring system. In addition, over 4 weeks of monitoring, we demonstrate that our MDP policy can extend the battery life on average of more than 2x while fulfilling the same average probability of misdetection compared to the myopic method. We illustrate results comparing myopic , MDP, and nonadaptive methods to monitor 14 subjects over 1 month.
Delaram Amiri, Arman Anzanpour, Iman Azimi, Marco Levorato, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani
ACM Trans. Comput. Heal.6
2020 Self-Awareness for Autonomous Systems
abstract
The articles in this month’s special issue cover concepts and fundamentals, architectures and techniques, and applications and case studies in the exciting area of self-awareness in autonomous systems.
Nikil Dutt, Carlo S. Regazzoni, Bernhard Rinner, Xin Yao 0001
Proc. IEEE1
2020 Embodied Self-Aware Computing Systems
abstract
Embodied self-aware computing systems are embedded in a physical environment with a rich set of sensors and actuators to interact both with their environment and with their own embodiment. Through this interaction, they learn about their situation, their own state, and their performance. Although they are application specific like traditional embedded systems (ESs), they are significantly more flexible, robust, and autonomous; they can adapt to a wide range of environmental variation and can cope with deterioration and shortcomings of their own performance. As such, embodied self-aware computing systems are an evolution of traditional embedded and cyber-physical systems into the direction of more autonomy, robustness, and flexibility. When traditional ESs operate in a changing world by demanding unchanging and fully characterized computing resources, embodied self-aware computing systems adapt to a changing world and changing computing resources. This article surveys the methods and methodologies used for embodied self-aware computing systems structured along with the faculties of: 1) sensory observation and abstraction; 2) self-aware assessment; and 3) hierarchical goals and control. The discussion is exemplified by application cases in the areas of systems-on-chip, control systems, health monitoring, and condition monitoring in industrial production systems.
Henry Hoffmann, Axel Jantsch, Nikil Dutt
Proc. IEEE3
2020 CAST: Content-Aware STT-MRAM Cache Write Management for Different Levels of Approximation
abstract
Spin transfer torque magnetic RAM (STT-MRAM) technology is one of the most promising alternative for static RAM (SRAM) for implementing on-chip memories. Compared with SRAMs, STT-MRAMs benefit from higher density and near-zero leakage power, nonetheless they impose high energy consumption for reliable write operations. However, in many applications, absolute data integrity is not required; thus, acting on the current applied in the write operations may represent a novel knob for disciplined approximate computing to obtain energy saving with a minimal quality loss in applications' outputs. This article proposes CAST, a hardware/software approach to adjust the energy/quality of write operations in STT-MRAM caches in multicore systems based on the content of requested write operations. CAST utilizes fine-grained cache-line-level actuation knobs with different levels of quality for individual write operations. This unique feature of STT-MRAMs allows to avoid interapplication actuation interference suffered by SRAMs, and makes the approach particularly suitable for systems running multiple applications with mixed accuracy sensitivity. Moreover, CAST exploits another peculiarity of STT-MRAMs represented by the asymmetry and transition-dependency of the write error rate, to further tune in a fine-grained manner the write current to achieve an additional energy saving, even in full-accurate applications. Our evaluations on workloads of full-approximate, mixed-criticality, and full-accurate applications demonstrate up to 57%, 34%, and 21% energy savings over a baseline STT-MRAM cache, respectively, with an acceptable quality of the generated outputs.
Amir Mahdi Hosseini Monazzah, Amir-Mohammad Rahmani, Antonio Miele, Nikil Dutt
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 Data Reuse for Accelerated Approximate Warps
abstract
Many data-driven applications, including computer vision, machine learning, speech recognition, and medical diagnostics show tolerance to computation error. These applications are often accelerated on GPUs, but the performance improvements require high energy usage. In this article, we present DRAAW, an approximate computing technique capable of accelerating GPGPU applications at a warp level. In GPUs, warps are groups of threads which issued together across multiple cores. The slowest thread dictates the pace of the warp, so DRAAW identifies these bottlenecks and avoids them during approximation. We alleviate computation costs by using an approximate lookup table which tracks recent operations and reuses them to exploit temporal locality within applications. To improve neural network performance, we propose neuron aware approximation, a technique which profiles operations within network layers and automatically configures DRAAW to ensure computations with more impact on the output accuracy are subject to less approximation. We evaluate our design by placing DRAAW within each core of an Nvidia Kepler Architecture Titan. DRAAW improves throughput by up to 2.8× and improves energy-delay product (EDP) by 5.6× for six GPGPU applications while maintaining less than 5% output error. We show neuron aware approximation accelerates the inference of six neutral networks by 2.9× and improves EDP by 6.2× with less than 1% impact on prediction accuracy.
Daniel Peroni, Mohsen Imani, Hamid Nejatollahi, Nikil Dutt, Tajana Rosing
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 Self-aware Cyber-Physical Systems
abstract
In this article, we make the case for the new class of Self-aware Cyber-physical Systems. By bringing together the two established fields of cyber-physical systems and self-aware computing, we aim at creating systems with strongly increased yet managed autonomy, which is a main requirement for many emerging and future applications and technologies. Self-aware cyber-physical systems are situated in a physical environment and constrained in their resources, and they understand their own state and environment and, based on that understanding, are able to make decisions autonomously at runtime in a self-explanatory way. In an attempt to lay out a research agenda, we bring up and elaborate on five key challenges for future self-aware cyber-physical systems: (i) How can we build resource-sensitive yet self-aware systems? (ii) How to acknowledge situatedness and subjectivity? (iii) What are effective infrastructures for implementing self-awareness processes? (iv) How can we verify self-aware cyber-physical systems and, in particular, which guarantees can we give? (v) What novel development processes will be required to engineer self-aware cyber-physical systems? We review each of these challenges in some detail and emphasize that addressing all of them requires the system to make a comprehensive assessment of the situation and a continual introspection of its own state to sensibly balance diverse requirements, constraints, short-term and long-term objectives. Throughout, we draw on three examples of cyber-physical systems that may benefit from self-awareness: a multi-processor system-on-chip, a Mars rover, and an implanted insulin pump. These three very different systems nevertheless have similar characteristics: limited resources, complex unforeseeable environmental dynamics, high expectations on their reliability, and substantial levels of risk associated with malfunctioning. Using these examples, we discuss the potential role of self-awareness in both highly complex and rather more simple systems, and as a main conclusion we highlight the need for research on above listed topics.
Kirstie L. Bellman, Christopher Landauer, Nikil Dutt, Lukas Esterle, Andreas Herkersdorf, Axel Jantsch, Nima Taherinejad, Peter R. Lewis 0001, Marco Platzner, Kalle Tammemäe
ACM Trans. Cyber Phys. Syst.3
2020 Introduction to the Special Issue on Self-Aware Cyber-physical Systems
abstract
No abstract available.
Axel Jantsch, Peter R. Lewis 0001, Nikil Dutt
ACM Trans. Cyber Phys. Syst.3
2020 Synthesis of Flexible Accelerators for Early Adoption of Ring-LWE Post-quantum Cryptography
abstract
The advent of the quantum computer makes current public-key infrastructure insecure. Cryptography community is addressing this problem by designing, efficiently implementing, and evaluating novel public-key algorithms capable of withstanding quantum computational power. Governmental agencies, such as NIST, are promoting standardization of quantum-resistant algorithms that is expected to run for 7 years. Several modern applications must maintain permanent data secrecy; therefore, they ultimately require the use of quantum-resistant algorithms. Because algorithms are still under scrutiny for eventual standardization, the deployment of the hardware implementation of quantum-resistant algorithms is still in early stages. In this article, we propose a methodology to design programmable hardware accelerators for lattice-based algorithms, and we use the proposed methodology to implement flexible and energy efficient post-quantum cache-based accelerators for NewHope , Kyber , Dilithium , Key Consensus from Lattice ( KCL ), and R.EMBLEM submissions to the NIST standardization contest. To the best of our knowledge, we propose the first efficient domain-specific, programmable cache-based accelerators for lattice-based algorithms. We design a single accelerator for a common kernel among various schemes with different kernel sizes, i.e., loop count, and data types. This is in contrast to the traditional approach of designing one special purpose accelerators for each scheme. We validate our methodology by integrating our accelerators into an HLS-based SoC infrastructure based on the X86 processor and evaluate overall performance. Our experiments demonstrate the suitability of the approach and allow us to collect insightful information about the performance bottlenecks and the energy efficiency of the explored algorithms. Our results provide guidelines for hardware designers, highlighting the optimization points to address for achieving the highest energy minimization and performance increase. At the same time, our proposed design allows us to specify and execute new variants of lattice-based schemes with superior energy efficiency compared to the main application processor without changing the hardware acceleration platform. For example, we manage to reduce the energy consumption up to 2.1× and energy-delay product (EDP) up to 5.2× and improve the speedup up to 2.5×.
Hamid Nejatollahi, Felipe Valencia, Subhadeep Banik, Francesco Regazzoni 0001, Rosario Cammarota, Nikil Dutt
ACM Trans. Embed. Comput. Syst.6
2020 Edge-Assisted Control for Healthcare Internet of Things: A Case Study on PPG-Based Early Warning Score
abstract
Recent advances in pervasive Internet of Things technologies and edge computing have opened new avenues for development of ubiquitous health monitoring applications. Delivering an acceptable level of usability and accuracy for these healthcare Internet of Things applications requires optimization of both system-driven and data-driven aspects, which are typically done in a disjoint manner. Although decoupled optimization of these processes yields local optima at each level, synergistic coupling of the system and data levels can lead to a holistic solution opening new opportunities for optimization. In this article, we present an edge-assisted resource manager that dynamically controls the fidelity and duration of sensing w.r.t. changes in the patient’s activity and health state, thus fine-tuning the trade-off between energy efficiency and measurement accuracy. The cornerstone of our proposed solution is an intelligent low-latency real-time controller implemented at the edge layer that detects abnormalities in the patient’s condition and accordingly adjusts the sensing parameters of a reconfigurable wireless sensor node. We assess the efficiency of our proposed system via a case study of the photoplethysmography-based medical early warning score system. Our experiments on a real full hardware-software early warning score system reveal up to 49% power savings while maintaining the accuracy of the sensory data.
Arman Anzanpour, Delaram Amiri, Iman Azimi, Marco Levorato, Nikil Dutt, Pasi Liljeberg, Amir-Mohammad Rahmani
ACM Trans. Internet Things5
2020 Mapping Spiking Neural Networks to Neuromorphic Hardware
abstract
Neuromorphic hardware implements biological neurons and synapses to execute a spiking neural network (SNN)-based machine learning. We present SpiNeMap, a design methodology to map SNNs to crossbar-based neuromorphic hardware, minimizing spike latency and energy consumption. SpiNeMap operates in two steps: SpiNeCluster and SpiNePlacer. SpiNeCluster is a heuristic-based clustering technique to partition an SNN into clusters of synapses, where intracluster local synapses are mapped within crossbars of the hardware and intercluster global synapses are mapped to the shared interconnect. SpiNeCluster minimizes the number of spikes on global synapses, which reduces spike congestion and improves application performance. SpiNePlacer then finds the best placement of local and global synapses on the hardware using a metaheuristic-based approach to minimize energy consumption and spike latency. We evaluate SpiNeMap using synthetic and realistic SNNs on a state-of-the-art neuromorphic hardware. We show that SpiNeMap reduces average energy consumption by 45% and spike latency by 21%, compared to the best-performing SNN mapping technique.
Adarsha Balaji, Francky Catthoor, Anup Das 0001, Yuefeng Wu, Khanh Huynh, Francesco Dell'Anna, Giacomo Indiveri, Jeffrey L. Krichmar, Nikil Dutt, Siebren Schaafsma
IEEE Trans. Very Large Scale Integr. Syst.9
2019 ARGA: Approximate Reuse for GPGPU Acceleration
abstract
Many data-driven applications including computer vision, speech recognition, and medical diagnostics show tolerance to error during computation. These applications are often accelerated on GPUs, but high computational costs limit performance and increase energy usage. In this paper, we present ARGA, an approximate computing technique capable of accelerating GPGPU applications. ARGA provides an approximate lookup table to GPGPU cores to avoid recomputing instructions with identical or similar values. We propose multi-table parallel lookup which enables computational reuse to significantly speed-up GPGPU computation by checking incoming instructions in parallel. The inputs of each operation are searched for in a lookup table. Matches resulting in an exact or low error are removed from the floating point pipeline and used directly as output. Matches producing highly inaccurate results are computed on exact hardware to minimize application error. We simulate our design by placing ARGA within each core of an Nvidia Kepler Architecture Titan and an AMD Southern Island 7970. We show our design improves performance throughput by up to 2.7× and improves EDP by 5.3× for 6 GPGPU applications while maintaining less than 5% output error. We also show ARGA accelerates inference of a LeNet NN by 2.1× and improves EDP by 3.7× without significantly impacting classification accuracy.
Daniel Peroni, Mohsen Imani, Hamid Nejatollahi, Nikil Dutt, Tajana Rosing
DAC4
2019 The Case for Exploiting Underutilized Resources in Heterogeneous Mobile Architectures
abstract
Heterogeneous architectures are ubiquitous in mobile platforms, with mobile SoCs typically integrating multiple processors along with accelerators such as GPUs (for data-parallel kernels) and DSPs (for signal processing kernels). This strict partitioning of application execution on heterogeneous compute resources often results in underutilization of resources such as DSPs. We present a case study executing a mix of popular data-parallel workloads such as convolutional neural networks (CNNs), computer vision filters and graphics rendering kernels on mobile devices, and show that both performance and energy consumption of mobile platforms can be improved by synergistically deploying these underutilized compute resources. Our experiments on a mobile Snapdragon 835 platform under both single and multiple application scenarios executing the aforementioned workloads demonstrates average performance and energy improvements of 15-46% and 18-80%, respectively, by synergistically deploying all available compute resources, especially the underutilized DSP.
Chen-Ying Hsieh, Ardalan Amiri Sani, Nikil Dutt
DATE3
2019 Goal-Driven Autonomy for Efficient On-chip Resource Management: Transforming Objectives to Goals
abstract
Run-time resource allocation of heterogeneous multi-core systems is challenging with varying workloads and limited power and energy budgets. User interaction within these systems changes the performance requirements, often conflicting with concurrent applications' objective and system constraints. Current resource allocation approaches focus on optimizing fixed objective, ignoring the variation in system and applications' objective at run-time. For an efficient resource allocation, the system has to operate autonomously by formulating a hierarchy of goals. We present goal-driven autonomy (GDA) for on-chip resource allocation decisions, which allows systems to generate and prioritize goals in response to the workload and system dynamic variation. We implemented a proof-of-concept resource management framework that integrates the proposed goal management control to meet power, performance and user requirements simultaneously. Experimental results on an Exynos platform containing ARM's big.LITTLE-based heterogeneous multi-processor (HMP) show the effectiveness of GDA in efficient resource allocation in comparison with existing fixed objective policies.
Elham Shamsa, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt
DATE6
2019 DNN-Assisted Sensor for Energy-Efficient ECG Monitoring
abstract
The quasi-periodic nature of electrocardiogram (ECG) signals enables the use of compression techniques to minimize communications and reduce energy intake for diagnostic and preventive health monitoring. However, compression often degrades signal quality and may impair analysis by means of machine learning algorithms for the detection of anomalies. In this paper, we present an approach to pre-select relevant portions of the ECG signal at the sensor to reduce network load while satisfying a predefined diagnostic sensitivity requirement. We deploy a Deep Neural Network (DNN) to filter-out the signal's normal rhythms and reduce the amount of data stored or transmitted for further processing. Our extensive experiments covering a wide range of DNN hyper-parameters illustrate the tradeoff between diagnostic sensitivity, channel usage, energy consumption and computational complexity.
Tao-Yi Lee, Marco Levorato, Nikil Dutt
GLOBECOM3
2019 Dynamic Computation Migration at the Edge: Is There an Optimal Choice?
abstract
In the era of Fog computing where one can decide to compute certain time-critical tasks at the edge of the network, designers often encounter a question whether the sensor layer provides the optimal response time for a service, or the Fog layer, or their combination. In this context, minimizing the total response time using computation migration is a communication-computation co-optimization problem as the response time does not depend only on the computational capacity of each side. In this paper, we aim at investigating this question and addressing it in certain situations. We formulate this question as a static or dynamic computation migration problem depending on whether certain communication and computation characteristics of the underlying system is known at design-time or not. We first propose a static approach to find the optimal computation migration strategy using models known at design-time. We then make a more realistic assumption that several sources of variation can affect the system's response latency (e.g., the change in computation time, bandwidth, transmission channel reliability, etc.), and propose a dynamic computation migration approach which can adaptively identify the latency optimal computation layer at runtime. We evaluate our solution using a case-study of artificial neural network based arrhythmia classification using a simulation environment as well as a real test-bed.
Sina Shahhosseini, Iman Azimi, Arman Anzanpour, Axel Jantsch, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani
ACM Great Lakes Symposium on VLSI6
2019 Flexible NTT Accelerators for RLWE Lattice-Based Cryptography
abstract
In this work, we propose methods to design flexible and energy-efficient hardware accelerators for ring learning with error (RLWE) lattice-based cryptographic protocols, such as key agreement and digital signature. We apply the proposed methods to design the first programmable DMA-based family of accelerators for the Number Theoretic Transform (NTT), a commonly used kernel inside variants of RLWE protocols NewHope and Kyber. We validate our methods by integrating the accelerators into an HLS-based System on Chip (SoC) simulator. Experiments confirm the suitability of the flexible DMA-based accelerators for their use as part of lattice-based schemes. Our proposed designs are capable of executing new variants of lattice-based schemes with superior energy efficiency compared to executing the scheme entirely on the main processor, but without modifying the hardware acceleration platform. Performance improvements are up to 2x, energy consumption improves up to 1.9x, and energy-delay product (EDP) improves up to 3.9x. Together with such improved energy efficiency and performance, the flexibility inherent in our accelerators provides insights for, while reducing the risk of early adoption of lattice-based PQC cryptographic protocols in hardware products.
Hamid Nejatollahi, Rosario Cammarota, Nikil Dutt
ICCD3
2019 SOSA: Self-Optimizing Learning with Self-Adaptive Control for Hierarchical System-on-Chip Management
abstract
Resource management strategies for many-core systems dictate the sharing of resources among applications such as power, processing cores, and memory bandwidth in order to achieve system goals. System goals require consideration of both system constraints (e.g., power envelope) and user demands (e.g., response time, energy-efficiency). Existing approaches use heuristics, control theory, and machine learning for resource management. They all depend on static system models, requiring a priori knowledge of system dynamics, and are therefore too rigid to adapt to emerging workloads or changing system dynamics.
Bryan Donyanavard, Tiago Rogério Mück, Amir-Mohammad Rahmani, Nikil Dutt, Armin Sadighi, Florian Maurer 0003, Andreas Herkersdorf
MICRO4
2019 SURF: Self-aware Unified Runtime Framework for Parallel Programs on Heterogeneous Mobile Architectures
abstract
The following topics are dealt with: multiprocessing systems; microprocessor chips; logic design; power aware computing; field programmable gate arrays; integrated circuit design; low-power electronics; integrated circuit reliability; logic gates; radiation hardening (electronics).
Chen-Ying Hsieh, Ardalan Amiri Sani, Nikil Dutt
VLSI-SoC3
2019 The power impact of hardware and software actuators on self-adaptable many-core systems
Andre L. M. Martins, Rafael Garibotti, Nikil Dutt, Fernando Gehm Moraes
J. Syst. Archit.3
2019 Hierarchical adaptive Multi-objective resource management for many-core systems
Andre L. M. Martins, Alzemiro Henrique Lucas da Silva, Amir-Mohammad Rahmani, Nikil Dutt, Fernando Gehm Moraes
J. Syst. Archit.4
2019 Neural correlates of sparse coding and dimensionality reduction
abstract
Supported by recent computational studies, there is increasing evidence that a wide range of neuronal responses can be understood as an emergent property of nonnegative sparse coding (NSC), an efficient population coding scheme based on dimensionality reduction and sparsity constraints.We review evidence that NSC might be employed by sensory areas to efficiently encode external stimulus spaces, by some associative areas to conjunctively represent multiple behaviorally relevant variables, and possibly by the basal ganglia to coordinate movement.In addition, NSC might provide a useful theoretical framework under which to understand the often complex and nonintuitive response properties of neurons in other brain areas.Although NSC might not apply to all brain areas (for example, motor or executive function areas) the success of NSC-based models, especially in sensory areas, warrants further investigation for neural correlates in other regions. Author summaryBrains face the fundamental challenge of extracting relevant information from high-dimensional external stimuli in order to form the neural basis that can guide an organism's behavior and its interaction with the world.One potential approach to addressing this challenge is to reduce the number of variables required to represent a particular input space (i.e., dimensionality reduction).We review compelling evidence that a range of neuronal responses can be understood as an emergent property of nonnegative sparse coding (NSC)-a form of efficient population coding due to dimensionality reduction and sparsity constraints.
Michael Beyeler, Emily L. Rounds, Kristofor D. Carlson, Nikil Dutt, Jeffrey L. Krichmar
PLoS Comput. Biol.4
2019 Optimal Application Mapping and Scheduling for Network-on-Chips with Computation in STT-RAM Based Router
abstract
Spin-Torque Transfer Magnetic RAM (STT-RAM), one of the emerging nonvolatile memory (NVM) technologies explored as the replacement for SRAM memory architectures, is particularly promising due to the fast access speed, high integration density, and zero standby power consumption. Recently, hybrid deigns with SRAM and STT-RAM buffers for routers in Network-on-Chip (NoC) systems have been widely implemented to maximize the mutually complementary characteristics of different memory technologies, and leverage the efficiency of intra-router latency and system power consumption. With the realization of Processing-in-Memory enabled by STT-RAM, in this paper, we novelly offload the execution from processors to the STT-RAM based on-chip routers to improve the application performance. On top of the hybrid buffer design in routers, we further present system-level approaches, including an ILP model and polynomial-time heuristic algorithms, to fine-tune the application mapping and scheduling on NoCs, with the objectives of improving system performance-energy efficiency. Network overhead caused by flit conflict in conventional communication circumstances can be ideally avoided by computing the contended flits in intermediate routers; meanwhile, the pressure of heavy workload on processors can be relieved by transferring partial operations to routers, such that network latency and system power consumption can be significantly reduced. Experimental results demonstrate that application schedule length and system energy consumption can be reduced by 35.62, 32.87 percent on average, respectively, in extensive evaluation experiments on PARSEC benchmark applications. In particular, the achievements of application performance and energy efficiency, averagely 36.44 and 33.19 percent, for the CNN application AlexNet have verified the practicability and effectiveness of our presented approaches.
Lei Yang 0018, Weichen Liu 0001, Nan Guan, Nikil Dutt
IEEE Trans. Computers4
2019 HESSLE-FREE: <u>He</u>terogeneou<u>s</u> <u>S</u>ystems <u>Le</u>veraging <u>F</u>uzzy Control for <u>R</u>untim<u>e</u> Resourc<u>e</u> Management
abstract
As computing platforms increasingly embrace heterogeneity, runtime resource managers need to efficiently, dynamically, and robustly manage shared resources (e.g., cores, power budgets, memory bandwidth). To address the complexities in heterogeneous systems, state-of-the-art techniques that use heuristics or machine learning have been proposed. On the other hand, conventional control theory can be used for formal guarantees, but may face unmanageable complexity for modeling system dynamics of complex heterogeneous systems. We address this challenge through HESSLE-FREE (Heterogeneous Systems Leveraging Fuzzy Control for Runtime Resource Management): an approach leveraging fuzzy control theory that combines the strengths of classical control theory together with heuristics to form a light-weight, agile, and efficient runtime resource manager for heterogeneous systems. We demonstrate the efficacy of HESSLE-FREE executing on a NVIDIA Jetson TX2 platform (containing a heterogeneous multi-processor with a GPU) to show that HESSLE-FREE: 1) provides opportunity for optimization in the controller and stability analysis to enhance the confidence in the reliability of the system; 2) coordinates heterogeneous compute units to achieve desired objectives (e.g., QoS, optimal power references, FPS) efficiently and with lower complexity , and 3) eases the burden of system specification.
Kasra Moazzemi, Biswadip Maity, Saehanseul Yi, Amir-Mohammad Rahmani, Nikil Dutt
ACM Trans. Embed. Comput. Syst.5
2018 SPECTR: Formal Supervisory Control and Coordination for Many-core Systems Resource Management
abstract
Resource management strategies for many-core systems need to enable sharing of resources such as power, processing cores, and memory bandwidth while coordinating the priority and significance of system- and application-level objectives at runtime in a scalable and robust manner. State-of-the-art approaches use heuristics or machine learning for resource management, but unfortunately lack formalism in providing robustness against unexpected corner cases. While recent efforts deploy classical control-theoretic approaches with some guarantees and formalism, they lack scalability and autonomy to meet changing runtime goals. We present SPECTR, a new resource management approach for many-core systems that leverages formal supervisory control theory (SCT) to combine the strengths of classical control theory with state-of-the-art heuristic approaches to efficiently meet changing runtime goals. SPECTR is a scalable and robust control architecture and a systematic design flow for hierarchical control of many-core systems. SPECTR leverages SCT techniques such as gain scheduling to allow autonomy for individual controllers. It facilitates automatic synthesis of the high-level supervisory controller and its property verification. We implement SPECTR on an Exynos platform containing ARM»s big.LITTLE-based heterogeneous multi-processor (HMP) and demonstrate that SPECTR»s use of SCT is key to managing multiple interacting resources (e.g., chip power and processing cores) in the presence of competing objectives (e.g., satisfying QoS vs. power capping). The principles of SPECTR are easily applicable to any resource type and objective as long as the management problem can be modeled using dynamical systems theory (e.g., difference equations), discrete-event dynamic systems, or fuzzy dynamics.
Amir-Mohammad Rahmani, Bryan Donyanavard, Tiago Rogério Mück, Kasra Moazzemi, Axel Jantsch, Onur Mutlu, Nikil Dutt
ASPLOS7
2018 Approximation-aware coordinated power/performance management for heterogeneous multi-cores
abstract
Run-time resource management of heterogeneous multi-core systems is challenging due to i) dynamic workloads, that often result in ii) conflicting knob actuation decisions, which potentially iii) compromise on performance for thermal safety. We present a runtime resource management strategy for performance guarantees under power constraints using functionally approximate kernels that exploit accuracy-performance trade-offs within error resilient applications. Our controller integrates approximation with power knobs - DVFS, CPU quota, task migration - in coordinated manner to make performance-aware decisions on power management under variable workloads. Experimental results on Odroid XU3 show the effectiveness of this strategy in meeting performance requirements without power violations compared to existing solutions.
Anil Kanduri, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Cristiana Bolchini, Nikil Dutt
DAC6
2018 Gain scheduled control for nonlinear power management in CMPs
abstract
Dynamic voltage and frequency scaling (DVFS) is a well-established technique for power management of thermal-or energy-sensitive chip multiprocessors (CMPs). In this context, linear control theoretic solutions have been successfully implemented to control the voltage-frequency knobs. However, modern CMPs with a large range of operating frequencies and multiple voltage levels display nonlinear behavior in the relationship between frequency and power. State-of-the-art linear controllers therefore under-optimize DVFS operation. We propose a Gain Scheduled Controller (GSC) for nonlinear runtime power management of CMPs that simplifies the controller implementation of systems with varying dynamic properties by utilizing an adaptive control theoretic approach in conjunction with static linear controllers. Our design improves the accuracy of the controller over a static linear controller with minimal overhead. We implement our approach on an Exynos platform containing ARM's big.LITTLE-based heterogeneous multi-processor (HMP) and demonstrate that the system's response to changes in target power is improved by 2x while operating up to 12% more efficiently for tracking accuracy.
Bryan Donyanavard, Amir-Mohammad Rahmani, Tiago Rogério Mück, Kasra Moazemmi, Nikil Dutt
DATE5
2018 Design methodologies for enabling self-awareness in autonomous systems
abstract
This paper deals with challenges and possible solutions for incorporating self-awareness principles in EDA design flows for autonomous systems. We present a holistic approach that enables self-awareness across the software/hardware stack, from systems-on-chip to systems-of-systems (autonomous car) contexts. We use the Information Processing Factory (IPF) metaphor as an exemplar to show how self-awareness can be achieved across multiple abstraction levels, and discuss new research challenges. The IPF approach represents a paradigm shift in platform design by envisioning the move towards a consequent platform-centric design in which the combination of self-organizing learning and formal reactive methods guarantee the applicability of such cyber-physical systems in safety-critical and high-availability applications.
Armin Sadighi, Bryan Donyanavard, Thawra Kadeed, Kasra Moazzemi, Tiago Rogério Mück, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Thomas Wild, Nikil Dutt, Rolf Ernst, Andreas Herkersdorf, Fadi J. Kurdahi
DATE9
2018 Trends in On-chip Dynamic Resource Management
abstract
The Complexity of emerging multi/many-core architectures and diversity of modern workloads demands coordinated dynamic resource management methods. We introduce a classification for these methods capturing the utilized resources and metrics. In this work, we use this classification to survey the key efforts in dynamic resource management. We first cover heuristic and optimization methods used to manage resources such as power, energy, temperature, Quality-of-Service (QoS) and reliability of the system. We then identify some of the machine learning based methods used in tuning architectural parameters in computer systems. In many cases, resource managers need to enforce design constraints during runtime with a certain level of guarantee. Hence, we also study the trend in deploying formal control theoretic approaches in order to achieve efficient and robust dynamic resource management.
Kasra Moazzemi, Anil Kanduri, David Juhasz, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt
DSD8
2018 Edge-Assisted Sensor Control in Healthcare IoT
abstract
The Internet of Things is a key enabler of mobile health-care applications. However, the inherent constraints of mobile devices, such as limited availability of energy, can impair their ability to produce accurate data and, in turn, degrade the output of algorithms processing them in real-time to evaluate the patient's state. This paper presents an edge-assisted framework, where models and control generated by an edge server inform the sensing parameters of mobile sensors. The objective is to maximize the probability that anomalies in the collected signals are detected over extensive periods of time under battery-imposed constraints. Although the proposed concept is general, the control framework is made specific to a use-case where vital signs -heart rate, respiration rate and oxygen saturation- are extracted from a Photoplethysmogram (PPG) signal to detect anomalies in real-time. Experimental results show a 16.9% reduction in sensing energy consumption in comparison to a constant energy consumption with the maximum misdetection probability of 0.17 in a 24-hour health monitoring system.
Delaram Amiri, Arman Anzanpour, Iman Azimi, Marco Levorato, Amir-Mohammad Rahmani, Pasi Liljeberg, Nikil Dutt
GLOBECOM7
2018 Self-Awareness for Heterogeneous MPSoCs: A Case Study using Adaptive, Reflective Middleware
abstract
Self-awareness has a long history in biology, psychology, medicine, engineering and (more recently) computing. In the past decade this has inspired new self-aware strategies for emerging computing substrates (e.g., complex heterogeneous MPSoCs) that must cope with the (often conflicting) challenges of resiliency, energy, heat, cost, performance, security, etc. in the face of highly dynamic operational behaviors and environmental conditions. Earlier we had championed the concept of CyberPhysical-Systems-on-Chip (CPSoC), a new class of sensor-actuator rich many-core computing platforms that intrinsically couples on-chip and cross-layer sensing and actuation to enable self-awareness. Unlike traditional MPSoCs, CPSoC is distinguished by an intelligent co-design of the control, communication, and computing (C3) system that interacts with the physical environment in real-time in order to modify the system's behavior so as to adaptively achieve desired objectives and Quality-of-Service (QoS). The CPSoC design paradigm enables self-awareness (i.e., the ability of the system to observe its own internal and external behaviors such that it is capable of making judicious decision) and (opportunistic) adaptation using the concept of cross-layer physical and virtual sensing and actuations applied across different layers of the hardware/software system stack. The closed loop control used for adaptation to dynamic variation -- commonly known as the observe-decide-act (ODA) loop -- is implemented using an adaptive, reflective middleware layer.
Nikil Dutt
ACM Great Lakes Symposium on VLSI1
2018 CARLsim 4: An Open Source Library for Large Scale, Biologically Detailed Spiking Neural Network Simulation using Heterogeneous Clusters
abstract
Large-scale spiking neural network (SNN) simulations are challenging to implement, due to the memory and computation required to iteratively process the large set of neural state dynamics and updates. To meet these challenges, we have developed CARLsim 4, a user-friendly SNN library written in C++ that can simulate large biologically detailed neural networks. Improving on the efficiency and scalability of earlier releases, the present release allows for the simulation using multiple GPUs and multiple CPU cores concurrently in a heterogeneous computing cluster. Benchmarking results demonstrate simulation of 8.6 million neurons and 0.48 billion synapses using 4 GPUs and up to 60x speedup for multi-GPU implementations over a single-threaded CPU implementation, making CARLsim 4 well-suited for large-scale SNN models in the presence of real-time constraints. Additionally, the present release adds new features, such as leaky-integrate-and-fire (LIF), 9-parameter Izhikevich, multi-compartment neuron models, and fourth order Runge Kutta integration.
Ting-Shuo Chou, Hirak J. Kashyap, Jinwei Xing, Stanislav Listopad, Emily L. Rounds, Michael Beyeler, Nikil Dutt, Jeffrey L. Krichmar
IJCNN7
2018 A Recurrent Neural Network Based Model of Predictive Smooth Pursuit Eye Movement in Primates
abstract
A predictive mechanism in the brain enables primates to visually track a target with almost zero lag smooth pursuit eye movements, overcoming the delays in processing retinal inputs. Interestingly, it also allows pursuit of occluded targets with nonlinear motion patterns. We propose a recurrent neural network (RNN) model that rapidly learns the target velocity sequence and generates eye velocity signals to eliminate the initial lag between target and eye velocities, and to track occluded targets with nonlinear velocity. Moreover, the model is able to adapt to unpredictable perturbation and phase shift of target velocity and qualitatively reproduce the initial pursuit acceleration in experimentally observed timescales. We propose that the frontal eye field (FEF) region of the primate brain is homologous to the proposed RNN based on its persistent predictive activities during pursuit and location on the pursuit pathway.
Hirak J. Kashyap, Georgios Detorakis, Nikil Dutt, Jeffrey L. Krichmar, Emre Neftci
IJCNN3
2018 Unsupervised heart-rate estimation in wearables with Liquid states and a probabilistic readout
Anup Das 0001, Paruthi Pradhapan, Willemijn Groenendaal, Prathyusha Adiraju, Raj Thilak Rajan, Francky Catthoor, Siebren Schaafsma, Jeffrey L. Krichmar, Nikil Dutt, Chris Van Hoof
Neural Networks9
2018 Platform-Centric Self-Awareness as a Key Enabler for Controlling Changes in CPS
abstract
Future cyber-physical systems will host a large number of coexisting distributed applications on hardware platforms with thousands to millions of networked components communicating over open networks. These applications and networks are subject to continuous change. The current separation of design process and operation in the field will be superseded by a life-long design process of adaptation, infield integration, and update. Continuous change and evolution, application interference, environment dynamics and uncertainty lead to complex effects which must be controlled to serve a growing set of platform and application needs. Self-adaptation based on self-awareness and self-configuration has been proposed as a basis for such a continuous in-field process. Research is needed to develop automated in-field design methods and tools with the required safety, availability, and security guarantees. The paper shows two complementary use cases of self-awareness in architectures, methods, and tools for cyber-physical systems. The first use case focuses on safety and availability guarantees in self-aware vehicle platforms. It combines contracting mechanisms, tool based self-analysis and self-configuration. A software architecture and a runtime environment executing these tools and mechanisms autonomously are presented including aspects of self-protection against failures and security threats. The second use case addresses variability and long term evolution in networked MPSoC integrating hardware and software mechanisms of surveillance, monitoring, and continuous adaptation. The approach resembles the logistics and operation principles of manufacturing plants which gave rise to the metaphoric term of an Information Processing Factory that relies on incremental changes and feedback control. Both use cases are investigated by larger research groups. Despite their different approaches, both use cases face similar design and design automation challenges which will be summarized in the end. We will argue that seemingly unrelated research challenges, such as in machine learning and security, could also profit from the methods and superior modeling capabilities of self-aware systems.
Mischa Möstl, Johannes Schlatow, Rolf Ernst, Nikil Dutt, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Fadi J. Kurdahi, Thomas Wild, Armin Sadighi, Andreas Herkersdorf
Proc. IEEE4
2018 Thermal-Aware Task Mapping on Dynamically Reconfigurable Network-on-Chip Based Multiprocessor System-on-Chip
abstract
Dark silicon is the phenomenon that a fraction of many-core chip has to be turned off or run in a low-power state in order to maintain the safe chip temperature. System-level thermal management techniques normally map application on non-adjacent cores, while communication efficiency among these cores will be oppositely affected over conventional network-on-chip (NoC). Recently, SMART NoC architecture is proposed, enabling single-cycle multi-hop bypass channels to be built between distant cores at runtime, to reduce communication latency. However, communication efficiency of SMART NoC will be diminished by communication contention, which will in turn decrease system performance. In this paper, we first propose an Integer-Linear Programming (ILP) model to properly address communication problem, which generates the optimal solutions with the consideration of inter-processor communication. We further present a novel heuristic algorithm for task mapping in dark silicon many-core systems, called TopoMap, on top of SMART architecture, which can effectively solve communication contention problem in polynomial time. With fine-grained consideration of chip thermal reliability and inter-processor communication, presented approaches are able to control the reconfigurability of NoC communication topology in task mapping and scheduling. Thermal-safe system is guaranteed by physically decentralized active cores, and communication overhead is reduced by the minimized communication contention and maximized bypass routing. Performance evaluation on PARSEC shows the applicability and effectiveness of the proposed techniques, which achieve on average 42.5 and 32.4 percent improvement in communication and application performance, and 32.3 percent reduction in system energy consumption, compared with state-of-the-art techniques. TopoMap only introduces 1.8 percent performance difference compared to ILP model and is more scalable to large-size NoCs.
Weichen Liu 0001, Lei Yang 0018, Weiwen Jiang, Liang Feng 0001, Nan Guan, Wei Zhang 0012, Nikil Dutt
IEEE Trans. Computers7
2018 Synergistic CPU-GPU Frequency Capping for Energy-Efficient Mobile Games
abstract
Mobile platforms are increasingly using Heterogeneous Multiprocessor Systems-on-Chip (HMPSoCs) with differentiated processing cores and GPUs to achieve high performance for graphics-intensive applications such as mobile games. Traditionally, separate CPU and GPU governors are deployed in order to achieve energy efficiency through Dynamic Voltage Frequency Scaling (DVFS) but miss opportunities for further energy savings through coordinated system-level application of DVFS. We present a cooperative CPU-GPU DVFS strategy (called Co-Cap) that orchestrates energy-efficient CPU and GPU DVFS through synergistic CPU and GPU frequency capping to avoid frequency overprovisioning while maintaining desired performance. Unlike traditional approaches that target a narrow set of mobile games, our Co-Cap approach is applicable across a wide range of microbenchmarks and mobile games. Our methodology employs a systematic training phase using fine-grained refinement steps with evaluations of frequency capping tables followed by a deployment phase, allowing deployment across a wide range of microbenchmarks and mobile games with varying graphics workloads. Our experimental results across multiple sets of over 200 microbenchmarks and 40 mobile games show that Co-Cap improves energy per frame by on average 8.9% (up to 18.3%) and 7.8% (up to 27.6%) (16.6% and 15.7% in CPU-dominant applications) and achieves minimal frames-per-second (FPS) loss by 0.9% and 0.85% (1.3% and 1.5% in CPU-dominant applications) on average in training and deployment sets, respectively, compared to the default CPU and GPU governors, with negligible overhead in execution time and power consumption on the ODROID-XU3 platform.
Jurn-Gyu Park, Chen-Ying Hsieh, Nikil Dutt, Sung-Soo Lim
ACM Trans. Embed. Comput. Syst.3
2018 ShaVe-ICE: Sharing Distributed Virtualized SPMs in Many-Core Embedded Systems
abstract
Traditional approaches for managing software-programmable memories (SPMs) do not support sharing of distributed on-chip memory resources and, consequently, miss the opportunity to better utilize those memory resources. Managing on-chip memory resources in many-core embedded systems with distributed SPMs requires runtime support to share memory resources between various threads with different memory demands running concurrently. Runtime SPM managers cannot rely on prior knowledge about the dynamically changing mix of threads that will execute and therefore should be designed in a way that enables SPM allocations for any unpredictable mix of threads contending for on-chip memory space. This article proposes ShaVe-ICE , an operating-system-level solution, along with hardware support, to virtualize and ultimately share SPM resources across a many-core embedded system to reduce the average memory latency. We present a number of simple allocation policies to improve performance and energy. Experimental results show that sharing SPMs could reduce the average execution time of the workload up to 19.5% and reduce the dynamic energy consumed in the memory subsystem up to 14%.
Majid Namaki-Shoushtari, Bryan Donyanavard, Luis Angel D. Bathen, Nikil Dutt
ACM Trans. Embed. Comput. Syst.4
2017 Quality-configurable memory hierarchy through approximation: special session
abstract
The memory subsystem is a major contributor to the overall performance and energy consumption of embedded computing platforms. The emergence of "killer" applications such as data-intensive recognition, mining, and synthesis (RMS) applications puts even more stress on the memory subsystem and exacerbates its energy consumption. Traditional mechanisms to ensure data integrity deploy overdesign (e.g., redundancy and error detection/correction) and/or guard-banding that consumes a significant part of the energy consumed in the memory subsystem. We explore opportunities for energy efficiency by exploiting the intrinsic tolerance of a vast class of approximate computing applications to some level of error in the on-chip memory hierarchy. We present two exemplars outlining the typical software and hardware mechanisms that are required for different components in the memory hierarchy, implemented in varying technologies such as SRAM and STT-MRAM.
Majid Namaki-Shoushtari, Amir-Mohammad Rahmani, Nikil Dutt
CASES3
2017 Self-awareness in remote health monitoring systems using wearable electronics
abstract
In healthcare, effective monitoring of patients plays a key role in detecting health deterioration early enough. Many signs of deterioration exist as early as 24 hours prior having a serious impact on the health of a person. As hospitalization times have to be minimized, in-home or remote early warning systems can fill the gap by allowing in-home care while having the potentially problematic conditions and their signs under surveillance and control. This work presents a remote monitoring and diagnostic system that provides a holistic perspective of patients and their health conditions. We discuss how the concept of self-awareness can be used in various parts of the system such as information collection through wearable sensors, confidence assessment of the sensory data, the knowledge base of the patient's health situation, and automation of reasoning about the health situation. Our approach to self-awareness provides (i) situation awareness to consider the impact of variations such as sleeping, walking, running, and resting, (ii) system personalization by reflecting parameters such as age, body mass index, and gender, and (iii) the attention property of self-awareness to improve the energy efficiency and dependability of the system via adjusting the priorities of the sensory data collection. We evaluate the proposed method using a full system demonstration.
Arman Anzanpour, Iman Azimi, Maximilian Götzinger, Amir-Mohammad Rahmani, Nima Taherinejad, Pasi Liljeberg, Axel Jantsch, Nikil Dutt
DATE8
2017 QuARK: Quality-configurable approximate STT-MRAM cache by fine-grained tuning of reliability-energy knobs
abstract
Emerging STT-MRAM memories are promising alternatives for SRAM memories to tackle their low density and high static power consumption, but impose high energy consumption for reliable read/write operations. However, absolute data integrity is not required for many approximate computing applications, allowing energy savings with minimal quality loss. This paper proposes QuARK, a hardware/software approach for trading reliability of STT-MRAM caches for energy savings in the on-chip memory hierarchy of multi- and many-core systems running approximate applications. In contrast to SRAM-based cache-way-level actuators, QuARK utilizes fine-grained cache-line-level actuation knobs with different levels of reliability for individual read and write accesses which are unique to STT-MRAM and suitable for systems running multiple applications with mixed accuracy sensitivity, thus avoiding interapplication actuation interference. Our experimental results with a set of recognition, mining and synthesis (RMS) benchmarks demonstrate up to 40% energy savings over a fully-protected STT-MRAM cache, with negligible loss in the quality of the generated outputs.
Amir Mahdi Hosseini Monazzah, Majid Namaki-Shoushtari, Seyed Ghassem Miremadi, Amir-Mohammad Rahmani, Nikil Dutt
ISLPED5
2017 PoIiCym: rapid prototyping of resource management policies for HMPs
abstract
Heterogeneous Multiprocessors (HMPs) are becoming pervasive in current modern embedded platforms (e.g. mobile devices). These platforms often provide better power-performance tradeoffs than their homogeneous predecessors; however, novel and intelligent resource management policies are required to manage the added complexity of heterogeneous platforms and exploit their power-performance benefits. In this paper we propose PoliCym, a framework for the prototyping, validating, and deploying resource management policies for heterogeneous platforms. PoliCym provides two main benefits to resource management policy developers and to the research community: 1) a trace-based offline simulator allows policies to be quickly prototyped, debugged, and validated on top of arbitrary platform configurations; and 2) a light-weight sensing-actuation interface allows the same policies to be efficiently deployed on top of Linux-based systems without the need for implementation changes or additional development cycles. We evaluate our light-weight interface in terms of overhead and validate the PoliCym offline simulator for an ARM big.LITTLE based HMP platform running Linux.
Tiago Rogério Mück, Bryan Donyanavard, Nikil Dutt
RSP3
2017 HiCH: Hierarchical Fog-Assisted Computing Architecture for Healthcare IoT
abstract
The Internet of Things (IoT) paradigm holds significant promises for remote health monitoring systems. Due to their life- or mission-critical nature, these systems need to provide a high level of availability and accuracy. On the one hand, centralized cloud-based IoT systems lack reliability, punctuality and availability (e.g., in case of slow or unreliable Internet connection), and on the other hand, fully outsourcing data analytics to the edge of the network can result in diminished level of accuracy and adaptability due to the limited computational capacity in edge nodes. In this paper, we tackle these issues by proposing a hierarchical computing architecture, HiCH, for IoT-based health monitoring systems. The core components of the proposed system are 1) a novel computing architecture suitable for hierarchical partitioning and execution of machine learning based data analytics, 2) a closed-loop management technique capable of autonomous system adjustments with respect to patient’s condition. HiCH benefits from the features offered by both fog and cloud computing and introduces a tailored management methodology for healthcare IoT systems. We demonstrate the efficacy of HiCH via a comprehensive performance assessment and evaluation on a continuous remote health monitoring case study focusing on arrhythmia detection for patients suffering from CardioVascular Diseases (CVDs).
Iman Azimi, Arman Anzanpour, Amir-Mohammad Rahmani, Tapio Pahikkala, Marco Levorato, Pasi Liljeberg, Nikil Dutt
ACM Trans. Embed. Comput. Syst.7
2017 Accuracy-Aware Power Management for Many-Core Systems Running Error-Resilient Applications
abstract
Power capping techniques based on dynamic voltage and frequency scaling (DVFS) and power gating (PG) are oriented toward power actuation, compromising on performance and energy. Inherent error resilience of emerging application domains, such as Internet-of-Things (IoT) and machine learning, provides opportunities for energy and performance gains. Leveraging accuracy-performance tradeoffs in such applications, we propose approximation (APPX) as another knob for closelooped power management, to complement power knobs with performance and energy gains. We design a power management framework, APPEND+, that can switch between accurate and approximate modes of execution subject to system throughput requirements. APPEND+ considers the sensitivity of the application to error to make disciplined alteration between levels of APPX such that performance is maximized while error is minimized. We implement a power management scheme that uses APPX, DVFS, and PG knobs hierarchically. We evaluated our proposed approach over machine learning and signal processing applications along with two case studies on IoT-early warning score system and fall detection. APPEND+ yields 1.9× higher throughput, improved latency up to five times, better performance per energy, and dark silicon mitigation compared with the state-of-the-art power management techniques over a set of applications ranging from high to no error resilience.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.7
2016 Cross-layer virtual/physical sensing and actuation for resilient heterogeneous many-core SoCs
abstract
We introduce the concepts of cross-layer virtual/physical sensing and actuation to achieve resiliency for the emerging class of heterogeneous many-core Systems-on-Chip (SoCs). Using the CyberPhysical System-on-Chip (CPSoC) concept as an exemplar sensor-rich many-core heterogeneous computing platform, we illustrate how to intrinsically couple on-chip and cross-layer physical and virtual sensing and actuation applied across different layers of the hardware/software system stack to adaptively achieve desired objectives and Quality-of-Service (QoS). We present two sample use cases that exemplify the cross-layer virtual/physical sensing and actuation approach. First, we present SmartBalance, a cross-layer sensing-driven Linux load balancer for energy efficient task execution on hetergoenous MPSOCs. Second, we present Partially Forgetful Memories, a software/hardware approach that achieves dynamic memory guard-banding for memory resilience and its application for approximate computing.
Santanu Sarma, Tiago Rogério Mück, Majid Namaki-Shoushtari, Abbas BanaiyanMofrad, Nikil Dutt
ASP-DAC5
2016 Approximation knob: power capping meets energy efficiency
abstract
Power Capping techniques are used to restrict power consumption of computer systems to a thermally safe limit. Current many-core systems employ dynamic voltage and frequency scaling (DVFS), power gating (PG) and scheduling methods as actuators for power capping. These knobs arc oriented towards power actuation, while the need for performance and energy savings are increasing in the dark silicon era. To address this, we propose approximation (APPX) as another knob for close-looped power management, lending performance and energy efficiency to existing power capping techniques. We use approximation in a pro-active way for long-term performance-energy objectives, complementing the short-term reactive power objectives. We implement an approximation-enabled power management framework, APPEND, that dynamically chooses an application with appropriate level of approximation from a set of variable accuracy implementations. Subject to the system dynamics, our power manager chooses an effective combination of knobs - APPX, DVFS and PG, in a hierarchical way to ensure power capping with performance and energy gains. Our proposed approach yields 1.5× higher throughput, improved latency upto 5×, better performance per energy and dark silicon mitigation compared to state-of-the-art power management techniques over a set of applications ranging from high to no error resilience.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt, Hannu Tenhunen
ICCAD6
2016 HiCAP: Hierarchical FSM-based Dynamic Integrated CPU-GPU Frequency Capping Governor for Energy-Efficient Mobile Gaming
abstract
Contemporary mobile platforms use software governors to achieve high performance with energy-efficiency for heterogeneous CPU-GPU based architectures that execute mobile games and other graphics-intensive applications. Mobile games typically exhibit inherent behavioral dynamism, which existing governor policies are unable to exploit effectively to manage CPU/GPU DVFS policies. To overcome this problem, we present HiCAP: a Hierarchical Finite State Machine (HFSM) based CPU-GPU governor that models the dynamic behavior of mobile gaming workloads, and applies a cooperative, dynamic CPU-GPU frequency-capping policy to yield energy efficiency adapting to the game's inherent dynamism. Our experiments on a large set of 37 mobile games exhibiting dynamic behavior show that our CAP dynamic governor policy achieved substantial energy efficiency gains of up to 18% improvement in energy-per-frame over existing governor policies, with minimal degradation in quality.
Jurn-Gyu Park, Nikil Dutt, Hoyeonjiki Kim, Sung-Soo Lim
ISLPED2
2016 HAMEX: heterogeneous architecture and memory exploration framework
abstract
The increasing amount of computation in heterogeneous architectures (including CPU and GPU cores) puts a big burden on memory subsystem. With the gap between compute units and the memory performance getting wider, designing a platform with a responsive memory system becomes more challenging. This issue is exacerbated when memory systems have to satisfy a high volume of traffic generated from heterogeneous compute units. Furthermore, as emerging memory technologies are being introduced to address these issues, a rapid and flexible mechanism is needed to evaluate these technologies in the context of heterogeneous architectures. This paper proposes HAMEX, a framework that enables early design space exploration of heterogeneous systems with a focus on resolving memory access bottle-necks. This framework first allows system designers to easily model heterogeneous architectures that can run both CPU and GPU workloads. Next, given a set of workloads partitioned on various compute units, traffic generated by these units are captured in order to explore different memory systems. We show the feasibility of design space exploration using HAMEX by simulating a contemporary commercial heterogeneous platform and explore the opportunities for power and performance improvements by adopting different memory technologies.
Kasra Moazzemi, Chen-Ying Hsieh, Nikil Dutt
RSP3
2016 Toward Smart Embedded Systems: A Self-aware System-on-Chip (SoC) Perspective
abstract
Embedded systems must address a multitude of potentially conflicting design constraints such as resiliency, energy, heat, cost, performance, security, etc., all in the face of highly dynamic operational behaviors and environmental conditions. By incorporating elements of intelligence, the hope is that the resulting “smart” embedded systems will function correctly and within desired constraints in spite of highly dynamic changes in the applications and the environment, as well as in the underlying software/hardware platforms. Since terms related to “smartness” (e.g., self-awareness, self-adaptivity, and autonomy) have been used loosely in many software and hardware computing contexts, we first present a taxonomy of “self-x” terms and use this taxonomy to relate major “smart” software and hardware computing efforts. A major attribute for smart embedded systems is the notion of self-awareness that enables an embedded system to monitor its own state and behavior, as well as the external environment, so as to adapt intelligently. Toward this end, we use a System-on-Chip perspective to show how the CyberPhysical System-on-Chip (CPSoC) exemplar platform achieves self-awareness through a combination of cross-layer sensing, actuation, self-aware adaptations, and online learning. We conclude with some thoughts on open challenges and research directions.
Nikil Dutt, Axel Jantsch, Santanu Sarma
ACM Trans. Embed. Comput. Syst.1
2016 SPMPool: Runtime SPM Management for Memory-Intensive Applications in Embedded Many-Cores
Hossein Tajik, Bryan Donyanavard, Nikil Dutt, Janmartin Jahn, Jörg Henkel
ACM Trans. Embed. Comput. Syst.3
2015 Models, abstractions, and architectures: the missing links in cyber-physical systems
abstract
Bridging disparate realms of physical and cyber system components requires models and methods that enable rapid evaluation of design alternatives in cyber-physical systems (CPS). The diverse intellectual traditions of physical and mathematical sciences makes this task exceptionally hard. This paper seeks to explore potential solutions by examining specific examples of CPS applications in automobiles and smart buildings. Both smart buildings and automobiles are complex systems with embedded knowledge across several domains. We present our experiences with development of CPS applications to illustrate the challenges that arise when expertise across domains is integrated into the system, and show that creation of models, abstractions, and architectures that address these challenges are key to next generation CPS applications.
Bharathan Balaji, Mohammad Abdullah Al Faruque, Nikil Dutt, Rajesh K. Gupta 0001, Yuvraj Agarwal
DAC3
2015 SmartBalance: a sensing-driven linux load balancer for energy efficiency of heterogeneous MPSoCs
abstract
Due to increased demand for higher performance and better energy efficiency, MPSoCs are deploying heterogeneous architectures with architecturally differentiated core types. However, the traditional Linux-based operating system is unable to exploit this heterogeneity since existing kernel load balancing and scheduling approaches lack support for aggressively heterogeneous architectural configurations (e.g. beyond two core types). In this paper we present SmartBalance: a sensing-driven closed-loop load balancer for aggressively heterogeneous MPSoCs that performs load balancing using a sense-predict-balance paradigm. SmartBalance can efficiently manage the chip resources while opportunistically exploiting the workload variations and performance-power trade-offs of different core types. When compared to the standard vanilla Linux kernel load balancer, our per-thread and per-core performance-power-aware scheme shows an improvement in energy efficiency (throughput/Watt) of over 50% for benchmarks from the PARSEC benchmark suite executing on a heterogeneous MPSoC with 4 different core types and over 20% w.r.t. state-of-the-art ARM's global task scheduling (GTS) scheme for octa-core big.Little architecture.
Santanu Sarma, Tiago Rogério Mück, Luis Angel D. Bathen, Nikil Dutt, Alexandru Nicolau
DAC4
2015 Cyberphysical-system-on-chip (CPSoC): a self-aware MPSoC paradigm with cross-layer virtual sensing and actuation
Santanu Sarma, Nikil Dutt, Puneet Gupta 0001, Nalini Venkatasubramanian, Alexandru Nicolau
DATE2
2015 Protecting caches against multi-bit errors using embedded erasure coding
abstract
Technology scaling advancement coupled with operational and environmental effects make embedded memories more vulnerable to both manufacturing and transient errors including multi-bit upsets. Conventional error correcting codes incur high latency, area, and power overheads to correct multi-bit errors. In this paper, we propose Embedded Erasure Coding (EEC), a low-cost technique that can correct multi-bit errors with low overheads. This technique employs interleaved parity bits to provide a fast and low-cost multi-bit error detection. Using the erasure coding concept, the error correction is done by reconstructing the contents of the erroneous cache blocks within each cache set. Our proposed technique trades the performance for higher reliability by reserving a part of the cache (e.g. one way) to store the erasure codes. Our simulation results show that EEC provides high reliability (100% error detection and correction) with lower area overhead as compared to other state-of-the-art techniques while imposing negligible performance overhead (3%).
Abbas BanaiyanMofrad, Mojtaba Ebrahimi, Fabian Oboril, Mehdi Baradaran Tahoori, Nikil Dutt
ETS5
2015 Self-Aware Cyber-Physical Systems-on-Chip
abstract
Self-awareness has a long history in biology, psychology, medicine, and more recently in engineering and computing, where self-aware features are used to enable adaptivity to improve a system's functional value, performance and robustness. With complex many-core Systems-on-Chip (SoCs) facing the conflicting requirements of performance, resiliency, energy, heat, cost, security, etc. - in the face of highly dynamic operational behaviors coupled with process, environment, and workload variabilities - there is an emerging need for self-awareness in these complex SoCs. Unlike traditional MultiProcessor Systems-on-Chip (MPSoCs), self-aware SoCs must deploy an intelligent co-design of the control, communication, and computing infrastructure that interacts with the physical environment in real-time in order to modify the system's behavior so as to adaptively achieve desired objectives and Quality-of-Service (QoS). Self-aware SoCs require a combination of ubiquitous sensing and actuation, health-monitoring, and statistical model-building to enable the SoC's adaptation over time and space. After defining the notion of self-awareness in computing, this paper presents the Cyber-Physical System-on-Chip (CPSoC) concept as an exemplar of a self-aware SoC that intrinsically couples on-chip and cross-layer sensing and actuation using a sensor-actuator rich fabric to enable self-awareness.
Nikil Dutt, Axel Jantsch, Santanu Sarma
ICCAD1
2015 CARLsim 3: A user-friendly and highly optimized library for the creation of neurobiologically detailed spiking neural networks
abstract
Spiking neural network (SNN) models describe key aspects of neural function in a computationally efficient manner and have been used to construct large-scale brain models. Large-scale SNNs are challenging to implement, as they demand high-bandwidth communication, a large amount of memory, and are computationally intensive. Additionally, tuning parameters of these models becomes more difficult and time-consuming with the addition of biologically accurate descriptions. To meet these challenges, we have developed CARLsim 3, a user-friendly, GPU-accelerated SNN library written in C/C++ that is capable of simulating biologically detailed neural models. The present release of CARLsim provides a number of improvements over our prior SNN library to allow the user to easily analyze simulation data, explore synaptic plasticity rules, and automate parameter tuning. In the present paper, we provide examples and performance benchmarks highlighting the library's features.
Michael Beyeler, Kristofor D. Carlson, Ting-Shuo Chou, Nikil Dutt, Jeffrey L. Krichmar
IJCNN4
2015 Large-Scale Spiking Neural Networks using Neuromorphic Hardware Compatible Models
abstract
Neuromorphic engineering is a fast growing field with great potential in both understanding the function of the brain, and constructing practical artifacts that build upon this understanding. For these novel chips and hardware to be useful, hardware compatible applications and simulation tools are needed. We argue that the neural circuit approach, in which networks of neuronal elements model brain circuitry are constructed, allows the development of practical applications and the exploration of brain function. At this level of abstraction, networks of 10 5 neurons or larger can be efficiently simulated, but still preserve the neuronal and synaptic dynamics that appear to be important for brain function. Because the neural circuit level supports spiking neural networks and the prevalent Addressable Event Representation (AER) communication scheme, it fits well with many existing neuromorphic hardware and simulation tools. To show how this approach can be applied, we present case studies of spiking neural networks in vision and recognition tasks based on one instantiation of a simulation environment. However, there are now many hardware options, simulation environments, and applications in this emerging field. These approaches and other considerations are discussed.
Jeffrey L. Krichmar, Philippe Coussy, Nikil Dutt
ACM J. Emerg. Technol. Comput. Syst.3
2015 A GPU-accelerated cortical neural network model for visually guided robot navigation
Michael Beyeler, Nicolas Oros, Nikil Dutt, Jeffrey L. Krichmar
Neural Networks3
2015 DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era
abstract
Fault-Tolerant Voltage-Scalable (FTVS) SRAM cache architectures are a promising approach to improve energy efficiency of memories in the presence of nanoscale process variation. Complex FTVS schemes are commonly proposed to achieve very low minimum supply voltages, but these can suffer from high overheads and thus do not always offer the best power/capacity trade-offs. We observe on our 45nm test chips that the “fault inclusion property” can enable lightweight fault maps that support multiple runtime supply voltages. Based on this observation, we propose a simple and low-overhead FTVS cache architecture for power/capacity scaling. Our mechanism combines multilevel voltage scaling with optional architectural support for power gating of blocks as they become faulty at low voltages. A static (SPCS) policy sets the runtime cache VDD once such that a only a few cache blocks may be faulty in order to minimize the impact on performance. We describe a Static Power/Capacity Scaling (SPCS) policy and two alternate Dynamic Power/Capacity Scaling (DPCS) policies that opportunistically reduce the cache voltage even further for more energy savings. This architecture achieves lower static power for all effective cache capacities than a recent more complex FTVS scheme. This is due to significantly lower overheads, despite the inability of our approach to match the min-VDD of the competing work at a fixed target yield. Over a set of SPEC CPU2006 benchmarks on two system configurations, the average total cache (system) energy saved by SPCS is 62% (22%), while the two DPCS policies achieve roughly similar energy reduction, around 79% (26%). On average, the DPCS approaches incur 2.24% performance and 6% area penalties.
Mark Gottscho, Abbas BanaiyanMofrad, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001
ACM Trans. Archit. Code Optim.3
2015 ViPZonE: Hardware Power Variability-Aware Virtual Memory Management for Energy Savings
abstract
Hardware variability is predicted to increase dramatically over the coming years as a consequence of continued technology scaling. In this paper, we apply the Underdesigned and Opportunistic Computing (UnO) paradigm by exposing system-level power variability to software to improve energy efficiency. We present ViPZonE, a memory management solution in conjunction with application annotations that opportunistically performs memory allocations to reduce DRAM energy. ViPZonE's components consist of a physical address space with DIMM-aware zones, a modified page allocation routine, and a new virtual memory system call for dynamic allocations from userspace. We implemented ViPZonE in the Linux kernel with GLIBC API support, running on a real x86-64 testbed with significant access power variation in its DDR3 DIMMs. We demonstrate that on our testbed, ViPZonE can save up to 27.80 percent memory energy, with no more than 4.80 percent performance degradation across a set of PARSEC benchmarks tested with respect to the baseline Linux software. Furthermore, through a hypothetical “what-if” extension, we predict that in future non-volatile memory systems which consume almost no idle power, ViPZonE could yield even greater benefits, demonstrating the ability to exploit memory hardware variability through opportunistic software.
Mark Gottscho, Luis Angel D. Bathen, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001
IEEE Trans. Computers3
2015 Using a Flexible Fault-Tolerant Cache to Improve Reliability for Ultra Low Voltage Operation
abstract
Caches are known to consume a large part of total microprocessor power. Traditionally, voltage scaling has been used to reduce both dynamic and leakage power in caches. However, aggressive voltage reduction causes process-variation--induced failures in cache SRAM arrays, which compromise cache reliability. In this article, we propose FFT-Cache, a flexible fault-tolerant cache that uses a flexible defect map to configure its architecture to achieve significant reduction in energy consumption through aggressive voltage scaling while maintaining high error reliability. FFT-Cache uses a portion of faulty cache blocks as redundancy—using block-level or line-level replication within or between sets—to tolerate other faulty caches lines and blocks. Our configuration algorithm categorizes the cache lines based on degree of conflict between their blocks to reduce the granularity of redundancy replacement. FFT-Cache thereby sacrifices a minimal number of cache lines to avoid impacting performance while tolerating the maximum amount of defects. Our experimental results on a processor executing SPEC2K benchmarks demonstrate that the operational voltage of both L1/L2 caches can be reduced down to 375 mV, which achieves up to 80% reduction in the dynamic power and up to 48% reduction in the leakage power. This comes with only a small performance loss (<%5) and 13% area overhead.
Abbas BanaiyanMofrad, Houman Homayoun, Nikil Dutt
ACM Trans. Embed. Comput. Syst.3
2014 GPGPU accelerated simulation and parameter tuning for neuromorphic applications
abstract
Neuromorphic engineering takes inspiration from biology to design brain-like systems that are extremely low-power, fault-tolerant, and capable of adaptation to complex environments. The design of these artificial nervous systems involves both the development of neuromorphic hardware devices and the development neuromorphic simulation tools. In this paper, we describe a simulation environment that can be used to design, construct, and run spiking neural networks (SNNs) quickly and efficiently using graphics processing units (GPUs). We then explain how the design of the simulation environment utilizes the parallel processing power of GPUs to simulate large-scale SNNs and describe recent modeling experiments performed using the simulator. Finally, we present an automated parameter tuning framework that utilizes the simulation environment and evolutionary algorithms to tune SNNs. We believe the simulation environment and associated parameter tuning framework presented here can accelerate the development of neuromorphic software and hardware applications by making the design, construction, and tuning of SNNs an easier task.
Kristofor D. Carlson, Michael Beyeler, Nikil Dutt, Jeffrey L. Krichmar
ASP-DAC3
2014 Multi-Layer Memory Resiliency
abstract
With memories continuing to dominate the area, power, cost and performance of a design, there is a critical need to provision reliable, high-performance memory bandwidth for emerging applications. Memories are susceptible to degradation and failures from a wide range of manufacturing, operational and environmental effects, requiring a multi-layer hardware/software approach that can tolerate, adapt and even opportunistically exploit such effects. The overall memory hierarchy is also highly vulnerable to the adverse effects of variability and operational stress. After reviewing the major memory degradation and failure modes, this paper describes the challenges for dependability across the memory hierarchy, and outlines research efforts to achieve multi-layer memory resilience using a hardware/software approach. Two specific exemplars are used to illustrate multilayer memory resilience: first we describe static and dynamic policies to achieve energy savings in caches using aggressive voltage scaling combined with disabling faulty blocks; and second we show how software characteristics can be exposed to the architecture in order to mitigate the aging of large register files in GPGPUs. These approaches can further benefit from semantic retention of application intent to enhance memory dependability across multiple abstraction levels, including applications, compilers, run-time systems, and hardware platforms.
Nikil Dutt, Puneet Gupta 0001, Alexandru Nicolau, Abbas BanaiyanMofrad, Mark Gottscho, Majid Namaki-Shoushtari
DAC1
2014 Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant Caches
abstract
Complicated approaches to fault-tolerant voltage-scalable (FTVS) SRAM cache architectures can suffer from high overheads. We propose static (SPCS) and dynamic (DPCS) variants of power/capacity scaling, a simple and low-overhead fault-tolerant cache architecture that utilizes insights gained from our 45nm SOI test chip. Our mechanism combines multi-level voltage scaling with power gating of blocks that become faulty at each voltage level. The SPCS policy sets the runtime cache VDD statically such that almost all of the cache blocks are not faulty. The DPCS policy opportunistically reduces the voltage further to save more power than SPCS while limiting the impact on performance caused by additional faulty blocks. Through an analytical evaluation, we show that our approach can achieve lower static power for all effective cache capacities than a recent complex FTVS work. This is due to significantly lower overheads, despite the failure of our approach to match the min-VDD of the competing work at fixed yield. Through architectural simulations, we find that the average energy saved by SPCS is 55%, while DPCS saves an average of 69% of energy with respect to baseline caches at 1 V. Our approach incurs no more than 4% performance and 5% area penalties in the worst case cache configuration.
Mark Gottscho, Abbas BanaiyanMofrad, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001
DAC3
2014 Sense-making from Distributed and Mobile Sensing Data: A Middleware Perspective
abstract
This paper presents a scalable and collaborative mobile crowdsensing framework for efficient collective understanding of users, contexts, and their environments. Collaborative mobile crowdsensing enables information to be gathered and shared by users who are directly involved (participatory sensing) or integrated seamlessly as needed (opportunistic sensing) through user mobile platforms. To address the scalability needs of the mobile ecosystem, we additionally employ compressive sensing techniques for approximate gathering and processing of sensor data - this requires new mechanisms for sensor data collection, tunable approximate processing, and mobile networking architecture, to create a compressive collaborative mobile crowdsensing platform called SenseDroid. The proposed framework is build using a multi-tired hierarchical architecture to sense spatial variations of a parameter of interest, perceive spatio-temporal fields, and enable energy efficient local mobile sensing with a small number of measurements. This approximate, yet tunable approach combines different sensing approaches opportunistically while trading scalability (and coverage) for data accuracy (and energy efficiency). In this paper we propose and discuss the framework and the challenges associated with compressive and collaborative mobile sensing for multi-tired hierarchical mobile network architecture for emerging mobile collaborative applications.
Santanu Sarma, Nalini Venkatasubramanian, Nikil Dutt
DAC3
2014 Minimal sparse observability of complex networks: Application to MPSoC sensor placement and run-time thermal estimation & tracking
abstract
This paper addresses the fundamental and practically useful question of identifying a minimum set of sensors and their locations through which a large complex dynamical network system and its time-dependent states can be observed. The paper defines the minimal sparse observability problem (MSOP) and provides analytical tools with necessary and sufficient conditions to make an arbitrary complex dynamic network system completely observable. The mathematical tools are then used to develop effective algorithms to find the sparsest measurement vector that provides the ability to estimate the internal states of a complex dynamic network system from experimentally accessible outputs. The developed algorithms are further used in the design of a sparse Kalman filter (SKF) to estimate the time-dependent internal states of a linear time-invariant (LTI) dynamical network system. The approach is applied to illustrate the minimum sensor in-situ run-time thermal estimation and robust hotspot tracking for dynamic thermal management (DTM) of high performance processors and MPSoCs.
Santanu Sarma, Nikil Dutt
DATE2
2014 FPGA emulation and prototyping of a cyberphysical-system-on-chip (CPSoC)
abstract
Cyber-Physical Systems-on-Chip (CPSoC) are a new class of sensor- and actuator-rich multiprocessor system-on-chips (MPSoCs) whose operations are monitored, coordinated, and controlled using a computing-communication-control (C3) centric core with additional on-chip and cross-layer sensing and actuation capabilities that enable self-awareness within the observe-decide-act (ODA) paradigm. In order to build, evaluate, and illustrate the effectiveness of various features of this new MPSoC paradigm in a fast and cost effective way, a rapid prototyping and emulation platform along with the tool chains is absolutely necessary. In this paper, we present a design library and an FPGA emulation and prototyping platform to build and investigate self-aware adaptive computing using CPSoC paradigm. Our example implementation of CPSoC prototyping using Xilinx FPGAs includes ring-oscillator (RO) based multipurpose sensors integrated with a sensor network-on-chip (sNoC) which in turn is interfaced either to a bus based shared memory architecture or to a communication and computation network-on- chip (cNoC) distributed fabric supporting several actuation mechanism in the software and hardware stack. We also briefly discuss few applications of the CPSoC design library and the platform.
Santanu Sarma, Nikil Dutt
RSP2
2014 A Reliability-Aware Address Mapping Strategy for NAND Flash Memory Storage Systems
abstract
The increasing density of NAND flash memory leads to a dramatic increase in the bit error rate of flash, which greatly reduces the ability of error correcting codes (ECC) to handle multibit errors. NAND flash memory is normally used to store the file system metadata and page mapping information. Thus, a broken physical page containing metadata may cause an unintended and severe change in functionality of the entire flash. This paper presents Meta-Cure, a novel hardware and file system interface that transparently protects metadata in the presence of multibit faults. Meta-Cure exploits built-in ECC and replication in order to protect pages containing critical data, such as file system metadata. Redundant pairs are formed at run time and distributed to different physical pages to protect against failures. Meta-Cure requires no changes to the file system, on-chip hierarchy, or hardware implementation of flash memory chip. We evaluate Meta-Cure under a real-embedded platform using a variety of I/O traces. The evaluation platform adopts dual ARM Cortex A9 processor cores with 64 Gb NAND flash memory. We have evaluated the effectiveness of Meta-Cure on the new technology file system file system. Experimental results show that the proposed technique can reduce uncorrectable page errors by 70.38% with less than 7.86% time overhead in comparison with conventional error correction techniques.
Yi Wang 0003, Min Huang 0002, Zili Shao, Henry C. B. Chan, Luis Angel D. Bathen, Nikil Dutt
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2014 NoC-based fault-tolerant cache design in chip multiprocessors
abstract
Advances in technology scaling increasingly make emerging Chip MultiProcessor (CMP) platforms more susceptible to failures that cause various reliability challenges. In such platforms, error-prone on-chip memories (caches) continue to dominate the chip area. Also, Network-on-Chip (NoC) fabrics are increasingly used to manage the scalability of these architectures. We present a novel solution for efficient implementation of fault-tolerant design of Last-Level Cache (LLC) in CMP architectures. The proposed approach leverages the interconnection network fabric to protect the LLC cache banks against permanent faults in an efficient and scalable way. During an LLC access to a faulty block, the network detects and corrects the faults, returning the fault-free data to the requesting core. Leveraging the NoC interconnection fabric, designers can implement any cache fault-tolerant scheme in an efficient, modular, and scalable manner for emerging multicore/manycore platforms. We propose four different policies for implementing a remapping-based fault-tolerant scheme leveraging the NoC fabric in different settings. The proposed policies enable design trade-offs between NoC traffic (packets sent through the network) and the intrinsic parallelism of these communication mechanisms, allowing designers to tune the system based on design constraints. We perform an extensive design space exploration on NoC benchmarks to demonstrate the usability and efficacy of our approach. In addition, we perform sensitivity analysis to observe the behavior of various policies in reaction to improvements in the NoC architecture. The overheads of leveraging the NoC fabric are minimal: on an 8-core, 16-cache-bank CMP we demonstrate reliable access to LLCs with additional overheads of less than 3% in area and less than 7% in power.
Abbas BanaiyanMofrad, Gustavo Girão, Nikil Dutt
ACM Trans. Embed. Comput. Syst.3
2014 Embedded RAIDs-on-chip for bus-based chip-multiprocessors
abstract
The dual effects of larger die sizes and technology scaling, combined with aggressive voltage scaling for power reduction, increase the error rates for on-chip memories. Traditional on-chip memory reliability techniques (e.g., ECC) incur significant power and performance overheads. In this article, we propose a low-power-and-performance-overhead Embedded RAID (E-RAID) strategy and present Embedded RAIDs-on-Chip (E-RoC), a distributed dynamically managed reliable memory subsystem for bus-based Chip-Multiprocessors. E-RoC achieves reliability through redundancy by optimizing RAID-like policies tuned for on-chip distributed memories. We achieve on-chip reliability of memories through the use of Distributed Dynamic ScratchPad Allocatable Memories (DSPAMs) and their allocation policies. We exploit aggressive voltage scaling to reduce power consumption overheads due to parallel DSPAM accesses, and rely on the E-RoC Manager to automatically handle any resulting voltage-scaling-induced errors. We demonstrate how E-RAIDs can further enhance the fault tolerance of traditional memory reliability approaches by designing E-RAID levels that exploit ECC. Finally, we show the power and flexibility of the E-RoC concept by showing the benefits of having a heterogeneous E-RAID levels that fit each application's needs (fault tolerance, power/energy, performance). Our experimental results on CHStone/Mediabench II benchmarks show that our E-RAID levels converge to 100% error-free data rates much faster than traditional ECC approaches. Moreover, E-RAID levels that exploit ECC can guarantee 99.9% error-free data rates at ultra low Vdd on average, where as traditional ECC approaches were able to attain at most 99.1% error-free data rates. We observe an average of 22% dynamic power consumption increase by using traditional ECC approaches with respect to the baseline (non-voltage scaled SPMs), whereas our E-RAID levels are able to save dynamic power consumption by an average of 27% (w.r.t. the same non-voltage scaled SPMs baseline), while incurring worst-case 2% higher performance overheads than traditional ECC approaches. By voltage scaling the memories, we see that traditional ECC approaches are able to save static energy by 6.4% (average), where as our E-RAID approaches achieve 23.4% static energy savings (average). Finally, we observe that mixing E-RAID levels allows us to further reduce the dynamic power consumption by up to 55.5% at the cost of an average 5.6% increase in execution time over traditional approaches.
Luis Angel D. Bathen, Nikil Dutt
ACM Trans. Embed. Comput. Syst.2
2014 Multicopy Cache: A Highly Energy-Efficient Cache Architecture
abstract
Caches are known to consume a large part of total microprocessor energy. Traditionally, voltage scaling has been used to reduce both dynamic and leakage power in caches. However, aggressive voltage reduction causes process-variation-induced failures in cache SRAM arrays, thus compromising cache reliability. We present MultiCopy Cache (MC 2 ), a new cache architecture that achieves significant reduction in energy consumption through aggressive voltage scaling while maintaining high error resilience (reliability) by exploiting multiple copies of each data item in the cache. Unlike many previous approaches, MC 2 does not require any error map characterization and therefore is responsive to changing operating conditions (e.g., Vdd noise, temperature, and leakage) of the cache. MC 2 also incurs significantly lower overheads compared to other ECC-based caches. Our experimental results on embedded benchmarks demonstrate that MC 2 achieves up to 60% reduction in energy and energy-delay product (EDP) with only 3.5% reduction in IPC and no appreciable area overhead.
Arup Chakraborty, Houman Homayoun, Amin Khajeh, Nikil Dutt, Ahmed M. Eltawil, Fadi J. Kurdahi
ACM Trans. Embed. Comput. Syst.4
2014 Introduction to Special Issue on Cross-layer Dependable Embedded Systems
abstract
No abstract available.
Nikil Dutt, Mehdi Baradaran Tahoori
ACM Trans. Embed. Comput. Syst.1
2014 SPMCloud: Towards the Single-Chip Embedded ScratchPad Memory-Based Storage Cloud
abstract
The era of cloud computing on-a-chip is enabled by the aggressive move towards many-core platforms and the rapid adoption of Network-on-Chips. As a result, there is a need for large-scale distributed on-chip shared memories that are reliable, low power, and seamlessly manageable. In this work, we propose SPMCloud , a novel scratchpad-memory-based cloud-inspired volatile storage subsystem designed to meet the needs of future-generation many-core platforms. SPMCloud is composed of several concepts, including: (1) a highly scalable data-center-like memory subsystem that exploits two enterprise-network-inspired memory configurations, namely, embedded Network Attached Storage ( eNAS ) and embedded Storage Area Network ( eSAN ), and (2) on-demand allocation of reliable memory space through memory virtualization and the use of embedded RAIDs. Our experimental results on Mediabench/CHStone benchmarks show that the SPMCloud 's fully distributed reliable memory subsystems can achieve 48% energy savings and 70% latency reduction on average over state-of-the-art NoC memory reliability techniques. We then evaluate the scalability of the SPMCloud and compare it with traditional SPM allocation policies. The SPMCloud 's dynamic allocator outperforms the best competition by an average 60% ( eNAS ) and 46% ( eSAN ) when the platform runs at 250 MHz and by an average 80% ( eNAS ) and 40% when running at 1 GHz. Moreover, the SPMCloud achieves an average 83% energy savings across all configurations (number of cores) with respect to the best competitors when running at 250 MHz and 1 GHz. We then studied the SPM hit ratio across the various allocation policies discussed in this article and showed that on average the SPMCloud 's priority-driven dynamic allocation policy achieves 93.5% SPM hit ratio, 0.6% higher hit ratio than the closest allocation policy. We then showed that the eNAS and eSAN achieve an average of 67.9% and 29% reduction in execution time, respectively, over the best competitor. Similarly, the eNAS and eSAN achieve an average of 82.7% and 82.3% energy savings, respectively, over the best competitor. Furthermore, we evaluated the scalability of the SPMCloud and its performance/energy efficiency when providing support for some of the heavier E-RAID levels, and showed that the eNAS / eSAN configurations with SECDED achieve an average of 51.5% and 34.9% reduction in execution time, respectively, over the best competitor with SECDED. Similarly, the eNAS / eSAN configurations with E-RAID Level 1, + SECDED achieve an average of 82.3% and 75.6% energy savings, respectively, over the best competitor.
Luis Angel D. Bathen, Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.2
2014 A Reliability Enhanced Address Mapping Strategy for Three-Dimensional (3-D) NAND Flash Memory
abstract
The linear scaling down of NAND flash memory is approaching its physical, electrical, and reliability limitations. To maintain the current trend of increasing bit density and reducing bit per cost, 3-D flash memory is emerging as a viable solution to fulfill the ever-increasing demands of storage capacity. In 3-D NAND flash memory, multiple layers are stacked to provide ultrahigh density storage devices. However, the physical architecture of 3-D flash memory leads to a higher probability of disturbance to adjacent physical pages and greatly increases bit error rates. This paper presents a novel physical-location-aware address mapping strategy for 3-D NAND flash memory. It permutes the physical mapping of pages and maximizes the distance between the consecutively logical pages, which can significantly reduce the disturbance to adjacent physical pages and effectively enhance the reliability. The proposed mapping strategy is applied to a representative flash storage system. Experimental results show that the proposed scheme can reduce uncorrectable page errors by 70.16% with less than 10.01% space overhead in comparison with the baseline scheme.
Yi Wang 0003, Zili Shao, Henry C. B. Chan, Luis Angel D. Bathen, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.5
2013 Variability-aware memory management for nanoscale computing
abstract
As the semiconductor industry continues to push the limits of sub-micron technology, the ITRS expects hardware (e.g., die-to-die, wafer-to-wafer, and chip-to-chip) variations to continue increasing over the next few decades. As a result, it is imperative for designers to build variation-aware software stacks that may adapt and opportunistically exploit said variations to increase system performance/responsiveness as well as minimize power consumption. The memory subsystem is one of the largest components in today's computing system, a main contributor to the overall power consumption of the system, and therefore one of the most vulnerable components to the effects of variations (e.g., power). This paper discusses the concept of variability-aware memory management for nanoscale computing systems. We show how to opportunistically exploit the hardware variations in on-chip and off-chip memory at the system level through the deployment of variation-aware software stacks.
Nikil Dutt, Puneet Gupta 0001, Alexandru Nicolau, Luis Angel D. Bathen, Mark Gottscho
ASP-DAC1
2013 VISA synthesis: Variation-aware Instruction Set Architecture synthesis
abstract
We present VISA: a novel Variation-aware Instruction Set Architecture synthesis approach that makes effective use of process variation from both software and hardware points of view. To achieve an efficient speedup, VISA selects custom instructions based on statistical static timing analysis (SSTA) for aggressive clocking. Furthermore, with minimum performance overhead, VISA dynamically detects and corrects timing faults resulting from aggressive clocking of the underlying processor. This hybrid software/hardware approach generates significant speedup without degrading the yield. Our experimental results on commonly used ISA synthesis benchmarks demonstrate that VISA achieves significant performance improvement compared with a traditional deterministic worst case-based approach (up to 78.0%) and an existing SSTA-based approach (up to 49.4%).
Yuko Hara-Azumi, Takuya Azumi, Nikil Dutt
ASP-DAC3
2013 Reliable on-chip systems in the nano-era: lessons learnt and future trends
abstract
Reliability concerns due to technology scaling have been a major focus of researchers and designers for several technology nodes. Therefore, many new techniques for enhancing and optimizing reliability have emerged particularly within the last five to ten years. This perspective paper introduces the most prominent reliability concerns from today's points of view and roughly recapitulates the progress in the community so far. The focus of this paper is on perspective trends from the industrial as well as academic points of view that suggest a way for coping with reliability challenges in upcoming technology nodes.
Jörg Henkel, Lars Bauer, Nikil Dutt, Puneet Gupta 0001, Sani R. Nassif, Muhammad Shafique 0001, Mehdi Baradaran Tahoori, Norbert Wehn
DAC3
2013 VAWOM: temperature and process variation aware wearout management in 3D multicore architecture
abstract
Three dimensional (3D) integration attempts to address challenges and limitations of new technologies such as interconnect delay and power consumption. However, high power density and increased temperature in 3D architectures accelerate wearout failure mechanisms such as Negative Bias Temperature Instability (NBTI). In this paper we present VAWOM (Variation Aware WearOut Management), an approach that reduces the NBTI effect by exploiting temperature and process variation in 3D architectures. We demonstrate the efficacy of VAWOM on a two-layer 3D architecture with 4x4 cores on the first layer and 4x4 last level caches on the second layer, and show that VAWOM reduces NBTI induced threshold voltage degradation by 30% with only a small degradation in performance.
Hossein Tajik, Houman Homayoun, Nikil Dutt
DAC3
2013 Modeling and analysis of fault-tolerant distributed memories for networks-on-chip
abstract
Advances in technology scaling increasingly make Network-on-Chips (NoCs) more susceptible to failures that cause various reliability challenges. With increasing area occupied by different on-chip memories, strategies for maintaining fault-tolerance of distributed on-chip memories become a major design challenge. We propose a system-level design methodology for scalable fault-tolerance of distributed on-chip memories in NoCs. We introduce a novel reliability clustering model for fault-tolerance analysis and shared redundancy management of on-chip memory blocks. We perform extensive design space exploration applying the proposed reliability clustering on a block-redundancy fault-tolerant scheme to evaluate the tradeoffs between reliability, performance, and overheads. Evaluations on a 64-core chip multiprocessor (CMP) with an 8x8 mesh NoC show that distinct strategies of our case study may yield up to 20% improvements in performance gains and 25% improvement in energy savings across different benchmarks, and uncover interesting design configurations.
Abbas BanaiyanMofrad, Nikil Dutt, Gustavo Girão
DATE2
2013 Outlook for many-core systems: Cloudy with a chance of virtualization
abstract
The emergence of many-core platforms increases the need for high memory bandwidth, which in turn creates the need for vast amounts of on-chip memory space. Designers must carefully provision the on-chip memory resources to meet application needs. Efficient memory management is extremely critical since it has a great impact on the system's power consumption and throughput. While memory hierarchies have traditionally been based on SRAM-based on-chip caches, the demands of predictability, low power/energy, as well as the emergence of non-volatile memories (NVMs) and mixed-criticality systems, have led to increasing use of software-controlled on-chip memories. The talk presents strategies for efficiently managing software-controlled memories in the many-core domain, while addressing the disparate challenges faced by designers in deploying such memory subsystems (e.g., sharing memory resources, handling variability, and deploying heterogeneous memory families). The overall approach revisits and extends the classical notion of clouds and memory virtualization to handle scalable on-chip memory organizations for reduced power consumption, security, reliability and yield management.
Nikil Dutt
ETS1
2013 Biologically plausible models of homeostasis and STDP: Stability and learning in spiking neural networks
abstract
Spiking neural network (SNN) simulations with spike-timing dependent plasticity (STDP) often experience runaway synaptic dynamics and require some sort of regulatory mechanism to stay within a stable operating regime. Previous homeostatic models have used L1or L2normalization to scale the synaptic weights but the biophysical mechanisms underlying these processes remain undiscovered. We propose a model for homeostatic synaptic scaling that modifies synaptic weights in a multiplicative manner based on the average postsynaptic firing rate as observed in experiments. The homeostatic mechanism was implemented with STDP in conductance-based SNNs with Izhikevich-type neurons. In the first set of simulations, homeostatic synaptic scaling stabilized weight changes in STDP and prevented runaway dynamics in simple SNNs. During the second set of simulations, homeostatic synaptic scaling was found to be necessary for the unsupervised learning of V1 simple cell receptive fields in response to patterned inputs. STDP, in combination with homeostatic synaptic scaling, was shown to be mathematically equivalent to non-negative matrix factorization (NNMF) and the stability of the homeostatic update rule was proven. The homeostatic model presented here is novel, biologically plausible, and capable of unsupervised learning of patterned inputs, which has been a significant challenge for SNNs with STDP.
Kristofor D. Carlson, Micah Richert, Nikil Dutt, Jeffrey L. Krichmar
IJCNN3
2013 Categorization and decision-making in a neurobiologically plausible spiking network using a STDP-like learning rule
Michael Beyeler, Nikil Dutt, Jeffrey L. Krichmar
Neural Networks2
2013 Underdesigned and Opportunistic Computing in Presence of Hardware Variability
abstract
Microelectronic circuits exhibit increasing variations in performance, power consumption, and reliability parameters across the manufactured parts and across use of these parts over time in the field. These variations have led to increasing use of overdesign and guardbands in design and test to ensure yield and reliability with respect to a rigid set of datasheet specifications. This paper explores the possibility of constructing computing machines that purposely expose hardware variations to various layers of the system stack including software. This leads to the vision of underdesigned hardware that utilizes a software stack that opportunistically adapts to a sensed or modeled hardware. The envisioned underdesigned and opportunistic computing (UnO) machines face a number of challenges related to the sensing infrastructure and software interfaces that can effectively utilize the sensory data. In this paper, we outline specific sensing mechanisms that we have developed and their potential use in building UnO machines.
Puneet Gupta 0001, Yuvraj Agarwal, Lara Dolecek, Nikil Dutt, Rajesh K. Gupta 0001, Rakesh Kumar 0002, Subhasish Mitra, Alexandru Nicolau, Tajana Rosing, Mani Srivastava 0001, Steven Swanson, Dennis Sylvester
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2013 MultiMaKe: Chip-multiprocessor driven memory-aware kernel pipelining
abstract
The increasing demand for low-power and high-performance multimedia embedded systems has motivated the need for effective solutions to satisfy application bandwidth and latency requirements under a tight power budget. As technology scales, it is imperative that applications are optimized to take full advantage of the underlying resources and meet both power and performance requirements. We propose MultiMaKe, an application mapping design flow capable of discovering and enabling parallelism opportunities via code transformations, efficiently distributing the computational load across resources, and minimizing unnecessary data transfers. Our approach decomposes the application's tasks into smaller units of computations called kernels, which are distributed and pipelined across the different processing resources. We exploit the ideas of inter-kernel data reuse to minimize unnecessary data transfers between kernels, early execution edges to drive performance, and kernel pipelining to increase system throughput. Our experimental results on JPEG and JPEG2000 show up to 97% off-chip memory access reduction, and up to 80% execution time reduction over standard mapping and task-level pipelining approaches.
Luis Angel D. Bathen, Yongjin Ahn, Sudeep Pasricha, Nikil Dutt
ACM Trans. Embed. Comput. Syst.4
2012 HaVOC: a hybrid memory-aware virtualization layer for on-chip distributed ScratchPad and non-volatile memories
abstract
Hybrid on-chip memories that combine Non-Volatile Memories (NVMs) with SRAMs promise to mitigate the increasing leakage power of traditional on-chip SRAMs. We present HaVOC: a run-time memory manager that virtualizes the hybrid on-chip memory space and supports efficient sharing of distributed ScratchPad Memories (SPMs) and NVMs. HaVOC allows programmers and the compiler to partition the application's address space and generate data/instruction block layouts considering virtualized hybrid address spaces. We define a data volatility metric used by our hybrid memory-aware compilation flow to generate memory allocation policies that are enforced at run-time by a filter-inspired dynamic memory algorithm. Our experimental results with a set of embedded benchmarks executing simultaneously on a Chip-Multiprocessor with hybrid NVM/SPMs show that HaVOC is able to reduce execution time and energy by 60.8% and 74.7% respectively with respect to traditional multitasking based SPM allocation policies.
Luis Angel D. Bathen, Nikil Dutt
DAC2
2012 Meta-Cure: a reliability enhancement strategy for metadata in NAND flash memory storage systems
abstract
The increasing density of NAND flash memory leads to a dramatic increase in the bit error rate of flash, which greatly reduces the ability of error correcting codes (ECC) to handle multi-bit errors. To ensure the functionality and reliability of flash memory, the pages containing address mapping information and other metadata should be carefully stored in flash memory. This paper presents Meta-Cure, a novel hardware and file system interface that transparently protects metadata in the presence of multi-bit faults. Meta-Cure exploits built-in ECC and replication in order to protect pages containing critical data. Redundant pairs are formed at run time and distributed to different physical pages to protect against failures. Meta-Cure requires no changes to the file system, on-chip hierarchy, or hardware implementation of flash memory chip. Experimental results show that the proposed technique can reduce uncorrectable page errors by 92% with less than 1% space overhead in comparison with conventional error correction techniques.
Yi Wang 0003, Luis Angel D. Bathen, Nikil Dutt, Zili Shao
DAC3
2012 VaMV: Variability-aware Memory Virtualization
abstract
Power consumption variability of both on-chip SRAMs and off-chip DRAMs is expected to continue to increase over the next decades. We opportunistically exploit this variability through a novel Variability-aware Memory Virtualization (VaMV) layer that allows programmers to partition their application's address space (through annotations) into virtual address regions and create mapping policies for each region. Each policy has different requirements (e.g., power, fault-tolerance) and is exploited by our dynamic memory management module (VaMVisor), which adapts to the underlying hardware, prioritizes the memory resources according to their characteristics (e.g., power consumption), and selectively maps data to the best-fitting memory resource (e.g., high-utilization data to low-power memory space). Our experimental results on embedded benchmarks show that VaMV is capable of reducing dynamic power consumption by 63% on average while reducing total execution time by an average of 34% by exploiting: 1) SRAM voltage scaling, 2) DRAM power variability, and 3) Efficient dynamic policy-driven variability-aware memory allocation.
Luis Angel D. Bathen, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001
DATE2
2012 3D-FlashMap: A physical-location-aware block mapping strategy for 3D NAND flash memory
abstract
Three-dimensional (3D) flash memory is emerging to fulfil the ever-increasing demands of storage capacity. In 3D NAND flash memory, multiple layers are stacked to increase bit density and reduce bit cost of flash memory. However, the physical architecture of 3D flash memory leads to a higher probability of disturbance to adjacent physical pages and greatly increases bit error rates. This paper presents 3D-FlashMap, a novel physical-location-aware block mapping strategy for three-dimensional NAND flash memory. 3D-FlashMap permutes the physical mapping of blocks and maximizes the distance between consecutively logical blocks, which can significantly reduce the disturbance to adjacent physical pages and effectively enhance the reliability. We apply 3D-FlashMap to a representative flash storage system. Experimental results show that the proposed scheme can reduce uncorrectable page errors by 85% with less than 2% space overhead in comparison with the baseline scheme.
Yi Wang 0003, Luis Angel D. Bathen, Zili Shao, Nikil Dutt
DATE4
2012 Spiking neuron model of basal forebrain enhancement of visual attention
abstract
Attentional mechanisms allow the brain to enhance the representation and transmission of certain signals at the expense of others. The basal forebrain has been shown to play an important role in attention through its diverse set of interactions with sensory and associational areas. A recent empirical study indicates that the nucleus basalis, a subset of neurons located in the basal forebrain, is important for improving sensory processing by increasing reliability and decreasing redundancy in the cortex and thalamus [1, 2]. We developed a spiking neural network model that simulates the nucleus basalis' interaction with the thalamus and visual cortex. In this model, we simulated two modes of action by which it is thought that the nucleus basalis may be influencing sensory processing: (1) inhibitory projections from the nucleus basalis to the thalamic reticular nucleus, which disinhibit the lateral geniculate nucleus (LGN) and gate information into the cortex, and (2) cholinergic excitation of inhibitory neurons in the visual cortex. We showed that the inhibition of the thalamic reticular nucleus GABAergic neurons leads to an increase in the reliability of spikes in the LGN and cortex. We observed that a decrease in the burst to tonic firing ratio in the LGN, coupled with the cholinergic system increasing inhibition in the visual cortex caused decorrelation in the cortex. These findings will help us better understand the mechanisms behind the control of attention by the basal forebrain and shed light on how the orchestrated action of the basal forebrain on multiple target areas can improve information processing in the brain.
Michael C. Avery, Jeffrey L. Krichmar, Nikil Dutt
IJCNN3
2012 Keynote speach
abstract
While designers have traditionally dealt with unreliable embedded memories through standard fault-tolerant techniques in the past, aggressive technology scaling and low-voltage operation (to save power) pose significant reliability challenges for distributed embedded memories. I briefly survey traditional hardware and software schemes for addressing reliability to achieve a low-power, fault-tolerant memory space, and keep production yield at tolerable levels. I then present strategies to virtualize the on-chip memory space to cope with the issues of low power, security, reliability and performance. I present the notions of Embedded Raids-on-Chip (E-RoC) and SPMVisor, holistic hardware/software solutions that virtualize the user memory space and which exploit unreliable distributed embedded memories for reduced power consumption, security, reliability and yield management.
Nikil Dutt
RSP1
2012 Software Controlled Memories for Scalable Many-Core Architectures
abstract
Technology scaling along with the ever evolving demand for media-rich software stacks have motivated the need for many-core platforms. With the increase in compute power and its inherent demand for high memory bandwidth comes the need for vast amounts of on-chip memory space. Thus, designers must carefully provision the memory real-estate to meet their application's needs. It has been shown in the embedded systems domain that both software controlled memories (e.g., scratchpad memories) and hardware-controlled memories (e.g., caches) have their pros and cons, some application domains such as multimedia fit very well in the software-controlled memory model, while other domains such as databases work well with caches. As a result, efficient memory management is extremely critical as it has a great impact on the system's power consumption and throughput. Traditional memory hierarchies primarily consist of SRAM-based on-chip caches, however, with the emergence of non-volatile memories (NVMs) and mixed-criticality systems, on-chip memories will be heterogeneous, not only in type (cache vs. scratchpad) but also in technology (e.g., SRAM vs. NVM). This paper surveys the state of the art in memory subsystems for many-core platforms, and presents strategies for efficiently managing software-controlled memories in the many-core domain, while addressing the various challenges designers face in deploying such memory subsystems (e.g., sharing the memory resources, accounting for variations in the subsystem, etc.).
Luis Angel D. Bathen, Nikil Dutt
RTCSA2
2012 Integrated Kernel Partitioning and Scheduling for Coarse-Grained Reconfigurable Arrays
abstract
Coarse-grained reconfigurable arrays (CGRAs) are a promising class of architectures conjugating flexibility and efficiency. Devising effective methodologies to map applications onto CGRAs is a challenging task, due to their parallel execution paradigm and constrained hardware resources. In order to handle complex applications, it is important to devise efficient strategies to partition a kernel into pieces that obey resource constraint and methodologies to schedule them on the underlying hardware. In this paper, we tackle these problems by proposing algorithms to address partitioning based on recursive searches over abstract trees. A novel scheduling strategy is also described that, leveraging differences in delays of various operations, is able to efficiently map operations on CGRA architectures. Experimental evidence on kernels derived from a diverse set of data flow graphs and EEMBC benchmarks demonstrate the efficacy of the described methods, which, when combined, achieve a higher runtime performance on a given mesh size than state-of-the-art approaches (as much as 38% for the benchmark applications considered).
Giovanni Ansaloni, Kazuyuki Tanimura, Laura Pozzi 0001, Nikil Dutt
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2012 Introduction to special section SCPS'09
abstract
No abstract available.
Robert P. Dick, Nikil Dutt
ACM Trans. Embed. Comput. Syst.3
2012 Combining code reordering and cache configuration
abstract
The instruction cache is a popular optimization target due to the cache's high impact on system performance and power and because of the cache's predictable temporal and spatial locality. This article is an in depth study on the interaction of code reordering (a long-known technique) and cache configuration (a relatively new technique). Experimental results show that code reordering coupled with cache configuration reveals additional energy savings as high as 10--15% for several benchmarks with reduced cache area as high as 48%. To exploit these additional benefits, we architect and evaluate several design exploration heuristics for combining these two methods.
Ann Gordon-Ross, Frank Vahid, Nikil Dutt
ACM Trans. Embed. Comput. Syst.3
2012 Error-Aware Algorithm/Architecture Coexploration for Video Over Wireless Applications
abstract
In this article, we propose a cross-layer algorithm/architecture coexploration for wireless multimedia systems to coordinate interactions among sublayer optimizers for improvements in energy/QoS/reliability. By exploiting the inherent redundancy in wireless multimedia systems, we generate an expanded design space over traditional layer-specific approaches. Specifically, we control the error resilient encoder at the application layer to provide awareness of architectural exploration at the physical layer allowing new design points with lower power consumption via aggressive voltage scaling. While trying to reduce energy consumption, the fault tolerant technique compensates the effect of the hardware and network errors due to aggressive voltage scaling and lossy transmission, respectively. Our experiments on H.263 video over a WCDMA communication system demonstrate that coexploration enlarges the feasible design space, which results in significant power savings of more than 20% in the WCDMA modem.
Amin Khajeh, Minyoung Kim 0002, Nikil Dutt, Ahmed M. Eltawil, Fadi J. Kurdahi
ACM Trans. Embed. Comput. Syst.3
2012 xTune: A formal methodology for cross-layer tuning of mobile embedded systems
abstract
Resource-limited mobile embedded systems can benefit greatly from dynamic adaptation of system parameters. We propose a novel approach that employs iterative tuning using lightweight formal verification at runtime with feedback for dynamic adaptation. One objective of this approach is to enable trade-off analysis across multiple layers (e.g., application, middleware, OS) and predict the possible property violations as the system evolves dynamically over time. Specifically, an executable formal specification is developed for each layer of the mobile system under consideration. The formal specification is then analyzed using statistical property checking and statistical quantitative analysis, to determine the impact of various resource management policies for achieving desired timing/QoS properties. Integration of formal analysis with dynamic behavior from system execution results in a feedback loop that enables model refinement and further optimization of policies and parameters. We demonstrate the applicability of this approach to the adaptive provisioning of resource-limited distributed real-time systems using a mobile multimedia case study.
Minyoung Kim 0002, Mark-Oliver Stehr, Carolyn L. Talcott, Nikil Dutt, Nalini Venkatasubramanian
ACM Trans. Embed. Comput. Syst.4
2012 EAVE: Error-Aware Video Encoding Supporting Extended Energy/QoS Trade-offs for Mobile Embedded Systems
abstract
Energy/QoS provisioning is challenging for video applications over lossy wireless network with power-constrained mobile handheld devices. In this work, we exploit the inherent error tolerance of video data to generate a range of acceptable operating points by controlling the amount of errors in the system. In particular, we propose an error-aware video encoding technique, EAVE , that intentionally injects errors while ensuring acceptable QoS. The expanded trade-off space generated by EAVE allows system designers to comparatively evaluate different operating points with varying QoS and energy consumption by aggressively exploiting error-resilience attributes, and could potentially result in significant energy savings. The novelty of our approach resides in active exploitation of errors to vary the operating conditions for further optimization of system parameters. Moreover, we present the adaptivity of our approach by incorporating the feedback from the decoding side to achieve the QoS requirement under the dynamic network status. Our experiments show that EAVE can reduce the energy consumption for an encoding device by up to 37% for a video conferencing application over a wireless network without quality degradation, compared to a standard video encoding technique over test video streams. Further, our experimental results demonstrate that EAVE can expand the design space by 14 times with respect to energy consumption and by 13 times with respect to video quality (compared to a traditional approach without active error exploitation) on average, over test video streams.
Kyoungwoo Lee, Nikil Dutt, Nalini Venkatasubramanian
ACM Trans. Embed. Comput. Syst.2
2011 FFT-cache: a flexible fault-tolerant cache architecture for ultra low voltage operation
abstract
Caches are known to consume a large part of total microprocessor power. Traditionally, voltage scaling has been used to reduce both dynamic and leakage power in caches. However, aggressive voltage reduction causes process-variation-induced failures in cache SRAM arrays, which compromise cache reliability. In this paper, we propose Flexible Fault-Tolerant Cache (FFT-Cache) that uses a flexible defect map to configure its architecture to achieve significant reduction in energy consumption through aggressive voltage scaling, while maintaining high error reliability. FFT-Cache uses a portion of faulty cache blocks as redundancy -- using block-level or line-level replication within or between sets to tolerate other faulty caches lines and blocks. Our configuration algorithm categorizes the cache lines based on degree of conflict of their blocks to reduce the granularity of redundancy replacement. FFT-Cache thereby sacrifices a minimal number of cache lines to avoid impacting performance while tolerating the maximum amount of defects. Our experimental results on SPEC2K benchmarks demonstrate that the operational voltage can be reduced down to 375mV, which achieves up to 80% reduction in dynamic power and up to 48% reduction in leakage power with small performance impact and area overhead.
Abbas BanaiyanMofrad, Houman Homayoun, Nikil Dutt
CASES3
2011 Slack-aware scheduling on Coarse Grained Reconfigurable Arrays
abstract
Coarse Grained Reconfigurable Arrays (CGRAs) are a promising class of architectures conjugating flexibility and efficiency. Devising effective methodologies to map applications onto CGRAs is a challenging task, due to their parallel execution paradigm and sparse interconnection topology. In this paper we present a scheduling framework that is able to efficiently map operations on CGRA architectures. It leverages differences in delays of various operations, which a reconfigurable architecture always exhibits at run-time, to effectively route data. We call this ability “slack-awareness”. Experimental evidence showcases the benefit of slack-aware scheduling in a coarse-grained re-configurable environment, as more complex applications can be mapped for a given mesh size and more efficient schedules can be achieved, compared to the state of the art methods.
Giovanni Ansaloni, Laura Pozzi 0001, Kazuyuki Tanimura, Nikil Dutt
DATE4
2011 E-RoC: Embedded RAIDs-on-Chip for low power distributed dynamically managed reliable memories
abstract
The dual effects of larger die sizes and technology scaling, combined with aggressive voltage scaling for power reduction, increase the error rates for on-chip memories. Traditional on-chip memory reliability techniques (e.g., ECC) incur significant power and performance overheads. In this paper, we propose a low-power-and-performance-overhead Embedded RAID (E-RAID) strategy and present Embedded RAIDs-on-Chip (E-RoC), a distributed dynamically managed reliable memory subsystem. E-RoC achieves reliability through redundancy by optimizing RAID-like policies tuned for on-chip distributed memories. We achieve on-chip reliability of memories through the use of distributed dynamic scratch pad allocatable memories (DSPAMs) and their allocation policies. We exploit aggressive voltage scaling to reduce power consumption overheads due to parallel DSP AM accesses, and rely on the E-RoC manager to automatically handle any resulting voltage-scaling-induced errors. Our experimental results on multimedia benchmarks show that E-RoC's fully distributed redundant reliable memory subsystem reduces power consumption by up to 85% and latency up to 61 % over traditional reliability approaches that use parity/cyclic hybrids for error checking and correction.
Luis Angel D. Bathen, Nikil Dutt
DATE2
2011 Neuromorphic modeling abstractions and simulation of large-scale cortical networks
abstract
Biological neural systems are well known for their robust and power-efficient operation in highly noisy environments. We outline key modeling abstractions for the brain and focus on spiking neural network models. We discuss aspects of neuronal processing and computational issues related to modeling these processes. Although many of these algorithms can be efficiently realized in specialized hardware, we present a case study of simulation of the visual cortex using a GPU based simulation environment that is readily usable by neuroscientists and computer scientists and efficient enough to construct very large networks comparable to brain networks.
Jeffrey L. Krichmar, Nikil Dutt, Jayram Moorkanikara Nageswaran, Micah Richert
ICCAD2
2011 Mapping Multi-Domain Applications Onto Coarse-Grained Reconfigurable Architectures
abstract
Coarse-grained reconfigurable architectures (CGRAs) have drawn increasing attention due to their performance and flexibility. However, their applications have been restricted to domains based on integer arithmetic since typical CGRAs support only integer arithmetic or logical operations. This paper introduces approaches to mapping applications onto CGRAs supporting both integer and floating-point arithmetic. After presenting an optimal formulation using integer linear programming, we present a fast heuristic mapping algorithm. Our experiments on randomly generated examples generate optimal mapping results using our heuristic algorithm for 97% of the examples within a few seconds. We observe similar results for practical examples from multimedia and 3-D graphics benchmarks. The applications mapped on a CGRA show up to 120 times performance improvement compared to software implementations, demonstrating the potential for application acceleration on CGRAs supporting floating-point operations.
Ganghee Lee, Kiyoung Choi, Nikil Dutt
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2011 A Multi-Granularity Power Modeling Methodology for Embedded Processors
abstract
With power becoming a major constraint for multiprocessor embedded systems, it is becoming important for designers to characterize and model processor power dissipation. It is critical for these processor power models to be useable across various modeling abstractions in an electronic system level (ESL) design flow, to guide early design decisions. In this paper, we propose a unified processor power modeling methodology for the creation of power models at multiple granularity levels that can be quickly mapped to an ESL design flow. Our experimental results based on applying the proposed methodology on the OpenRISC and MIPS processors demonstrate the usefulness of having multiple power models. The generated models range from very high-level two-state and architectural/instruction set simulator models that can be used in transaction level models, to extremely detailed cycle-accurate models that enable early exploration of power optimization techniques. These models offer a designer tremendous flexibility to trade off estimation accuracy with estimation/simulation effort.
Young-Hwan Park, Sudeep Pasricha, Fadi J. Kurdahi, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.4
2010 E < MC2: less energy through multi-copy cache
abstract
Caches are known to consume a large part of total microprocessor power. Traditionally, voltage scaling has been used to reduce both dynamic and leakage power in caches. However, aggressive voltage reduction causes process-variation-induced failures in cache SRAM arrays, which compromise cache reliability. We present Multi-Copy Cache (MC2), a new cache architecture that achieves significant reduction in energy consumption through aggressive voltage scaling, while maintaining high error resilience (reliability) by exploiting multiple copies of each data item in the cache. Unlike many previous approaches, MC2 does not require any error map characterization and therefore is responsive to changing operating conditions (e.g., Vdd-noise, temperature and leakage) of the cache. MC2 also incurs significantly lower overheads compared to other ECC-based caches. Our experimental results on embedded benchmarks demonstrate that MC2 achieves up to 60% reduction in energy and energy-delay product (EDP) with only 3.5% reduction in IPC and no appreciable area overhead.
Arup Chakraborty, Houman Homayoun, Amin Khajeh, Nikil Dutt, Ahmed M. Eltawil, Fadi J. Kurdahi
CASES4
2010 RELOCATE: Register File Local Access Pattern Redistribution Mechanism for Power and Thermal Management in Out-of-Order Embedded Processor
Houman Homayoun, Aseem Gupta, Alexander V. Veidenbaum, Avesta Sasan, Fadi J. Kurdahi, Nikil Dutt
HiPEAC6
2010 Towards reverse engineering the brain: Modeling abstractions and simulation frameworks
abstract
Biological neural systems are well known for their robust and power-efficient operation in highly noisy environments. Biological circuits are made up of low-precision, unreliable and massively parallel neural elements with highly reconfigurable and plastic connections. Two of the most interesting properties of the neural systems are its self-organizing capabilities and its template architecture. Recent research in spiking neural networks has demonstrated interesting principles about learning and neural computation. Understanding and applying these principles to practical problems is only possible if large-scale spiking neural simulators can be constructed. Recent advances in low-cost multiprocessor architectures make it possible to build large-scale spiking network simulators. In this paper we review modeling abstractions for neural circuits and frameworks for modeling, simulating and analyzing spiking neural networks.
Jayram Moorkanikara Nageswaran, Micah Richert, Nikil Dutt, Jeffrey L. Krichmar
VLSI-SoC3
2010 Partitioning techniques for partially protected caches in resource-constrained embedded systems
abstract
Increasing exponentially with technology scaling, the soft error rate even in earth-bound embedded systems manufactured in deep subnanometer technology is projected to become a serious design consideration. Partially protected cache (PPC) is a promising microarchitectural feature to mitigate failures due to soft errors in power, performance, and cost sensitive embedded processors. A processor with PPC maintains two caches, one protected and the other unprotected, both at the same level of memory hierarchy. The intuition behind PPCs is that not all data in the application is equally prone to soft errors. By finding and mapping the data that is more prone to soft errors to the protected cache, and error-resilient data to the unprotected cache, failures induced by soft errors can be significantly reduced at a minimal power and performance penalty. Consequently, the effectiveness of PPCs critically hinges on the compiler's ability to partition application data into error-prone and error-resilient data. The effectiveness of PPCs has previously been demonstrated on multimedia applications—where an obvious partitioning of data exists, the multimedia data is inherently resilient to soft errors, and the rest of the data and the entire code is assumed to be error-prone. Since the amount of multimedia data is a quite significant component of the entire application data, this obvious partitioning is quite effective. However, no such obvious data and code partitioning exists for general applications. This severely restricts the applicability of PPCs to data caches and instruction caches in general. This article investigates vulnerability-based partitioning schemes that are applicable to applications in general and effectively reduce failures due to soft errors at minimal power and performance overheads. Our experimental results on an HP iPAQ-like processor enhanced with PPC architecture, running benchmarks from the MiBench suite demonstrate that our partitioning heuristic efficiently finds page partitions for data PPCs that can reduce the failure rate by 48% at only 2% performance and 7% energy overhead, and finds page partitions for instruction PPCs that reduce the failure rate by 50% at only 2% performance and 8% energy overhead, on average.
Kyoungwoo Lee, Aviral Shrivastava, Nikil Dutt, Nalini Venkatasubramanian
ACM Trans. Design Autom. Electr. Syst.3
2010 Bandwidth Management in Application Mapping for Dynamically Reconfigurable Architectures
abstract
Partial dynamic reconfiguration (often referred to as partial RTR) enables true on-demand computing. In an on-demand computing environment, a dynamically invoked application is assigned resources such as data bandwidth, configurable logic. The limited logic resources are customized during application execution by exploiting partial RTR. In this article, we propose an approach that maximizes application performance when available bandwidth and logic resources are limited. Our proposed approach is based on theoretical principles of minimizing application schedule length under bandwidth and logic resource constraints. It includes detailed microarchitectural considerations on a commercially popular reconfigurable device, and it exploits partial RTR very effectively by utilizing data-parallelism property of common image-processing applications. We present extensive application case studies on a cycle-accurate simulation platform that includes detailed resource considerations of the Xilinx Virtex XC2V3000. Our experimental results demonstrate that applying our proposed approach to common image-filtering applications leads to 15--20% performance gain in scenarios with limited bandwidth, when compared to prior work that also exploits data-parallelism with RTR but includes simpler bandwidth considerations. Last but not the least, we also demonstrate how our proposed theoretical principles can be directly applied to solve related problems such as minimizing schedule length under logic resource and power constraints.
Sudarshan Banerjee, Elaheh Bozorgzadeh, Juanjo Noguera, Nikil Dutt
ACM Trans. Reconfigurable Technol. Syst.4
2010 Evaluating Carbon Nanotube Global Interconnects for Chip Multiprocessor Applications
abstract
In ultra-deep submicrometer (UDSM) technologies, the current paradigm of using copper (Cu) interconnects for on-chip global communication is rapidly becoming a serious performance bottleneck. In this paper, we perform a system level evaluation of Carbon Nanotube (CNT) interconnect alternatives that may replace conventional Cu interconnects. Our analysis explores the impact of using CNT global interconnects on the performance and energy consumption of several multi-core chip multiprocessor (CMP) applications. Results from our analysis indicate that with improvements in fabrication technology, CNT-based global interconnects can significantly outperform Cu-based global interconnects.
Sudeep Pasricha, Fadi J. Kurdahi, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.3
2010 CAPPS: A Framework for Power-Performance Tradeoffs in Bus-Matrix-Based On-Chip Communication Architecture Synthesis
abstract
On-chip communication architectures have a significant impact on the power consumption and performance of emerging chip multiprocessor (CMP) applications. However, customization of such architectures for an application requires the exploration of a large design space. Designers need tools to rapidly explore and evaluate relevant communication architecture configurations exhibiting diverse power and performance characteristics. In this paper, we present an automated framework for fast system-level, application-specific, power-performance tradeoffs in a bus matrix communication architecture synthesis (CAPPS). Our study makes two specific contributions. First, we develop energy models for system-level exploration of bus matrix communication architectures. Second, we incorporate these models into a bus matrix synthesis flow that enables designers to efficiently explore the power-performance design space of different bus matrix configurations. Experimental results show that our energy macromodels incur less than 5% average cycle energy error across 180-65 nm technology libraries. Our early system-level power estimation approach also shows a significant speedup ranging from 1000 to 2000× when compared with detailed gate-level power estimation. Furthermore, on applying our synthesis framework to three industrial networking CMP applications, a tradeoff space that exhibits up to 20% variation in power and up to 40% variation in performance is generated, demonstrating the usefulness of our approach.
Sudeep Pasricha, Young-Hwan Park, Fadi J. Kurdahi, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.4
2009 Dynamically reconfigurable on-chip communication architectures for multi use-case chip multiprocessor applications
abstract
The phenomenon of digital convergence and increasing application complexity today is motivating the design of chip multiprocessor (CMP) applications with multiple use cases. Most traditional on-chip communication architecture design techniques perform synthesis and optimization only for a single use-case, which may lead to sub-optimal design decisions for multi-use case applications. In this paper we present a framework to generate a dynamically reconfigurable crossbar-based on-chip communication architecture that can support multiple use-case bandwidth and latency constraints. Our framework generates on-chip communication architectures with a low cost, low power dissipation, and with minimal reconfiguration overhead. Results of applying our framework on several networking CMP applications show that our approach is able to generate a crossbar solution with significantly lower cost (2.4times to 3.8times), and lower power dissipation (1.5times to 3.1times), compared to the best previously proposed approach.
Sudeep Pasricha, Nikil Dutt, Fadi J. Kurdahi
ASP-DAC2
2009 TRAM: A tool for Temperature and Reliability Aware Memory Design
abstract
Memories are increasingly dominating Systems on Chip (SoC) designs and thus contribute a large percentage of the total system's power dissipation, area and reliability. In this paper, we present a tool which captures the effects of supply voltage Vddand temperature on memory performance and their interrelationships. We propose a Temperature- and Reliability- Aware Memory Design (TRAM) approach which allows designers to examine the effects of frequency, supply voltage, power dissipation, and temperature on reliability in a mutually interrelated manner. Our experimental results indicate that thermal unaware estimation of probability of error can be off by at least two orders of magnitude and up to five orders of magnitude from the realistic, temperature-aware cases. We also observed that thermal aware Vddselection using TRAM can reduce the total power dissipation by up to 2.5times while attaining an identical predefined limit on errors.
Amin Khajeh, Aseem Gupta, Nikil Dutt, Fadi J. Kurdahi, Ahmed M. Eltawil, Kamal S. Khouri, Magdy S. Abadir
DATE3
2009 Efficient simulation of large-scale Spiking Neural Networks using CUDA graphics processors
abstract
Neural network simulators that take into account the spiking behavior of neurons are useful for studying brain mechanisms and for engineering applications. Spiking neural network (SNN) simulators have been traditionally simulated on large-scale clusters, super-computers, or on dedicated hardware architectures. Alternatively, graphics processing units (GPUs) can provide a low-cost, programmable, and high-performance computing platform for simulation of SNNs. In this paper we demonstrate an efficient, Izhikevich neuron based large-scale SNN simulator that runs on a single GPU. The GPU-SNN model (running on an NVIDIA GTX-280 with 1 GB of memory), is up to 26 times faster than a CPU version for the simulation of 100 K neurons with 50 million synaptic connections, firing at an average rate of 7 Hz. For simulation of 100 K neurons with 10 million synaptic connections, the GPU-SNN model is only 1.5 times slower than real-time. Further, we present a collection of new techniques related to parallelism extraction, mapping of irregular communication, and compact network representation for effective simulation of SNNs on GPUs. The fidelity of the simulation results were validated against CPU simulations using firing rate, synaptic weight distribution, and inter-spike interval analysis. We intend to make our simulator available to the modeling community so that researchers will have easy access to large-scale SNN simulations.
Jayram Moorkanikara Nageswaran, Nikil Dutt, Jeffrey L. Krichmar, Alexandru Nicolau, Alexander V. Veidenbaum
IJCNN2
2009 Computing Spike-based Convolutions on GPUs
abstract
In spiking neural networks, asynchronous spike events are processed in parallel by neurons. Emulations of such networks are traditionally computed by CPUs or realized using dedicated neuromorphic hardware. In many neuromorphic systems, the Address-Event-Representation (AER) is used for spike communication. In this paper we present the acceleration of AER based spike processing using a Graphics Processing Unit (GPU). In our experiment we interface a 128×128 pixel AER vision sensor to a spiking neural network implemented on a GPU for real-time convolution-based nonlinear feature extraction with convolution kernel sizes ranging from 48×48 to 112×112 pixels. We show parallelism-performance trade-offs on GPUs for single spike per thread, multiple spikes per thread, and multiple objects parallelism techniques. Our implementation can achieve a kernel speedup of up to 35× on a single NVIDIA GTX280 board when compared to a CPU-only implementation.
Jayram Moorkanikara Nageswaran, Nikil Dutt, Tobi Delbruck
ISCAS2
2009 Live Demonstration: Computing Spike-based Convolutions on GPUs
abstract
This demonstration shows the first implementation of a real-time spike-based convolution processing system which combines a spike based dynamic vision sensor (DVS) with parallel graphics processor unit (GPU) computation. Moving objects with different features (shape and size) are presented to the system. In the first demo, the system responses in real time to recognize and keep track of one user specified object and ignore the others. In the second one, the system concurrently extracts several features, and labels the outputs with different colors. Users will enjoy the real-time response and learn about using spike-based sensors combined with conventional procedural processing.
Jayram Moorkanikara Nageswaran, Nikil Dutt, Tobi Delbruck
ISCAS2
2009 A Conservative Approximation Method for the Verification of Preemptive Scheduling Using Timed Automata
abstract
This paper presents a conservative approximation method for the real-time verification of asynchronous event-driven distributed systems. This problem is known to be undecidable in the generic setting. The proposed approach is based on composable timed automata models that provide a sufficient condition to determine schedulability. We demonstrate the method on a real-time CORBA avionics design.
Gabor Madl, Nikil Dutt, Sherif Abdelwahed
IEEE Real-Time and Embedded Technology and Applications Symposium2
2009 A configurable simulation environment for the efficient simulation of large-scale spiking neural networks on graphics processors
Jayram Moorkanikara Nageswaran, Nikil Dutt, Jeffrey L. Krichmar, Alexandru Nicolau, Alexander V. Veidenbaum
Neural Networks2
2009 Adaptive Scratch Pad Memory Management for Dynamic Behavior of Multimedia Applications
abstract
Exploiting runtime memory access traces can be a complementary approach to compiler optimizations for the energy reduction in memory hierarchy. This is particularly important for emerging multimedia applications since they usually have input-sensitive runtime behavior which results in dynamic and/or irregular memory access patterns. These types of applications are normally hard to optimize by static compiler optimizations. The reason is that their behavior stays unknown until runtime and may even change during computation. To tackle this problem, we propose an integrated approach of software [compiler and operating system (OS)] and hardware (data access record table) techniques to exploit data reusability of multimedia applications in Multiprocessor Systems on Chip. Guided by compiler analysis for generating scratch pad data layouts and hardware components for tracking dynamic memory accesses, the scratch pad data layout adapts to an input data pattern with the help of a runtime scratch pad memory manager incorporated in the OS. The runtime data placement strategy presented in this paper provides efficient scratch pad utilization for the dynamic applications. The goal is to minimize the amount of accesses to the main memory over the entire runtime of the system, which leads to a reduction in the energy consumption of the system. Our experimental results show that our approach is able to significantly improve the energy consumption of multimedia applications with dynamic memory access behavior over an existing compiler technique and an alternative hardware technique.
Doosan Cho, Sudeep Pasricha, Ilya Issenin, Nikil Dutt, Minwook Ahn, Yunheung Paek
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2009 Compiler-in-the-Loop Design Space Exploration Framework for Energy Reduction in Horizontally Partitioned Cache Architectures
abstract
Horizontally partitioned caches (HPCs) are a power-efficient architectural feature in which the processor maintains two or more data caches at the same level of hierarchy. HPCs help reduce cache pollution and thereby improve performance. Consequently, most previous research has focused on exploiting HPCs to improve performance and achieve energy reduction only as a byproduct of performance improvement. However, with energy consumption becoming the first class design constraint, there is an increasing need for compilation techniques aimed at energy reduction itself. This paper proposes and explores several low-complexity algorithms aimed at reducing the energy consumption. Acknowledging that the compiler has a significant impact on the energy consumption of the HPCs, Compiler-in-the-Loop Design Space Exploration methodologies are also presented to carefully choose the HPC parameters that result in minimum energy consumption for the application.
Aviral Shrivastava, Ilya Issenin, Nikil Dutt, Yunheung Paek
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2009 Hybrid-compiled simulation: An efficient technique for instruction-set architecture simulation
abstract
Instruction-set simulators are critical tools for the exploration and validation of new processor architectures. Due to the increasing complexity of architectures and time-to-market pressure, performance is the most important feature of an instruction-set simulator. Interpretive simulators are flexible but slow, whereas compiled simulators deliver speed at the cost of flexibility and compilation overhead. This article presents a hybrid instruction-set-compiled simulation (HISCS) technique for generation of fast instruction-set simulators that combines the benefit of both compiled and interpretive simulation. This article makes two important contributions: (i) it improves the interpretive simulation performance by applying compiled simulation at the instruction level using a novel template-customization technique to generate optimized decoded instructions during compile time; and (ii) it reduces the compile-time overhead by combining the benefits of both static and dynamic-compiled simulation. Our experimental results using two contemporary processors (ARM7 and SPARC) demonstrate an order-of-magnitude reduction in compilation time as well as a 70% performance improvement, on average, over the best-known published result in instruction-set simulation.
Mehrdad Reshadi, Prabhat Mishra 0001, Nikil Dutt
ACM Trans. Embed. Comput. Syst.3
2009 Cross-abstraction Functional Verification and Performance Analysis of Chip Multiprocessor Designs
abstract
This paper introduces thecross-abstractionreal-timeanalysis(Carta) framework for the model-based functional verification and performance estimation of chip multiprocessors (CMPs) utilizing bus matrix (crossbar switch) interconnection networks. We argue that the inherent complexity in CMP designs requires the synergistic use of various models of computation to efficiently manage the tradeoffs between accuracy and complexity. Our approach builds on domain-specific modeling languages (DSMLs) driving an open-source tool-chain that provides a cross-abstraction bridge between the finite-state machine (FSM), discrete-event (DE), and timed automata (TA) models of computation, and utilizes multiple model checkers to analyze formal properties at the cycle-accurate and transaction-level abstractions. The cross-abstraction analysis exploits accuracy for functional verification, and achieves significant speedups for performance estimation with marginal accuracy loss. We demonstrate results on an industrial strength networking CMP design utilizing a bus matrix interconnection network. To the best of our knowledge, the Carta framework is the first model-based tool-chain that utilizes multiple abstractions and model checkers for the comprehensive and formal functional verification, performance estimation, and real-time verification of bus matrix-based CMP designs.
Gabor Madl, Sudeep Pasricha, Nikil Dutt, Sherif Abdelwahed
IEEE Trans. Ind. Informatics3
2009 System-level PVT variation-aware power exploration of on-chip communication architectures
abstract
With the shift towards deep submicron (DSM) technologies, the increase in leakage power and the adoption of power-aware design methodologies have resulted in potentially significant variations in power consumption under different process, voltage, and temperature (PVT) corners. In this article, we first investigate the impact of PVT corners on power consumption at the system-on-chip (SoC) level, especially for the on-chip communication infrastructure. Given a target technology library, we then show how it is possible to “scale up” and abstract the PVT variability at the system level, allowing characterization of the PVT-aware design space early in the design flow. We conducted several experiments to estimate power for PVT corner cases, at the gate level, as well as at the higher system level. Our preliminary results are very interesting, and indicate that (i) there are significant variations in power consumption across PVT corners; and (ii) the PVT-aware power estimation problem may be amenable to a reasonably simple abstraction at the system level.
Sudeep Pasricha, Young-Hwan Park, Nikil Dutt, Fadi J. Kurdahi
ACM Trans. Design Autom. Electr. Syst.3
2009 Exploiting Application Data-Parallelism on Dynamically Reconfigurable Architectures: Placement and Architectural Considerations
abstract
Partial dynamic reconfiguration, often called run-time reconfiguration (RTR), is a key feature in modern reconfigurable platforms. In this paper, we present parallelism granularity selection (PARLGRAN), an application mapping approach that maximizes performance of application task chains on architectures with such capability. PARLGRAN essentially selects a suitable granularity of data-parallelism for individual data parallel tasks while considering key issues such as significant reconfiguration overhead and placement constraints. It integrates granularity selection very effectively in a joint scheduling and placement formulation, necessary due to constraints imposed by partial RTR. As a key step to validating PARLGRAN, we additionally present an exact strategy (integer linear programming formulation). We demonstrate that PARLGRAN generates high-quality schedules with: (1) a set of small test cases where we compare our results with the exact strategy; (2) a very large set of synthetic experiments with over a thousand data-points where we compare it with a simpler strategy that tries to statically maximize data-parallelism, i.e., only considers resource availability; and (3) a detailed application case study of JPEG encoding. The application case-study confirms that blindly maximizing data-parallelism can result in schedules even worse than that generated by a simple (but RTR-aware) approach oblivious to data-parallelism. Last, but very important, we demonstrate that our approach is well-suited for true on-demand computing with detailed execution time estimates on a typical embedded processor. Heuristic execution time is comparable to task execution time, i.e., it is feasible to integrate PARLGRAN in a run-time scheduler for dynamically reconfigurable architectures.
Sudarshan Banerjee, Elaheh Bozorgzadeh, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.3
2009 Fast Configurable-Cache Tuning With a Unified Second-Level Cache
abstract
Tuning a configurable cache subsystem to an application can greatly reduce memory hierarchy energy consumption. Previous tuning methods use a level one configurable cache only, or a second level with separate instruction and data configurable caches. We instead use a commercially-common unified second level cache, a seemingly minor difference that actually expands the configuration space from 500 to about 20 000. We develop additive way tuning for tuning a cache subsystem with this large space, yielding 61% energy savings and 9% performance improvements over a nonconfigurable cache, greatly outperforming an extension of a previous method.
Ann Gordon-Ross, Frank Vahid, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.3
2009 Partially Protected Caches to Reduce Failures Due to Soft Errors in Multimedia Applications
abstract
With advances in process technology, soft errors are becoming an increasingly critical design concern. Owing to their large area, high density, and low operating voltages, caches are worst hit by soft errors. Based on the observation that in multimedia applications, not all data require the same amount of protection from soft errors, we propose a partially protected cache (PPC) architecture, in which there are two caches, one protected and the other unprotected at the same level of memory hierarchy. We demonstrate that as compared to the existing unprotected cache architectures, PPC architectures can provide 47 times reduction in failure rate, at only 1% runtime and 3% power overheads. In addition, the failure rate reduction obtained by PPCs is very sensitive to the PPC cache configuration. Therefore, this observation provides an opportunity for further improvement of the solution by correctly parameterizing the PPC configurations. Consequently, we develop design space exploration (DSE) strategies to discover the best PPC configuration. Our DSE technique can reduce the exploration time by more than six times as compared to an exhaustive approach.
Kyoungwoo Lee, Aviral Shrivastava, Ilya Issenin, Nikil Dutt, Nalini Venkatasubramanian
IEEE Trans. Very Large Scale Integr. Syst.4
2008 Quo vadis, BTSoC (Billion Transistor SoC)?
abstract
Billion transistor systems-on-chip (BTSoCs) present designers with a classic case of the "embarrassment-of-riches" syndrome: with so many devices at one's disposal, designers may be tempted to integrate functionality willy-nilly, with no strategic rethinking of what this level of integration can both afford, as well as achieve. While many advocate "business-as- usual" - including ad-hoc integration of functionality to achieve application-specific or domain-dependent designs. The author believes that BTSoCs present us with some opportunities for a paradigm shift in the architectural strategies and design processes for designing such complex chips. The author summarizes some key principles and ideas here.
Nikil Dutt
ASP-DAC1
2008 ORB: An on-chip optical ring bus communication architecture for multi-processor systems-on-chip
abstract
As application complexity continues to increase, multiprocessor systems-on-chip (MPSoC) with tens to hundreds of processing cores are becoming the norm. While computational cores have become faster with each successive technology generation, communication between them has become a bottleneck that limits overall chip performance. On-chip optical interconnects can overcome this bottleneck by replacing electrical wires with optical waveguides. In this paper we propose an optical ring bus (ORB) based on-chip communication architecture for next generation MPSoCs. ORB uses an optical ring waveguide to replace global pipelined electrical interconnects while preserving the interface with today's bus protocol standards such as AMBA AXI. We present experiments to show how ORB has the potential to provide superior performance (more than 2times) and significantly lower power consumption (a reduction of more than 10times) compared to traditionally used pipelined, all-electrical bus-based communication architectures, for 65-22 nm technology nodes.
Sudeep Pasricha, Nikil Dutt
ASP-DAC2
2008 A Compiler-in-the-Loop framework to explore Horizontally Partitioned Cache architectures
abstract
Horizontally Partitioned Caches (HPCs) are a promising architectural feature to reduce the energy consumption of the memory subsystem. However, the energy reduction obtained using HPC architectures is very sensitive to the HPC parameters. Therefore it is very important to explore the HPC design space and carefully choose the HPC parameters that result in minimum energy consumption for the application. However, since in HPC architectures, the compiler has a significant impact on the energy consumption of the memory subsystem, it is extremely important to include compiler while deciding the HPC design parameters. While there has been no previous approaches to HPC design exploration, existing cache design space exploration methodologies do not include the compiler effects during DSE. In this paper, we present a Compiler-in- the-Loop (CIL) Design Space Exploration (DSE) methodology to explore and decide the HPC design parameters. Our experimental results on HP iPAQ h4300-like memory subsystem running benchmarks from the MiBench suite demonstrate that CIL DSE can discover HPC configurations with up to 80% lesser energy consumption than the HPC configuration in the iPAQ. In contrast, tradition simulation-only exploration can discover HPC design parameters that result in only 57% memory subsystem energy reduction. Finally our hybrid CIL DSE heuristic saves 67% of the exploration time as compared to the exhaustive exploration, while providing maximum possible energy savings on our set of benchmarks.
Aviral Shrivastava, Ilya Issenin, Nikil Dutt
ASP-DAC3
2008 ESL hand-off: fact or EDA fiction?
abstract
Moving up in the level of abstraction is the holy grail of EDA. Each transition to the next level of abstraction allows 100x improvement in simulation speed and 100x improvement in design productivity. ESL is being promoted as the next level above RTL but is it really happening?
Hiroyuki Yagi, Wolfgang Roesner, Tim Kogel, Eshel Haritan, Hidekazu Tangi, Michael McNamara, Gary Smith 0001, Nikil Dutt, Giovanni Mancini
DAC8
2008 Memory-aware NoC Exploration and Design
abstract
In the past decade, tremendous progress has been made in NoC research, spanning architectures, protocols and tools. In addition to a large number of academic and research projects, we are now seeing several commercial realizations of NoC- based chip designs. With chip capacities going well beyond the billion transistor mark, on one hand large amounts of the die are occupied by memory resources and on the other hand many complex applications being mapped to these chips are also memory-intensive. In such instances, memories dominate all the axes of traditional design constraints, including, but not limited to performance, area (cost), and power/energy. Furthermore, the move towards sub-nanometer technologies elevates another critical design consideration: process variability and thermal sensitivity, which in turn critically affect the reliability of memories as well. All of these trends make the case for a memory-aware NoC design methodology.
Nikil Dutt
DATE1
2008 Constraint Refinement for Online Verifiable Cross-Layer System Adaptation
abstract
Adaptive resource management is critical to ensuring the quality of real-time distributed applications, particularly for energy-constrained mobile handheld devices. In this context, an optimization that simultaneously considers multiple layers (e.g., application, middleware, operating system) needs to be developed for continuous adaptation of system parameters. The tuning of system parameters greatly affects the system's ability to meet QoS requirements, and also directly affects the energy consumption and system robustness. We present a novel approach to developing cross-layer optimization for resource limited real-time distributed systems, based on a constraint refinement technique combined with formal specification and feedback from system implementation. Our approach tunes the parameters in a compositional manner allowing coordinated interaction among sub-layer optimizers that enables holistic cross-layer optimization. We present experiments on a realistic multimedia application which demonstrate that constraint refinement enables us to generate robust and near optimal parameter settings. The constraint language can be used as an interface for composition by encapsulating the details of local optimization algorithms.
Minyoung Kim 0002, Mark-Oliver Stehr, Carolyn L. Talcott, Nikil Dutt, Nalini Venkatasubramanian
DATE4
2008 Compiler driven data layout optimization for regular/irregular array access patterns
abstract
Embedded multimedia applications consist of regular and irregular memory access patterns. Particularly, irregular pattern are not amenable to static analysis for extraction of access patterns, and thus prevent efficient use of a Scratch Pad Memory (SPM) hierarchy for performance and energy improvements. To resolve this, we present a compiler strategy to optimize data layout in regular/irregular multimedia applications running on embedded multiprocessor environments. The goal is to maximize the amount of accesses to the SPM over the entire system which leads to a reduction in the energy consumption of the system. This is achieved by optimizing data placement of application-wide reused data so that it resides in the SPMs of processing elements. Specifically, our scheme is based on a profiling that generates a memory access footprint. The memory access footprint is used to identify data elements with fine granularity that can profitably be placed in the SPMs to maximize performance and energy gains. We present a heuristic approach that efficiently exploits the SPMs using memory access footprint. Our experimental results show that our approach is able to reduce energy consumption by 30% and improve performance by 18% over cache based memory subsystems for various multimedia applications.
Doosan Cho, Sudeep Pasricha, Ilya Issenin, Nikil Dutt, Yunheung Paek, SunJun Ko
LCTES4
2008 Mitigating the impact of hardware defects on multimedia applications: a cross-layer approach
abstract
Increasing exponentially with each technology generation, hardware-induced soft errors pose a significant threat for the reliability of mobile multimedia devices. Since traditional hardware error protection techniques incur significant power and performance overheads, this paper proposes a cooperative cross-layer approach that exploits existing error control schemes at the application layer to mitigate the impact of hardware defects. Specifically, we propose error detection codes in hardware, drop and forward recovery in middleware, and error-resilient video encoding at the application level to effectively and efficiently combat soft errors with minimal overheads. Experimental evaluation on standard test video streams demonstrates that our cooperative error-aware method for video encoding improves performance by 60% and energy consumption by 58% with even better reliability at the cost of only 3% quality degradation on average, as compared to an error correction code based hardware protection technique. Combining intelligent schemes to select a recovery mechanism can guide system designers to trade off multiple constraints such as performance, power, reliability, and QoS.
Kyoungwoo Lee, Aviral Shrivastava, Minyoung Kim 0002, Nikil Dutt, Nalini Venkatasubramanian
ACM Multimedia4
2008 Data-Reuse-Driven Energy-Aware Cosynthesis of Scratch Pad Memory and Hierarchical Bus-Based Communication Architecture for Multiprocessor Streaming Applications
abstract
As technology advances, it becomes feasible to implement a large multiprocessor systems-on-chip (MPSoCs) to satisfy the increased performance demands of embedded applications. The increased complexity of systems leads to an increased power consumption. Reducing the consumption is an important task, considering that the available power may be limited in battery-operated embedded systems. The selection of memory and communication architectures affects the power efficiency of the design. In this paper, we propose a novel approach that enables the energy-aware cosynthesis of both memory and communication architectures for streaming applications. As opposed to earlier techniques, we propose a powerful compile-time analysis of memory access behavior in multiprocessor systems, which adds flexibility in selecting scratch-pad-based memory architectures. We propose and compare three memory/communication synthesis techniques, namely, an optimal mixed integer-linear-programming (ILP)-based cosynthesis technique, a mixed ILP (MILP)-based traditional two-step synthesis approach, where memory and communication synthesis is sequentially performed, and a cosynthesis heuristic that synthesizes energy-efficient hierarchical bus-based communication architectures with guaranteed throughput. Our experimental results on a number of streaming applications show that both the traditional two-step synthesis approach and heuristic result in up to 50% worse power consumption in comparison with our proposed cosynthesis approach. However, on some of the streaming benchmarks, our cosynthesis heuristic approach was able to find optimal or near-optimal results in a much shorter time than the MILP cosynthesis approach.
Ilya Issenin, Erik Brockmeyer, Bart Durinck, Nikil Dutt
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2008 Register File Power Reduction Using Bypass Sensitive Compiler
abstract
This paper explores, develops, and investigates several bypass-sensitive compilation techniques to reduce the register file power by reducing the access frequency to the register file. We study the effectiveness of our techniques on the Intel XScale processor, which is based on the previously proposed ldquoon-demand register fetch readrdquo architectural feature. Furthermore, we show that our bypass-sensitive compilation technique is effective on various partial bypass configurations.
Aviral Shrivastava, Nikil Dutt, Alexandru Nicolau, Yunheung Paek, Eugene Earlie
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2008 Energy-aware cosynthesis of real-time multimedia applications on MPSoCs using heterogeneous scheduling policies
abstract
Real-time multimedia applications are increasingly being mapped onto MPSoC (multiprocessor system-on-chip) platforms containing hardware--software IPs (intellectual property), along with a library of common scheduling policies such as EDF, RM. The choice of a scheduling policy for each IP is a key decision that greatly affects the design's ability to meet real-time constraints, and also directly affects the energy consumed by the design. We present a cosynthesis framework for design space exploration that considers heterogeneous scheduling while mapping multimedia applications onto such MPSoCs. In our approach, we select a suitable scheduling policy for each IP such that system energy is minimized—our framework also includes energy-reduction techniques utilizing dynamic power management. Experimental results on a realistic multimode multimedia terminal application demonstrate that our approach enables us to select design points with up to 60.5% reduced energy for a given area constraint, while meeting all real-time requirements. More importantly, our approach generates a tradeoff space between energy and cost allowing designers to comparatively evaluate multiple system level mappings.
Minyoung Kim 0002, Sudarshan Banerjee, Nikil Dutt, Nalini Venkatasubramanian
ACM Trans. Embed. Comput. Syst.3
2008 Fast exploration of bus-based communication architectures at the CCATB abstraction
abstract
Currently, system-on-chip (SoC) designs are becoming increasingly complex, with more and more components being integrated into a single SoC design. Communication between these components is increasingly dominating critical system paths and frequently becomes the source of performance bottlenecks. It, therefore, becomes imperative for designers to explore the communication space early in the design flow. Traditionally, system designers have used Pin-Accurate Bus Cycle Accurate (PA-BCA) models for early communication space exploration. These models capture all of the bus signals and strictly maintain cycle accuracy, which is useful for reliable performance exploration but results in slow simulation speeds for complex, designs, even when they are modeled using high-level languages. Recently, there have been several efforts to use the Transaction-Level Modeling (TLM) paradigm for improving simulation performance in BCA models. However, these transaction-based BCA (T-BCA) models capture a lot of details that can be eliminated when exploring communication architectures. In this paper, we extend the TLM approach and propose a new transaction-based modeling abstraction level (CCATB) to explore the communication design space. Our abstraction level bridges the gap between the TLM and BCA levels, and yields an average performance speedup of 120% over PA-BCA and 67% over T-BCA models, on average. The CCATB models are not only faster to simulate, but also extremely accurate and take less time to model compared to both T-BCA and PA-BCA models. We describe the mechanisms that produce the speedup in CCATB models and also analyze how the achieved simulation speedup scales with design complexity. To demonstrate the effectiveness of using CCATB for exploration, we present communication space exploration case studies from the broadband communication and multimedia application domains.
Sudeep Pasricha, Nikil Dutt, Mohamed Ben-Romdhane
ACM Trans. Embed. Comput. Syst.2
2008 Editorial
abstract
No abstract available.
Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.1
2008 Editorial
abstract
No abstract available.
Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.1
2008 Editorial
abstract
No abstract available.
Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.1
2008 Specification-driven directed test generation for validation of pipelined processors
abstract
Functional validation is a major bottleneck in pipelined processor design due to the combined effects of increasing design complexity and lack of efficient techniques for directed test generation. Directed test vectors can reduce overall validation effort, since shorter tests can obtain the same coverage goal compared to the random tests. This article presents a specification-driven directed test generation methodology. The proposed methodology makes three important contributions. First, a general graph model is developed that can capture the structure and behavior (instruction set) of a wide variety of pipelined processors. The graph model is generated from the processor specification. Next, we propose a functional fault model that is used to define the functional coverage for pipelined architectures. Finally, we propose two complementary test generation techniques: test generation using model checking, and test generation using template-based procedures. These test generation techniques accept the graph model of the architecture as input and generate test programs to detect all the faults in the functional fault model. Our experimental results on two pipelined processor models demonstrate several orders-of-magnitude reduction in overall validation effort by drastically reducing both test-generation time and number of test programs required to achieve a coverage goal.
Prabhat Mishra 0001, Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.2
2007 LEAF: A System Level Leakage-Aware Floorplanner for SoCs
abstract
Process scaling and higher leakage power have resulted in increased power densities and elevated die temperatures. Due to the interdependence of temperature and leakage power, we observe that the floorplan has an impact on both the temperatures and the leakage of the IP-blocks in a system on chip (SoC). Hence, in this paper we propose a novel system level leakage aware floorplanner (LEAF) which optimizes floorplans for temperature-aware leakage power along with the traditional metrics of area and wire length. Our floorplanner takes a SoC netlist and the dynamic power profile of functional blocks to determine a placement while optimizing for temperature dependent leakage power, area, and wire length. To demonstrate the effectiveness of LEAF, we implemented our methodology on ten industrial SoC designs from Freescale Semiconductor Inc. and evaluated the trade-off between leakage power and area. We observed up to 190% difference in the leakage power between leakage-unaware and leakage aware floorplanning.
Aseem Gupta, Nikil Dutt, Fadi J. Kurdahi, Kamal S. Khouri, Magdy S. Abadir
ASP-DAC2
2007 Software controlled memory layout reorganization for irregular array access patterns
abstract
Many embedded array-intensive applications have irregular access patterns that are not amenable to static analysis for extraction of access patterns, and thus prevent efficient use of a Scratch Pad Memory (SPM) hierarchy for performance and power improvement. We present a profiling based strategy that generates a memory access trace which can be used to identify data elements with fine granularity that can profitably be placed in the SPMs to maximize performance and energy gains. We developed an entire toolchain that allows incorporation of the code required to profitably move data to SPMs; visualization of the extracted access pattern after profiling; and evaluation/exploration of the generated application code to steer mapping of data to the SPM to yield performance and energy benefits.We present a heuristic approach that efficiently exploits the SPM using the profiler-driven access pattern behaviors. Experimental results on EEMBC and other industrial codes obtained with our framework show that we are able to achieve 36% energy reduction and reduce execution time by up to 22% compared to a cache based system.
Doosan Cho, Ilya Issenin, Nikil Dutt, Jonghee W. Yoon, Yunheung Paek
CASES3
2007 Selective Band width and Resource Management in Scheduling for Dynamically Reconfigurable Architectures
abstract
Partial dynamic reconfiguration (often referred to as partial RTR) enables true on-demand computing. A dynamically invoked application is assigned resources such as data bandwidth, configurable logic, and the limited logic resources are customized during application execution with partial RTR. In this work, we present key theoretical principles for maximizing application performance when available bandwidth is limited. We exploit bandwidth very effectively by selecting a suitable clock frequency for each task and maximize performance with partial RTR by exploiting data-parallelism property of common image-processing tasks. Our theoretical principles are integrated in our scheduling strategy, SCHEDRTR. We present detailed application case studies on a cycle-accurate simulation platform that addresses micro architectural concerns and includes detailed resource considerations of the Virtex XC2V3000 device. Our results demonstrate that applying SCHEDRTR to common image-filtering applications leads to 15--20% performance gain in scenarios with limited bandwidth, when compared to a sophisticated RTR scheduling strategy with data-parallelism but simpler bandwidth considerations.
Sudarshan Banerjee, Elaheh Bozorgzadeh, Nikil Dutt, Juanjo Noguera
DAC3
2007 Interactive presentation: Functional and timing validation of partially bypassed processor pipelines
Qiang Zhu 0008, Aviral Shrivastava, Nikil Dutt
DATE3
2007 Performance estimation of distributed real-time embedded systems by discrete event simulations
abstract
Key challenges in the performance estimation of distributed real-time embedded (DRE) systems include the systematic measurement of coverage by simulations, and the automated generation of directed test vectors. This paper investigates how DRE systems can be represented as discrete event systems (DES) in continuous time, and proposes an automated method for the performance evaluation of such systems. The proposed method also provides a way for the verification of dense time properties for a large class of DRE systems. This approach provides a formal executable model allowing to bridge the gap between simulations and formal verification. Our results show that the proposed DES-based evaluation method can achieve better coverage in large-scale DRE systems than alternative methods.
Gabor Madl, Nikil Dutt, Sherif Abdelwahed
EMSOFT2
2007 System level power estimation methodology with H.264 decoder prediction IP case study
abstract
This paper presents a methodology to generate a hierarchy of power models for power estimation of custom hardware IP blocks, enabling a trade-off between power estimation accuracy, modeling effort and estimation speed. Our power estimation approach enables several novel system-level explorations - such as observing the effect of clock gating, and the effects of tweaking application-level parameters on system power - with an estimation accuracy that is close to the gate-level. We implemented our methodology on an H.264 video decoder prediction IP case study, created power models, and evaluated the effects of varying design parameters (e.g., clock gating, IIP frame ratios, quantization), allowing rapid system-level power exploration of these design parameters.
Young-Hwan Park, Sudeep Pasricha, Fadi J. Kurdahi, Nikil Dutt
ICCD4
2007 Annotation Integration and Trade-off Analysis for Multimedia Applications
abstract
Multimedia applications for mobile devices, such as video/audio streaming, process streams of incoming data in a regular, predictable way. Content-aware optimizations through annotations allow us to highly improve the power savings at the various levels of abstraction: hardware/OS, network, application. However, in a typical system there is a continuous interaction between the components of the system at all levels, which requires a careful analysis of the combined effect of the aforementioned techniques. We investigate such an interaction and we describe metrics for estimating the effect various trade-off have on power and quality. By applying our metrics at the various abstraction levels we show how better energy savings can be achieved with lower quality degradations, through power-quality trade-offs and cross-layer interaction.
Radu Cornea, Alexandru Nicolau, Nikil Dutt
IPDPS3
2007 DYNAMO: A Cross-Layer Framework for End-to-End QoS and Energy Optimization in Mobile Handheld Devices
abstract
In this paper, we present the design and implementation of a cross-layer framework for evaluating power and performance tradeoffs for video streaming to mobile handheld systems. We utilize a distributed middleware layer to perform joint adaptations at all levels of system hierarchy - applications, middleware, OS, network and hardware for optimized performance and energy benefits. Our framework utilizes an intermediate server in close proximity of the mobile device to perform end-to-end adaptations such as admission control, intelligent network transmission and dynamic video transcoding. The knowledge of these adaptations are then used to drive "on-device" adaptations, which include CPU voltage scaling through OS based soft realtime scheduling, LCD backlight intensity adaptation and network card power management. We first present and evaluate each of these adaptations individually and subsequently report the performance of the joint adaptations. We have implemented our cross-layer framework (called DYNAMO) and evaluated it on Compaq iPaq running Linux using streaming video applications. Our experimental results show that such joint adaptations can result in energy savings as high as 54% over the case where no optimization are used while substantially enhancing the user experience on hand-held systems.
Shivajit Mohapatra, Nikil Dutt, Alexandru Nicolau, Nalini Venkatasubramanian
IEEE J. Sel. Areas Commun.2
2007 Introduction of Architecturally Visible Storage in Instruction Set Extensions
abstract
Instruction set extensions (ISEs) can be used effectively to accelerate the performance of embedded processors. The critical and difficult task of ISE selection is often performed manually by designers. A few automatic methods for ISE generation have shown good capabilities but are still limited in the handling of memory accesses, and so they fail to directly address the memory wall problem. We present here the first ISE identification technique that can automatically identify state-holding application-specific functional units (AFUs) comprehensively, thus being able to eliminate a large portion of memory traffic from cache and the main memory. Our cycle-accurate results obtained by the SimpleScalar simulator show that the identified AFUs with architecturally visible storage gain significantly more than previous techniques and achieve an average speedup of 2.8times over pure software execution with a little area overhead. Moreover, the number of required memory-access instructions is reduced by two thirds on average, suggesting corresponding benefits on energy consumption
Partha Biswas, Nikil Dutt, Laura Pozzi 0001, Paolo Ienne
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 A Framework for Cosynthesis of Memory and Communication Architectures for MPSoC
abstract
Memory and communication architectures have a significant impact on the cost, performance, and time-to-market of complex multiprocessor system-on-chip (MPSoC) designs. The memory architecture dictates most of the data traffic flow in a design, which in turn influences the design of the communication architecture. Thus, there is a need to cosynthesize the memory and communication architectures to avoid making suboptimal design decisions. This is in contrast to traditional platform-based design approaches where memory and communication architectures are synthesized separately. In this paper, the authors propose an automated application-specific cosynthesis framework for memory and communication architecture (COSMECA) in MPSoC designs. The primary objective is to design a communication architecture having the least number of buses, which satisfies performance and memory-area constraints, while the secondary objective is to reduce the memory-area cost. Results of applying COSMECA to several industrial strength MPSoC applications from the networking domain indicate a saving of as much as 40% in number of buses and 29% in memory area compared to the traditional approach
Sudeep Pasricha, Nikil Dutt
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 BMSYN: Bus Matrix Communication Architecture Synthesis for MPSoC
abstract
Modern multiprocessor system-on-chip designs have high bandwidth constraints which must be satisfied by the underlying communication architecture. Traditional hierarchical shared bus communication architectures can only support limited bandwidths and are not scalable for very high-performance designs. Bus matrix-based communication architectures consist of several parallel busses which provide a suitable backbone to support high-bandwidth systems but suffer from high-cost overhead due to extensive bus wiring inside the matrix. Manual traversal of the vast exploration space to synthesize a minimal cost bus matrix that also satisfies performance constraints is practically infeasible. In this paper, we address this problem by proposing an automated approach for synthesizing a bus matrix communication architecture, which satisfies all performance constraints in the design and minimizes wire congestion in the matrix. To validate our approach, we consider several industrial strength applications from the networking domain and show that our approach results in up to 9times component savings when compared to a full bus matrix, and up to 3.2times savings when compared to a maximally connected reduced bus matrix, while satisfying all performance constraints in the design.
Sudeep Pasricha, Nikil Dutt, Mohamed Ben-Romdhane
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Automatic Design Space Exploration of Register Bypasses in Embedded Processors
abstract
Register bypassing is a popular and powerful architectural feature to improve processor performance in pipelined processors by eliminating certain data hazards. However, extensive bypassing comes with a significant impact on cycle time, area, and power consumption of the processor. Recent research therefore advocates the use of partial bypassing in a processor. However, accurate performance evaluation of partially bypassed processors is still a challenge, primarily due to the lack of bypass-sensitive retargetable compilation techniques. No existing partial bypass exploration framework estimates the power and area overhead of partial bypassing. As a result, the designers end up making suboptimal design decisions during the exploration of partial bypass design space. This paper presents PBExplore - an automatic design-space-exploration framework for register bypasses. PBExplore accurately evaluates the performance of a partially bypassed processor using a bypass-sensitive compilation technique. It synthesizes the bypass control logic and estimates the area and energy overhead of each bypass configuration. PBExplore is thus able to effectively perform multidimensional exploration of the partial bypass design space. We present experimental results of benchmarks from the MiBench suite on the Intel XScale architecture on and demonstrate the need, utility, and exploration capabilities of PBExplore.
Aviral Shrivastava, Eugene Earlie, Nikil Dutt, Alexandru Nicolau, Yunheung Paek
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2007 Editorial
abstract
No abstract available.
Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.1
2007 DRDU: A data reuse analysis technique for efficient scratch-pad memory management
abstract
In multimedia and other streaming applications, a significant portion of energy is spent on data transfers. Exploiting data reuse opportunities in the application, we can reduce this energy by making copies of frequently used data in a small local memory and replacing speed- and power-inefficient transfers from main off-chip memory by more efficient local data transfers. In this article we present an automated approach for analyzing these opportunities in a program that allows modification of the program to use custom scratch-pad memory configurations comprising a hierarchical set of buffers for local storage of frequently reused data. Using our approach we are able to both reduce energy consumption of the memory subsystem when using a scratch-pad memory by about a factor of two, on average, and improve memory system performance compared to a cache of the same size.
Ilya Issenin, Erik Brockmeyer, Miguel Corbalan, Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.4
2007 Instruction set synthesis with efficient instruction encoding for configurable processors
abstract
Application-specific instructions can significantly improve the performance, energy-efficiency, and code size of configurable processors. While generating new instructions from application-specific operation patterns has been a common way to improve the instruction set (IS) of a configurable processor, automating the design of ISs for given applications poses new challenges---how to create as well as utilize new instructions in a systematic manner, and how to choose the best set of application-specific instructions considering the various effects the new instructions may have on the data path and the compilation? To address these problems, we present a novel IS synthesis framework that optimizes the IS through an efficient instruction encoding for the given application as well as for the given data path architecture. We first build a library of new instructions created with various encoding alternatives taking into account the data path architecture constraints, and then select the best set of instructions while satisfying the instruction bitwidth constraint. We formulate the problem using integer linear programming and also present an effective heuristic algorithm. Experimental results using our technique generate ISs that show improvements of up to about 40% over the native IS for several application benchmarks running on typical embedded RISC processors.
Jongeun Lee, Kiyoung Choi, Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.3
2006 PARLGRAN: parallelism granularity selection for scheduling task chains on dynamically reconfigurable architectures
abstract
Partial dynamic reconfiguration, often called RTR (run-time reconfiguration) is a key feature in modern reconfigurable platforms. While partial RTR enables additional application performance, it imposes physical constraints necessitating simultaneous scheduling and placement while mapping application task graphs onto such architectures. In this paper, we present PARLGRAN, an approach that maximizes performance of application task chains by selecting a suitable granularity of data-parallelism for individual data parallel tasks. Our approach focuses on reconfiguration delay overhead and placement-related issues (such as fragmentation) while selecting individual data-parallelism granularity as an integral part of simultaneous scheduling and placement. We demonstrate that our heuristic generates high-quality schedules on an extensive set of over a 1000 synthetic experiments by comparing the results with an approach that tries to statically maximize data-parallelism, i.e., does not consider the overheads and constraints associated with partial RTR. A detailed case-study on JPEG encoding additionally confirms that blindly maximizing data-parallelism can result in schedules even worse than that generated by a simple (but RTR-aware) approach oblivious to data-parallelism
Sudarshan Banerjee, Elaheh Bozorgzadeh, Nikil Dutt
ASP-DAC3
2006 Memory optimal single appearance schedule with dynamic loop count for synchronous dataflow graphs
abstract
In this paper, we propose a new single appearance schedule for synchronous dataflow programs to minimize data memory and code memory size simultaneously. While a single appearance schedule promises only one appearance of each node definition in the generated code, it requires significant amount of data memory overhead compared with a buffer optimal schedule allowing multiple appearance. The key idea of the proposed technique is to make a dynamic decision of loop count to make a schedule quasi-static. The proposed quasi-static schedule produces a single appearance schedule code with minimum data memory requirement. We prove that every buffer optimal schedule can be transformed to our single appearance schedule which requires optimal buffer size for arbitrary synchronous dataflow graphs. The only penalty for the proposed technique is slight performance overhead of computing loop counts dynamically. In order to minimize the overhead we propose optimization techniques. Experimental results show that the proposed algorithm reduces 20% total memory with less than 1% performance overhead compared with the previous single appearance schedule algorithms.
Hyunok Oh, Nikil Dutt, Soonhoi Ha
ASP-DAC2
2006 Constraint-driven bus matrix synthesis for MPSoC
abstract
Modern multi-processor system-on-chip (MPSoC) designs have high bandwidth constraints which must be satisfied by the underlying communication architecture. Bus matrix based communication architectures consist of several parallel buses which provide a suitable backbone to support high bandwidth systems, but suffer from high cost overhead due to extensive bus wiring inside the matrix. Manual traversal of the vast exploration space to synthesize a minimal cost bus matrix that also satisfies performance constraints is practically infeasible. In this paper, we address this problem by proposing an automated approach for synthesizing a bus matrix communication architecture which satisfies all performance constraints in the design and minimizes wire congestion in the matrix. To validate our approach, we consider several industrial strength applications from the networking domain and show that our approach results in up to 9times component savings when compared to a full bus matrix and up to 3.2times savings when compared to a maximally connected reduced bus matrix
Sudeep Pasricha, Nikil Dutt, Mohamed Ben-Romdhane
ASP-DAC2
2006 Mitigating soft error failures for multimedia applications by selective data protection
abstract
With advances in process technology, soft errors(SE)are becoming an increasingly critical design concern. Due to their large area and high density, caches are worst hit by soft errors. Although Error Correction Code based mechanisms protect the data in caches, they have high performance and power overheads. Since multimedia applications are increasingly being used in mission-critical embedded systems where both reliability and energy are a major concern, there is a de?nite need to improve reliability in embedded systems, without too much energy overhead. We observe that while a soft error in multimedia data may only result in a minor loss in QoS, a soft error in avariable that controls the execution ?ow of the program may be fatal. Consequently, we propose to partition the data space into failure critical and failure non-critical data, and provide a high-degree of soft error protection only to the failure critical data in Horizontally Partitioned Caches. Experimental results demonstrate that our selective data protection can achieve the failure rate close to that of a soft error protected cache system, while retaining the performance and energy consumption similar to those of a traditional cache system, with some degradation in QoS. For example, for conventional con?guration as in IntelXScale, our approach achieves the same failure rate, while improving performance by 28% and reducing energy consumption by 29%in comparison with a soft error protected cache.
Kyoungwoo Lee, Aviral Shrivastava, Ilya Issenin, Nikil Dutt, Nalini Venkatasubramanian
CASES4
2006 A backlight optimization scheme for video playback on mobile devices
abstract
For a typical portable handheld device, the backlight accounts for a significant percentage of the total energy consumption (e.g., around 30% for a Compaq iPAQ 3650). Substantial energy savings can be achieved by dynamically adapting backlight intensity levels on such low- power portable devices. In this paper, we analyze the characteristics of video streaming services and propose an adaptive scheme called Quality Adapted Backlight Scaling (QABS), to achieve backlight energy savings for video playback applications on handheld devices. Specifically, we present a fast algorithm to optimize backlight dimming while keeping the degradation in image quality to a minimum so that the overall service quality is close to a specified threshold. Additionally, we propose two effective techniques to prevent frequent backlight switching, which negatively affects user perception of video. Our initial experimental results indicate that the energy used for backlight is significantly reduced, while the desired quality is satisfied. The proposed algorithms can be realized in real time. I. INTRODUCTION With the widespread availability of 3G cellular networks, mobile hand-held devices are increasingly being designed to support stream- ing video content. These devices have stringent power constraints because they use batteries with finite lifetime. On the other hand, multimedia services are known to be very resource intensive and tend to exhaust battery resources quickly. Therefore, conserving power to prolong battery life is an important research problem that needs to be addressed, specifically for video streaming applications on mobile handheld devices. Most hand-held devices are equipped with a TFT (Thin-Film Transistor) LCD (Liquid Crystal Display). For these devices, the display unit is driven by the illumination of backlight. The backlight consumes a considerable percentage of the total energy usage of the handheld device; it consumes 20%-40% of the total system power (for Compaq iPAQ) (?). Dynamically dimming the backlight is considered an effective method to save energy (?), (?), (?) with scaling up of the pixel luminance to compensate for the reduced fidelity. The luminance scaling, however, tends to saturate the bright part of the picture, thereby affecting the fidelity of the video quality. In (?), a dynamic backlight luminance scaling (DLS) scheme is proposed. Based on different scenarios, three compensation strategies are discussed, i.e., brightness compensation, image enhancement, and context processing. However, their calculation of the distortion does not consider the fact that the clipped pixel values do not contribute equally to the quality distortion. In (?), a similar method, named concurrent brightness and contrast scaling (CBCS), is proposed. CBCS aims at conserving power by reducing the backlight illumi- nation while retaining the image fidelity through preservation of the image contrast. Their distortion definition and proposed compensation technique may be good for static image based applications, such as the graphic user interface (GUI) and maps, but might not be suitable for streaming video scenarios, because their contrast compensation further compromises the fidelity of the images. In addition, neither (?) nor (?) solves the problem associated with frequent backlight switching which can be quite distracting to the end user. In this paper, we explicitly incorporate video quality into the backlight switching strategy and propose a quality adaptive back- light scaling (QABS) scheme. The backlight dimming affects the brightness of the video. Therefore, we only consider the luminance compensation such that the lost brightness can be restored. The lumi- nance compensation, however, inevitably results in quality distortion. For the video streaming application, the quality is normally defined as the resemblance between the original and processed video. Hence, for the sake of simplicity and without loss of generality, we define the quality distortion function as the mean square error (MSE)(see Equation (1)) and the quality function as the peak signal to noise ratio (PSNR)(see Equation (2)), both of which are well accepted objective video quality measurements.
Liang Cheng 0002, Shivajit Mohapatra, Magda El Zarki, Nikil Dutt, Nalini Venkatasubramanian
CCNC4
2006 Multiprocessor system-on-chip data reuse analysis for exploring customized memory hierarchies
abstract
The increasing use of multiprocessor systems-on-chip (MPSoCs) for high performance demands of embedded applications results in high power dissipation. The memory subsystem is a large and critical contributor to both energy and performance, requiring system designers to perform exploration of low power memory organizations. In this paper we present a novel multiprocessor data reuse analysis technique that allows the system designer to explore a wide range of customized memory hierarchy organizations with different size and energy profiles. Our technique enables the system designer to explore feasible memory subsystem solutions that meet power and area constraints while maintaining the necessary performance level. Our experiments on the complex QSDPCM benchmark illustrate the exploration of a wide range of customized memory hierarchies for an MPSoC implementation.
Ilya Issenin, Erik Brockmeyer, Bart Durinck, Nikil Dutt
DAC4
2006 Automatic identification of application-specific functional units with architecturally visible storage
abstract
Instruction set extensions (ISEs) can be used effectively to accelerate the performance of embedded processors. The critical, and difficult task of ISE selection is often performed manually by designers. A few automatic methods for ISE generation have shown good capabilities, but are still limited in the handling of memory accesses, and so they fail to directly address the memory wall problem. We present here the first ISE identification technique that can automatically identify state-holding application-specific functional units (AFUs) comprehensively, thus being able to eliminate a large portion of memory traffic from cache and main memory. Our cycle-accurate results obtained by the SimpleScalar simulator show that the identified AFUs with architecturally visible storage gain significantly more than previous techniques, and achieve an average speedup of 2.8times over pure software execution. Moreover, the number of required memory-access instructions is reduced by two thirds on average, suggesting corresponding benefits on energy consumption
Partha Biswas, Nikil Dutt, Paolo Ienne, Laura Pozzi 0001
DATE2
2006 Software annotations for power optimization on mobile devices
abstract
Modern applications for mobile devices, such as multimedia video/audio, often exhibit a common behavior: they process streams of incoming data in a regular, predictable way. The runtime behavior of these applications can be accurately estimated most of the time by analyzing the data to be processed and annotating the stream with the information collected. We introduce a software annotation based approach to power optimization and demonstrate its application on a backlight adjustment technique for LCD displays during multimedia playback, for improved battery life and user experience. Results from analysis and simulation show that up to 65% of backlight power can be saved through our technique, with minimal or no visible quality degradation
Radu Cornea, Alexandru Nicolau, Nikil Dutt
DATE3
2006 Automatic generation of operation tables for fast exploration of bypasses in embedded processors
abstract
Customizing the bypasses in an embedded processor uncovers valuable trade-offs between the power, performance and the cost of the processor. Meaningful exploration of bypasses requires bypass-sensitive compiler. Operation tables (OTs) have been proposed to perform bypass-sensitive compilation. However, due to lack of automated methods to generate OTs, OTs are currently manually specified by the designer. Manual specification of OTs is not only an extremely time consuming task, but is also highly error-prone. In this paper, we present AutoOT, an algorithm to automatically generate OTs from a high-level processor description. Our experiments on the Intel XScale processor model running MiBench benchmarks demonstrate that AutoOT greatly reduces the time and effort of specification. Automatic generation of OTs makes it feasible to perform full bypass exploration on the Intel XScale and thus discover interesting alternate bypass configurations in a reasonable time. To further reduce the compile-time overhead of OT generation, we propose another novel algorithm, AutoOTDB. AutoOTDB is able to cut the compile-time overhead of OT generation by half
Eugene Earlie, Aviral Shrivastava, Alexandru Nicolau, Nikil Dutt, Yunheung Paek
DATE5
2006 COSMECA: application specific co-synthesis of memory and communication architectures for MPSoC
abstract
Memory and communication architectures have a significant impact on the cost, performance, and time-to-market of complex multi-processor system-on-chip (MPSoC) designs. The memory architecture dictates most of the data traffic flow in a design, which in turn influences the design of the communication architecture. Thus there is a need to co-synthesize the memory and communication architectures to avoid making sub-optimal design decisions. This is in contrast to traditional platform-based design approaches where memory and communication architectures are synthesized separately. In this paper, we propose an automated application specific co-synthesis methodology for memory and communication architectures (COSMECA) in MPSoC designs. The primary objective is to design a communication architecture having the least number of busses, which satisfies performance and memory area constraints, while the secondary objective is to reduce the memory area cost. Results of applying COSMECA to several industrial strength MPSoC applications from the networking domain indicate a saving of as much as 40% in number of busses and 29% in memory area compared to the traditional approach
Sudeep Pasricha, Nikil Dutt
DATE2
2006 Formal performance evaluation of AMBA-based system-on-chip designs
abstract
... (AMBA) is a widely used interconnection standard for SoC design. In order to support high-speed pipelined data transfers, AMBA supports a rich set of bus signals, making the analysis of AMBA-based embedded systems a challenging proposition. This paper makes two main contributions to the analysis and evaluation of AMBA-based SoC designs. The first contribution is to provide a method for the performance analysis and evaluation of AMBA-based SoC designs using formal models. This method provides a way to obtain the end-to-end execution bounds of AMBA-based SoC designs, and guarantees the correctness of the results. The second contribution is to use these formal models to prove the functional correctness of the SoC designs. Using our formal models, we were able to uncover an ambiguous case in the AMBA specification that can lead to deadlocks. This case has not been previously documented by methods focused on AMBA protocol verification. Finally, we validate the proposed performance analysis approach by comparing results with a SystemC implementation of a digital camera case study.
Gabor Madl, Sudeep Pasricha, Luis Angel D. Bathen, Nikil Dutt, Qiang Zhu 0008
EMSOFT4
2006 Minimizing peak power for application chains on architectures with partial dynamic reconfiguration
abstract
Power consumption is a key concern on modern reconfigurable architectures. In this paper, we address the problem of minimizing peak power while mapping application task chains onto reconfigurable architectures with partial dynamic reconfiguration capability. Our proposed methodology minimizes peak power for a given timing constraint. It is based on detailed data-parallelism considerations to ensure that tight timing constraints are met. Our methodology generates physically placed task execution schedules and includes selection of a suitable number of data-parallel instances for each task, a suitable clock frequency, and execution workload for each task instance. Case studies on real image-filtering applications demonstrate that our approach results in significant peak power savings (between 40%-50%) for tight as well as relaxed timing constraints
Sudarshan Banerjee, Elaheh Bozorgzadeh, Juanjo Noguera, Nikil Dutt
FPT4
2006 Video Stream Annotations for Energy Trade-offs in Multimedia Applications
abstract
Recent applications for distributed mobile devices, including multimedia video/audio streaming, typically process streams of incoming data in a regular, predictable way. The behavior of these applications during runtime can be accurately predicted most of the time by analyzing the data to be processed and annotating the stream with the information collected. We introduce an annotation-based approach to power-quality trade-offs and demonstrate its application on CPU frequency scaling during video decoding, for an improved user experience on portable devices. Our experiments show that up to 50% of the power consumed by the CPU during video decoding can be saved with this approach
Radu Cornea, Alexandru Nicolau, Nikil Dutt
ISPDC3
2006 Bypass aware instruction scheduling for register file power reduction
abstract
Since register files suffer from some of the highest power densities within processors, designers have investigated several architectural strategies for register file power reduction, including "On Demand RF Read" where the register file is read only if the operand value is not available from the bypasses. However, we show in this paper that significant additional reductions in the register file power consumption can be obtained by scheduling instructions so that they transfer the operands via bypasses, rather than reading from the register file. Such instruction scheduling requires the compiler to be cognizant of the bypasses in the processor pipeline. In this paper, we develop several bypass aware instruction scheduling heuristics varying in time complexity, and study their effectiveness on the Intel XScale processor pipeline running MiBench benchmarks. Our experimental results show additional power consumption reductions of up to 26% and on average 12% over and above the register file power reduction achieved through existing techniques.
Aviral Shrivastava, Nikil Dutt, Alexandru Nicolau, Yunheung Paek, Eugene Earlie
LCTES3
2006 Generic Processor Modeling for Automatically Generating Very Fast Cycle-Accurate Simulators
abstract
Detailed modeling of processors is required for validating processor behavior and evaluating parameters such as performance and power consumption. Fast cycle-accurate simulators are essential in handling today's complex hardware and software designs at a reasonable time. These problems are challenging enough by themselves and have seen many previous research efforts. Addressing both simultaneously is even more challenging, with many existing approaches focusing on one over another. Abstract models in fast simulators do not provide enough information required for different phases of the design. On the other hand, detailed models are very difficult to generate and result in very slow simulators. In this paper, a modeling approach based on reduced colored Petri net (RCPN) is proposed, which has the following three advantages: 1) it is very generic and support a wide range of processor features; 2) it offers a very simple and intuitive yet formal way of modeling pipelined processors; and 3) it can generate high-performance cycle-accurate simulators. RCPN inherits all useful features of colored Petri nets while avoiding their exponential growth in complexity. In this paper, it is shown how this approach is general enough to model features such as very long instruction word out-of-order execution, dynamic scheduling, register renaming, hazard detection, and branch prediction. Furthermore, the results of generating cycle-accurate simulators from RCPN models of XScale and StrongArm processors are shown, where an order of magnitude (~15 times on the average) speedup over the popular SimpleScalar advanced reduced instruction set computing machine simulator is achieved
Mehrdad Reshadi, Bita Gorjiara, Nikil Dutt
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2006 A retargetable framework for instruction-set architecture simulation
abstract
Instruction-set architecture (ISA) simulators are an integral part of today's processor and software design process. While increasing complexity of the architectures demands high-performance simulation, the increasing variety of available architectures makes retargetability a critical feature of an instruction-set simulator. Retargetability requires generic models while high-performance demands target specific customizations. To address these contradictory requirements, we have developed a generic instruction model and a generic decode algorithm that facilitates easy and efficient retargetability of the ISA-simulator for a wide range of processor architectures, such as RISC, CISC, VLIW, and variable length instruction-set processors. The instruction model is used to generate compact and easy to debug instruction descriptions that are very similar to that of architecture manual. These descriptions are used to generate high-performance simulators. Our retargetable framework combines the flexibility of interpretive simulation with the speed of compiled simulation. The generation of the simulator is completely separate from the simulation engine. Hence, we can incorporate any fast simulation technique in our retargetable framework without introducing any performance penalty. To demonstrate this, we have incorporated fast IS-CS simulation engine in our retargetable framework which has generated 70% performance improvement over the best known simulators in this category. We illustrate the retargetability of our approach using two popular, yet different, realistic architectures: the SPARC and the ARM.
Mehrdad Reshadi, Nikil Dutt, Prabhat Mishra 0001
ACM Trans. Embed. Comput. Syst.2
2006 Editorial
abstract
No abstract available.
Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.1
2006 Architecture description language (ADL)-driven software toolkit generation for architectural exploration of programmable SOCs
abstract
Advances in semiconductor technology permit increasingly complex applications to be realized using programmable systems-on-chips (SOCs). Furthermore, shrinking time-to-market demands, coupled with the need for product versioning through software modification of SOC platforms, have led to a significant increase in the software content of these SOCs. However, designer productivity is greatly hampered by the lack of automated software generation tools for the exploration and evaluation of different architectural configurations. Traditional hardware-software codesign flows do not support effective exploration and customization of the embedded processors used in programmable SOCs. The inherently application-specific nature of embedded processors and the stringent area, power, and performance constraints in embedded systems design critically require a fast and automated architecture exploration methodology. Architecture description language (ADL)-Driven design space exploration and software toolkit generation strategies present a viable solution to this problem, providing a systematic mechanism for a top-down design and validation of complex systems. The heart of this approach lies in the ability to automatically generate a software toolkit that includes an architecture-sensitive compiler, a cycle-accurate simulator, assembler, debugger, and verification/validation tools. This article illustrates a software toolkit generation methodology using the EXPRESSION ADL. Our exploration studies demonstrate the need for and usefulness of this approach, using as an example the problem of compiler-in-the-loop design space exploration of reduced instruction-set embedded processor architectures.
Prabhat Mishra 0001, Aviral Shrivastava, Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.3
2006 Compilation framework for code size reduction using reduced bit-width ISAs (rISAs)
abstract
For many embedded applications, program code size is a critical design factor. One promising approach for reducing code size is to employ a “dual instruction set”, where processor architectures support a normal (usually 32-bit) Instruction Set, and a narrow, space-efficient (usually 16-bit) Instruction Set with a limited set of opcodes and access to a limited set of registers. This feature however, requires compilers that can reduce code size by compiling for both Instruction Sets. Existing compiler techniques operate at the routine-level granularity and are unable to make the trade-off between increased register pressure (resulting in more spills) and decreased code size. We present a compilation framework for such dual instruction sets, which uses a profitability based compiler heuristic that operates at the instruction-level granularity and is able to effectively take advantage of both Instruction Sets. We demonstrate consistent and improved code size reduction (on average 22%), for the MIPS 32/16 bit ISA. We also show that the code compression obtained by this “dual instruction set” technique is heavily dependent on the application characteristics and the narrow Instruction Set itself.
Aviral Shrivastava, Partha Biswas, Ashok Halambi, Nikil Dutt, Alexandru Nicolau
ACM Trans. Design Autom. Electr. Syst.4
2006 Integrating Physical Constraints in HW-SW Partitioning for Architectures With Partial Dynamic Reconfiguration
abstract
Partial dynamic reconfiguration is a key feature of modern reconfigurable architectures such as the Xilinx Virtex series of devices. However, this capability imposes strict placement constraints such that even exact system-level partitioning (and scheduling) formulations are not guaranteed to be physically realizable due to placement infeasibility. We first present an exact approach for hardware-software (HW-SW) partitioning that guarantees correctness of implementation by considering placement implications as an integral aspect of HW-SW partitioning. Our exact approach is based on integer linear programming (ILP) and considers key issues such as configuration prefetch for minimizing schedule length on the target single-context device. Next, we present a physically aware HW-SW partitioning heuristic that simultaneously partitions, schedules, and does linear placement of tasks on such devices. With the exact formulation, we confirm the necessity of physically-aware HW-SW partitioning for the target architecture. We demonstrate that our heuristic generates high-quality schedules by comparing the results with the exact formulation for small tests and with a popular, but placement-uanaware scheduling heuristic for a large set of over a hundred tests. Our final set of experiments is a case study of JPEG encoding-we demonstrate that our focus on physical considerations along with our consideration of multiple task implementation points enables our approach to be easily extended to handle heterogenous architectures (with specialized resources distributed between general purpose programmable logic columns). The execution time of our heuristic is very reasonable-task graphs with hundreds of nodes are processed (partitioned, scheduled, and placed) in a couple of minutes
Sudarshan Banerjee, Elaheh Bozorgzadeh, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.3
2006 ISEGEN: an iterative improvement-based ISE generation technique for fast customization of processors
abstract
Customization of processor architectures through instruction set extensions (ISEs) is an effective way to meet the growing performance demands of embedded applications. A high-quality ISE generation approach needs to obtain results close to those achieved by experienced designers, particularly for complex applications that exhibit regularity: expert designers are able to exploit manually such regularity in the data flow graphs to generate high-quality ISEs. In this paper, we present ISEGEN, an approach that identifies high-quality ISEs by iterative improvement following the basic principles of the well-known Kernighan-Lin min-cut heuristic. Experimental results on a number of MediaBench, EEMBC, and cryptographic applications show that our approach matches the quality of the optimal solution obtained by exhaustive search. We also show that our ISEGEN technique is on average 20times faster than a genetic formulation that generates equivalent solutions. Furthermore, the ISEs identified by our technique exhibit 35% more speedup than the genetic solution on a large cryptographic application by effectively exploiting its regular structure
Partha Biswas, Sudarshan Banerjee, Nikil Dutt, Laura Pozzi 0001, Paolo Ienne
IEEE Trans. Very Large Scale Integr. Syst.3
2006 Energy efficient watermarking on mobile devices using proxy-based partitioning
abstract
Digital watermarking embeds an imperceptible signature or watermark in a digital file containing audio, image, text, or video data. The watermark can be used to authenticate the data file and for tamper detection. It is particularly valuable in the use and exchange of digital media, such as audio and video, on emerging handheld devices. However, watermarking is computationally expensive and adds to the drain of the available energy in handheld devices. In this paper, we first analyze the energy profile of various watermarking algorithms. We also study the impact of security and image quality on energy consumption. Second, we present an approach in which we partition the watermarking embedding and extraction algorithms and migrate some tasks to a proxy server. This leads to a lower energy consumption on the handheld without compromising the security of the watermarking process. Experimental results show that executing the watermarking tasks that are partitioned between the proxy and the handheld devices, reduces the total energy consumed by 80%, and improves performance by two orders of magnitude compared to running the application on only the handheld device
Arun Kejariwal, Alexandru Nicolau, Nikil Dutt, Rajesh K. Gupta 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2006 FABSYN: floorplan-aware bus architecture synthesis
abstract
As system-on-chip (SoC) designs become more complex, it is becoming harder to design communication architectures to handle the ever increasing volumes of inter-component communication. Manual traversal of the vast communication design space to synthesize a communication architecture that meets performance requirements becomes infeasible. In this paper, we address this problem by proposing an automated approach for floorplan-aware bus architecture synthesis (FABSYN) to synthesize cost-effective, bus-based communication architectures that satisfy the performance constraints in a design. Our synthesis approach incorporates a high-level floorplanning and wire delay estimation engine to evaluate the feasibility of the synthesized bus architecture and detect bus cycle time violations early in the design How, at the system level. We present case studies of network communication SoC subsystems for which we synthesized bus architectures, detected and eliminated timing violations, and generated core placements in a matter of hours instead of several days for a manual effort.
Sudeep Pasricha, Nikil Dutt, Elaheh Bozorgzadeh, Mohamed Ben-Romdhane
IEEE Trans. Very Large Scale Integr. Syst.2
2006 Retargetable pipeline hazard detection for partially bypassed processors
abstract
Register bypassing is a widely used feature in modern processors to eliminate certain data hazards. Although complete bypassing is ideal for performance, it has significant impact on the cycle time, area, and power consumption of the processor. Owing to the strict design constraints on the performance, cost, and the power consumption of embedded processor systems, architects seek a compromise between the design parameters by implementing partial bypassing in processors. However, partial bypassing in processors presents challenges for compilation. Traditional data hazard detection and/or avoidance techniques used in retargetable compilers that assume a constant value of operation latency, break down in the presence of partial bypassing. In this article, we present the concept of operation tables (OTs) that can be used to accurately detect data hazards, even in the presence of incomplete bypassing. OTs integrate the detection of all kinds of pipeline hazards in a unified framework, and can, therefore, be easily deployed in a compiler to generate better schedules. Our experimental results on the popular Intel XScale embedded processor running embedded applications from the MiBench suite, demonstrate that accurate pipeline hazard detection by OTs can result in up to 20% performance improvement over the best performing GCC generated code. Finally, we demonstrate the usefulness of OTs over various bypass configurations of the Intel XScale
Aviral Shrivastava, Eugene Earlie, Nikil Dutt, Alexandru Nicolau
IEEE Trans. Very Large Scale Integr. Syst.3
2005 Automated throughput-driven synthesis of bus-based communication architectures
abstract
As System-on-Chip (SoC) designs become more complex, it becomes increasingly harder to design communication architectures which satisfy design constraints. Manually traversing the vast communication design space for constraint-driven synthesis is not feasible anymore. In this paper we propose an approach that automates the synthesis of bus-based communication architectures for systems characterized by (possibly several) throughput constraints. Our approach accurately and effectively prunes the large communication design space to synthesize a feasible low-cost bus architecture which satisfies the constraints in a design.
Sudeep Pasricha, Nikil Dutt, Mohamed Ben-Romdhane
ASP-DAC2
2005 A generalized technique for energy-efficient operating voltage set-up in dynamic voltage scaled processors
abstract
Dynamic voltage scaling (DVS) which is an effective energy minimization technique has been well-studied in recent years. Yet the problem of selecting voltage levels for multiple voltage DVS systems remains an unresolved issue. In this paper, we present a novel technique for dealing with the problem of finding k operating voltages to minimize the energy consumption (voltage set-up problem). A new formulation of the voltage set-up problem is given to make our solution less dependent on the specific DVS scheme. Then it is solved optimally using dynamic programming in polynomial time. With almost the same time complexity we extend the proposed technique to explore the design space to determine the best number of voltage levels. It is confirmed from the experiments that the proposed voltage set-up solution reduces energy consumption by 19.2% on average over that of previous technique [7].
Jaewon Seo, Nikil Dutt
ASP-DAC2
2005 Single appearance schedule with dynamic loop count for minimum data buffer from synchronous dataflow graphs
abstract
In this paper, we propose a new single appearance schedule for synchronous dataflow programs to minimize data memory and code memory size at the same time. When the software code is automatically synthesized from the dataflow program graphs, a single appearance schedule promises only one appearance of each node definition in the generated code. While several heuristics have been developed to find a single appearance schedule, they all have to pay significant amount of data memory overhead compared with a buffer optimal schedule. The key idea of the proposed technique is to make a dynamic decision of loop count to make a schedule quasi-static. The proposed quasi-static static schedule produces a single appearance schedule code with minimum data memory requirement. We prove that the proposed scheduling technique is optimal for a chain-structured graph in terms of data memory requirement while maintaining the single appearance schedule. The only penalty for the proposed technique is slight performance overhead of computing loop counts dynamically. Experimental results show that the proposed algorithm reduces 20% total memory with less than 1% performance overhead compared with the previous single appearance schedule algorithms for CD2DAT and non uniform filter bank applications.
Hyunok Oh, Nikil Dutt, Soonhoi Ha
CASES2
2005 Compilation techniques for energy reduction in horizontally partitioned cache architectures
abstract
Horizontally partitioned data caches are a popular architectural feature in which the processor maintains two or more data caches at the same level of hierarchy. Horizontally partitioned caches help reduce cache pollution and thereby improve performance. Consequently most previous research has focused on exploiting horizontally partitioned data caches to improve performance, and achieve energy reduction only as a byproduct of performance improvement. In constrast, in this paper we show that optimizing for performance trades-off several opportunities for energy reduction. Our experiments on a HP iPAQ h4300-like memory subsystem demonstrate that optimizing for energy consumption results in up to 127% less memory subsystem energy consumption than the performance optimal solution. Furthermore, we show that energy optimal solution incurs on average only 1.7% performance penalty. Therefore, with energy consumption becoming a first class design constraint, there is a need for compilation techniques aimed at energy reduction. To achieve aforementioned energy savings we propose and explore several low-complexity algorithms aimed at reducing the energy consumption and show that very simple greedy heuristics achieve 97% of the possible memory subsystem energy savings.
Aviral Shrivastava, Ilya Issenin, Nikil Dutt
CASES3
2005 Physically-aware HW-SW partitioning for reconfigurable architectures with partial dynamic reconfiguration
abstract
Many reconfigurable architectures offer partial dynamic configurability, but current system-level tools cannot guarantee feasible implementations when exploiting this feature. We present a physically aware hardware-software (HW-SW) scheme for minimizing application execution time under HW resource constraints, where the HW is a reconfigurable architecture with partial dynamic reconfiguration capability. Such architectures impose strict placement constraints that lead to implementation infeasibility of even optimal scheduling formulations that ignore the nature of these constraints. We propose an exact and a heuristic formulation that simultaneously partition, schedule, and do linear placement of tasks on such architectures. With our exact formulation, we prove the critical nature of placement constraints. We demonstrate that our heuristic generates high-quality schedules by comparing the results with the exact formulation for small tests and a popular, but placement-uanaware scheduling heuristic for larger tests. With a case study, we demonstrate extension of our approach to handle heterogenous architectures with specialized resources distributed between general purpose programmable logic columns. The execution time of our heuristic is very reasonable- task graphs with hundreds of nodes are processed in a couple of minutes.
Sudarshan Banerjee, Elaheh Bozorgzadeh, Nikil Dutt
DAC3
2005 Floorplan-aware automated synthesis of bus-based communication architectures
abstract
As System-on-Chip (SoC) designs become more complex, it is becoming harder to design communication architectures to handle the ever increasing volumes of inter-component communication. Manual traversal of the vast communication design space to synthesize a communication architecture that meets performance requirements becomes infeasible. In this paper, we address this problem by proposing an automated approach for synthesizing cost-effective, bus-based communication architectures that satisfy the performance constraints in a design. Our synthesis flow also incorporates a high-level floorplanning and wire delay estimation engine to evaluate the feasibility of the synthesized bus architecture and detect timing violations early in the design flow. We present case studies of network communication SoC subsystems for which we synthesized bus architectures, detected timing violations and generated core placements in a matter of hours instead of several days it took for a manual effort.
Sudeep Pasricha, Nikil Dutt, Elaheh Bozorgzadeh, Mohamed Ben-Romdhane
DAC2
2005 ISEGEN: Generation of High-Quality Instruction Set Extensions by Iterative Improvement
abstract
Customization of processor architectures through instruction set extensions (ISEs) is an effective way to meet the growing performance demands of embedded applications. A high-quality ISE generation approach needs to obtain results close to those achieved by experienced designers, particularly for complex applications that exhibit regularity; expert designers are able to exploit manually such regularity in the data flow graphs to generate high-quality ISEs. We present ISEGEN, an approach that identifies high-quality ISEs by iterative improvement following the basic principles of the well-known Kernighan-Lin (K-L) min-cut heuristic. Experimental results on a number of MediaBench, EEMBC and cryptographic applications show that our approach matches the quality of the optimal solution obtained by exhaustive search. We also show that our ISEGEN technique is on average 20 times faster than a genetic formulation that generates equivalent solutions. Furthermore, the ISEs identified by our technique exhibit 35% more speedup than the genetic solution on a large cryptographic application (AES) by effectively exploiting its regular structure.
Partha Biswas, Sudarshan Banerjee, Nikil Dutt, Laura Pozzi 0001, Paolo Ienne
DATE3
2005 FORAY-GEN: Automatic Generation of Affine Functions for Memory Optimizations
abstract
In today's embedded applications, a significant portion of energy is spent in the memory subsystem. Several approaches have been proposed to minimize this energy, including the use of scratch pad memories, with many based on static analysis of a program. However, it is often not possible to perform static analysis and optimization of a program's memory access behavior unless the program is specifically written for this purpose. We introduce the FORAY model of a program that permits aggressive analysis of the application's memory behavior that further enables such optimization since it consists of 'for' loops and array accesses which are easily analyzable. We present FORAY-GEN, an automated profile-based approach for extraction of the FORAY model from the original program. We also demonstrate how FORAY-GEN enhances applicability of other memory subsystem optimization approaches, resulting in an average doubling in the number of memory references that can be analyzed by existing static approaches.
Ilya Issenin, Nikil Dutt
DATE2
2005 Functional Coverage Driven Test Generation for Validation of Pipelined Processors
abstract
Functional verification of microprocessors is one of the most complex and expensive tasks in the current system-on-chip design process. A significant bottleneck in the validation of such systems is the lack of a suitable functional coverage metric. The paper presents a functional coverage based test generation technique for pipelined architectures. The proposed methodology makes three important contributions. First, a general graph-theoretic model is developed that can capture the structure and behavior (instruction-set) of a wide variety of pipelined processors. Second, we propose a functional fault model that is used to define the functional coverage for pipelined architectures. Finally, test generation procedures are presented that accept the graph model of the architecture as input and generate test programs to detect all the faults in the functional fault model. Our experimental results on two pipelined processor models demonstrate that the number of test programs generated by our approach to obtain a fault coverage is an order of magnitude less than those generated by traditional random or constrained-random test generation techniques.
Prabhat Mishra 0001, Nikil Dutt
DATE2
2005 Generic Pipelined Processor Modeling and High Performance Cycle-Accurate Simulator Generation
abstract
Detailed modeling of processors and high performance cycle-accurate simulators are essential for today's hardware and software design. These problems are challenging enough by themselves and have seen many previous research efforts. Addressing both simultaneously is even more challenging, with many existing approaches focusing on one over another. In this paper, we propose the reduced colored Petri net (RCPN) model that has two advantages: first, it offers a very simple and intuitive way of modeling pipelined processors: second, it can generate high performance cycle-accurate simulators. RCPN benefits from all the useful features of colored Petri nets without suffering from their exponential growth in complexity. RCPN processor models are very intuitive since they are a mirror image of the processor pipeline block diagram. Furthermore, in our experiments on the generated cycle-accurate simulators for XScale and StrongArm processor models, we achieved an order of magnitude (/spl sim/15 times) speedup over the popular SimpleScalar ARM simulator.
Mehrdad Reshadi, Nikil Dutt
DATE2
2005 PBExplore: A Framework for Compiler-in-the-Loop Exploration of Partial Bypassing in Embedded Processors
abstract
Varying partial bypassing in pipelined processors is an effective way to make performance, area and energy tradeoffs in embedded processors. However, performance evaluation of partial bypassing in processors has been inaccurate, largely due to the absence of bypass-sensitive retargetable compilation techniques. Furthermore no existing partial bypass exploration framework estimates the power and cost overhead of partial bypassing. In this paper we present PBExplore: a framework for compiler-in-the-loop exploration of partial bypassing in processors. PBExplore accurately evaluates the performance of a partially bypassed processor using a generic bypass-sensitive compilation technique. It synthesizes the bypass control logic and estimates the area and energy overhead of each bypass configuration. PBExplore is thus able to effectively perform multidimensional exploration of the partial bypass design space. We present experimental results on the Intel XScale architecture on MiBench benchmarks and demonstrate the need, utility and exploration capabilities of PBExplore.
Aviral Shrivastava, Nikil Dutt, Alexandru Nicolau, Eugene Earlie
DATE2
2005 Considering Run-Time Reconfiguration Overhead in Task Graph Transformations for Dynamically Reconfigurable Architectures
abstract
In modern dynamic FPGA-based platforms where multiple processes may be executing concurrently, partial dynamic reconfiguration (RTR) is a key technique for maximizing application performance under resource constraints. For platforms with column-based partial RTR, we propose a new technique to statically transform linear task graphs (common in image processing applications). In our approach, the granularity of data parallelism for each task is determined while considering the reconfiguration overhead along with architectural constraints imposed by partial RTR. On JPEG applications, our technique can improve the execution time by up to 37% by choosing the right granularity of task parallelism.
Sudarshan Banerjee, Elaheh Bozorgzadeh, Nikil Dutt
FCCM3
2005 A first look at the interplay of code reordering and configurable caches
abstract
The instruction cache is a popular target for optimizations of microprocessor-based systems because of the cache's high impact on system performance and power, and because of the cache's predictable temporal and spatial locality. Optimization techniques can be designed based on this predictability. We explore for the first time the interplay of two popular instruction cache optimization techniques: the long-known technique of code reordering and the relatively-new technique of cache configuration. We address the question of whether those two optimizations complement each other or if one optimization dominates the other. Through experiments using embedded system benchmarks, we show that cache configuration dominates a particular category of code reordering techniques with respect to optimizing performance and energy, obviating the need for reordering. We also examine the modern scenario of synthesized custom caches, and show that combining cache configuration with code reordering results in cache size reductions of 13% on average, and up to 89% in some benchmarks, beyond just cache configuration alone.
Ann Gordon-Ross, Frank Vahid, Nikil Dutt
ACM Great Lakes Symposium on VLSI3
2005 Optimal integration of inter-task and intra-task dynamic voltage scaling techniques for hard real-time applications
abstract
It is generally accepted that the dynamic voltage scaling (DVS) is one of the most effective techniques for energy minimization. According to the granularity of units to which voltage scaling is applied, the DVS problem can be divided into two subproblems: (i) inter-task DVS problem; and (ii) intra-task DVS problem. A lot of effective DVS techniques have addressed either one of the two subproblems, but none of them have attempted to solve both simultaneously, which is mainly due to an excessive computation complexity to solve it optimally. This work addresses this core issue, that is, Can the combined problem be solved effectively and efficiently? More specifically, our work shows, for a set of inter-dependent tasks, that the combined DVS problem can be solved optimally in polynomial time. Experimental results indicate that the proposed integrated DVS technique is able to reduce energy consumption by 10.6% on average over the results by (Zhang et al., 2002 and Seo et al., 2004) (i.e., a straightforward combination of two optimal inter- and intra task DVS techniques).
Jaewon Seo, Nikil Dutt
ICCAD3
2005 An Experimental Study on Energy Consumption of Video Encryption for Mobile Handheld Devices
abstract
Secure video communication on mobile handheld devices is challenging mainly due to (a) the significant computational needs of both video coding and encryption algorithms and (b) the limited battery capacity of handheld devices. In this paper, we evaluate several video encryption schemes from the perspective of energy consumption both analytically and experimentally. Specifically, we implement video encryption schemes on mobile handhelds to support a H.263 based secure video application, and extensively measure the energy consumption due to encoding and encryption for several classes of video clips. Contrary to popular belief, our experiments show that energy overhead of full video encryption is insignificant compared to the energy consumed for video encoding (between 2% and 4% of total energy cost) in most cases
Kyoungwoo Lee, Nikil Dutt, Nalini Venkatasubramanian
ICME2
2005 Fast configurable-cache tuning with a unified second-level cache
abstract
Tuning a configurable cache subsystem to an application can greatly reduce memory hierarchy energy consumption. Previous tuning methods use a level one configurable cache only, or a second level with separate instruction and data configurable caches. We instead use a commercially-common unified second level, a seemingly minor difference that actually expands the configuration space from 500 to about 20,000. We develop additive way tuning for tuning a cache subsystem with this large space, yielding 62% energy savings and 35% performance improvements over a non-configurable cache, greatly outperforming an extension of a previous method
Ann Gordon-Ross, Frank Vahid, Nikil Dutt
ISLPED3
2005 Code Size Reduction in Heterogeneous-Connectivity-Based DSPs Using Instruction Set Extensions
abstract
Existing trend of processors shows a progress toward customizable and reconfigurable architectures. In this paper, we study the benefit of combining the architectural design of a VLIW DSP and the concepts of modern customizable processors like ASIPs (application specific instruction set processors) for code size reduction. VLIW DSP architectures exhibit heterogeneous connections between functional units and register files for speeding up special tasks. Such architectural characteristics can be effectively exploited through the use of complex instruction set extensions (ISEs). Although VLIWs are increasingly being used for DSP applications to achieve very high performance, such architectures are known to suffer from increased code size. This paper also addresses how to generate and use ISEs that can result in significant code size reduction in VLIW DSPs without degrading performance. Unfortunately, contemporary techniques for generation of ISEs when applied before resource-binding fail to generate legal ISEs for VLIW architectures with heterogeneous connectivity between the functional units and register files. We propose a heuristic-based approach to generate ISEs for a generalized heterogeneous-connectivity-based VLIW DSP architecture. We achieve an average code size reduction of 25 percent on the MiBench suite with no penalty in performance by applying our ISE generation algorithms on the Tl TMS320C6xx, a representative VLIW DSP. We also show that the overhead of the required architectural assists for our approach is minimal: The TMS320C6xx pipeline meets the required timing with only a limited overhead in area.
Partha Biswas, Nikil Dutt
IEEE Trans. Computers2
2005 Editorial
abstract
No abstract available.
Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.1
2004 Energy efficient code generation exploiting reduced bit-width instruction set architectures (rISA)
Aviral Shrivastava, Nikil Dutt
ASP-DAC2
2004 Introduction of local memory elements in instruction set extensions
abstract
Automatic generation of Instruction Set Extensions (ISEs), to be executed on a custom processing unit or a coprocessor is an important step towards processor customization. A typical goal of a manual designer is to combine a large number of atomic instructions into an ISE satisfying microarchitectural constraints. However, memory operations pose a challenge for previous ISE approaches by limiting the size of the resulting instruction. In this paper, we introduce memory elements into custom units which result in ISEs closer to those sought after by the designers. We consider two kinds of memory elements for mapping to the specialized hardware: small hardware tables and architecturally-visible state registers. We devised a genetic algorithm to specifically exploit opportunities of introducing memory elements during ISE generation. Finally, we demonstrate the effectiveness of our approach by a detailed study of the variation in performance, area and energy in the presence of the generated ISEs, on a number of MediaBench, EEMBC and cryptographic applications. With the introduction of memory, the average speedup varied from 2.7X to 5X depending on the architectural configuration with a nominal area overhead. Moreover, we obtained an average energy reduction of 26% with respect to a 32-KB cache.
Partha Biswas, Vinay Choudhary, Kubilay Atasu, Laura Pozzi 0001, Paolo Ienne, Nikil Dutt
DAC6
2004 Proxy-based task partitioning of watermarking algorithms for reducing energy consumption in mobile devices
abstract
Digital watermarking is a process that embeds an imperceptible signature or watermark in a digital file containing audio, image, text or video data. The watermark is later used to authenticate the data file and for tamper detection. It is particularly valuable in the use and exchange of digital media such as audio and video on emerging handheld devices. However, watermarking is computationally expensive and adds to the drain of the available energy in handheld devices. We present an approach in which we partition the watermarking embedding and extraction algorithms and migrate some tasks to a proxy server. This leads to a lower energy consumption on the handheld without compromising the security of the watermarking process. Our results show that executing watermarking partitioned between the proxy and the handheld reduces the total energy consumed by 80% over running it only on the handheld and improves performance by over two orders of magnitude.
Arun Kejariwal, Alexandru Nicolau, Nikil Dutt, Rajesh K. Gupta 0001
DAC4
2004 Extending the transaction level modeling approach for fast communication architecture exploration
abstract
System-on-Chip (SoC) designs are increasingly becoming more complex. Efficient on-chip communication architectures are critical for achieving desired performance in these systems. System designers typically use Bus Cycle Accurate (BCA) models written in high level languages such as C/C++ to explore the communication design space. These models capture all of the bus signals and strictly maintain cycle accuracy, which is useful for reliable performance exploration but results in slow simulation speeds for complex designs, even when they are modeled using high level languages. Recently there have been several efforts to use the Transaction Level Modeling (TLM) paradigm for improving simulation performance in BCA models. However these BCA models capture a lot of details that can be eliminated when exploring communication architectures.In this paper we extend the TLM approach and propose a new and faster transaction-based modeling abstraction level (CCATB) to explore the communication design space. Our abstraction level bridges the gap between the TLM and BCA levels, and yields an average performance speedup of 55 over BCA models. We demonstrate how fast and accurate exploration of tradeoffs is possible for high-performance shared bus architectures such as AMBA 2.0 and AMBA 3.0 (AXI) in industrial strength designs at the proposed abstraction level.
Sudeep Pasricha, Nikil Dutt, Mohamed Ben-Romdhane
DAC2
2004 Energy-Aware System Design for Wireless Multimedia
abstract
In this paper, we present various challenges that arise in the delivery and exchange of multimedia information to mobile devices. Specifically, we focus on techniques for maintaining QoS to end-user multimedia applications (e.g. video streaming, multimedia conferencing) while maximizing device lifetimes. In order to cope with the resource intensive nature of multimedia applications (in terms of computation, bandwidth and consequently power) and dynamic congestion levels in wireless networks, an end-to-end approach to QoS-aware power optimization is required. We discuss the trend towards such an integrated approach that couples the architectural, OS, middleware and application layers to achieve both user experience and device energy gains. We conclude with a discussion of tools for integrated system design and testing that will aid in rapid deployment of wireless multimedia.
Hans Van Antwerpen, Nikil Dutt, Rajesh K. Gupta 0001, Shivajit Mohapatra, Cristiano Pereira, Nalini Venkatasubramanian, Ralph von Vignau
DATE2
2004 Network Topology Exploration of Mesh-Based Coarse-Grain Reconfigurable Architectures
abstract
Several coarse-grain reconfigurable architectures proposed recently consist of a large number of processing elements (PEs) connected in a mesh-like network topology. We study the effects of three aspects of network topology exploration on the performance of applications on these architectures: (a) changing the interconnection between PEs; (b) changing the way the network topology is traversed while mapping operations to the PEs; and (c) changing the communication delays on the interconnects between PEs. We propose network topology traversal strategies that first schedule PEs that are spatially close and that have more interconnections among them. We use an interconnect aware list scheduling heuristic as a vehicle to perform the network topology exploration experiments on a set of designs derived from DSP applications. Our experimental results show that a spiral traversal strategy, coupled with a two neighbor interconnect topology leads to good performance for the DSP benchmarks considered. Our prototype framework thus provides an exploration environment for system architects to explore and tune coarse-grain reconfigurable architectures for particular application domains.
Nikhil Bansal 0003, Nikil Dutt, Alexandru Nicolau, Rajesh K. Gupta 0001
DATE3
2004 Automatic Tuning of Two-Level Caches to Embedded Applications
abstract
The power consumed by the memory hierarchy of a microprocessor can contribute to as much as 50% of the total microprocessor system power, and is thus a good candidate for optimizations. We present an automated method for tuning two-level caches to embedded applications for reduced energy consumption. The method is applicable to both a simulation-based exploration environment and a hardware-based system prototyping environment. We introduce the two-level cache tuner, or TCaT - a heuristic for searching the huge solution space of possible configurations. The heuristic interlaces the exploration of the two cache levels and searches the various cache parameters in a specific order based on their impact on energy. We show the integrity of our heuristic across multiple memory configurations and even in the presence of hardware/software partitioning - a common optimization capable of achieving significant speedups and/or reduced energy consumption. We apply our exploration heuristic to a large set of embedded applications. Our experiments demonstrate the efficacy of our heuristic: on average the heuristic examines only 7% of the possible cache configurations, but results in cache sub-system energy savings of 53%, only 1% more than the optimal cache configuration. In addition, the configured cache achieves an average speedup of 30% over the base cache configuration due to tuning of cache line size to the application's needs.
Ann Gordon-Ross, Frank Vahid, Nikil Dutt
DATE3
2004 Loop Shifting and Compaction for the High-Level Synthesis of Designs with Complex Control Flow
abstract
Emerging embedded system applications in multimedia and image processing are characterized by complex control flow consisting of deeply nested conditionals and loops. We present a technique called loop shifting that incrementally exploits loop level parallelism across iterations by shifting and compacting operations across loop iterations. Our experimental results show that loop shifting is particularly effective for the synthesis of designs with complex control especially when resource utilization is already high and/or under tight resource constraints. In situations when further loop unrolling (or initiating another iteration of the loop body) leads to a sharp increase in the longest combinational path in the circuit and the circuit area, loop shifting is able to achieve up to 20% reduction in the input-to-output delay in the synthesized circuit. We implemented loop shifting within the SPARK parallelizing high-level synthesis framework and present results for experiments on designs derived from multimedia and image processing applications.
Nikil Dutt, Rajesh K. Gupta 0001, Alexandru Nicolau
DATE2
2004 Data Reuse Analysis Technique for Software-Controlled Memory Hierarchies
abstract
In multimedia and other streaming applications a significant portion of energy is spent on data transfers. Exploiting data reuse opportunities in the application, we can reduce this energy by making copies of frequently used data in a small local memory and replacing speed and power inefficient transfers from main off-chip memory by more efficient local data transfers. In this paper we present an automated approach for analyzing these opportunities in a program that allows modification of the program to use custom scratch pad memory configurations comprising a hierarchical set of buffers for local storage of frequently reused data. Using our approach we are able to reduce energy consumption of the memory subsystem when using a scratch pad memory by a factor of two on average compare to a cache of the same size.
Ilya Issenin, Erik Brockmeyer, Miguel Corbalan, Nikil Dutt
DATE4
2004 Graph-Based Functional Test Program Generation for Pipelined Processors
abstract
Functional verification is widely acknowledged as a major bottleneck in microprocessor design. While early work on specification driven functional test program generation has proposed several promising ideas, many challenges remain in applying them to realistic embedded processors. We present a graph coverage based functional test program generation approach for pipelined processors. The proposed methodology makes three important contributions. First, it automatically generates the graph model of the pipelined processor from the specification using functional abstraction. Second, it generates functional test programs based on the coverage of the pipeline behaviour. Finally, the test generation time is drastically reduced due to the use of module level property checking. We applied this methodology on the DLX processor to demonstrate the usefulness of our approach.
Prabhat Mishra 0001, Nikil Dutt
DATE2
2004 Functional Validation of Programmable Architectures
abstract
Validation of programmable architectures, consisting of processor cores, coprocessors, and memory subsystems, is one of the major bottlenecks in current system-on-chip design methodology. A critical challenge in validation of such systems is the lack of a golden reference model. Traditional validation techniques employ different reference models depending on the abstraction level and verification task (e.g., functional simulation or property checking), resulting in potential inconsistencies between multiple reference models. This paper presents a validation methodology that uses an architecture description language (ADL) based specification as a golden reference model for validation of programmable architectures, and generation of executable models such as simulators and hardware prototypes. We present a validation framework that uses the generated hardware as a reference model to verify the hand-written implementation using a combination of symbolic simulation and equivalence checking. We also present functional coverage based test generation techniques for validation of pipelined processor architectures. Finally, the generated simulator and hardware models are also used for early exploration of programmable architectures.
Prabhat Mishra 0001, Nikil Dutt
DSD2
2004 Interconnect-Aware Mapping of Applications to Coarse-Grain Reconfigurable Architectures
Nikhil Bansal 0003, Nikil Dutt, Alexandru Nicolau, Rajesh K. Gupta 0001
FPL3
2004 FIFO power optimization for on-chip networks
abstract
As the design community moves towards architecting multiprocessor systems-on-chip (MPSoC), it is widely believed that an on-chip interconnection network is potentially the best candidate to satisfy the high aggregate throughput needed by dozens of IP blocks. In this context, power (energy) estimation and reduction techniques for switches and links, the core components of an interconnection network, gain added significance. FIFO buffers are a key component of a majority of network switches- buffers have been estimated to be the single largest power consumer for a typical switch in an on-chip network. In this report, we analyze energy-power characteristics of FIFOs for onchip networks and propose an optimization to reduce FIFO energy consumption in the context of an on-chip network. Our experimental results demonstrate promising reductions in energy consumptions (19-33 % for 256 and 512 bit wide links). Furthermore, our approach yields increasing
Sudarshan Banerjee, Nikil Dutt
ACM Great Lakes Symposium on VLSI2
2004 Using global code motions to improve the quality of results for high-level synthesis
abstract
The quality of synthesis results for most high-level synthesis approaches is strongly affected by the choice of control flow (through conditions and loops) in the input description. This leads to a need for high-level and compiler transformations that overcome the effects of programming style on the quality of generated circuits. To address this issue, we have developed a set of speculative code-motion transformations that enable movement of operations through, beyond, and into conditionals with the objective of maximizing performance. We have implemented these code transformations, along with supporting code-motion techniques and variable renaming techniques, in a high-level synthesis research framework called Spark. Spark takes a behavioral description in ANSI-C as input and generates synthesizable register-transfer level VHDL. We present results for experiments on designs derived from three real-life multimedia and image processing applications, namely, the MPEG-1 and -2 and GNU image manipulation program applications. We find that the speculative-code motions lead to reductions between 36% and 59% in the number of states in the finite-state machine (controller complexity) and the cycles on the longest path (performance) compared with the case when only nonspeculative code motions are employed. Also, logic synthesis results show fairly constant critical path lengths (clock period) and a marginal increase in area.
Nicolae Savoiu, Nikil Dutt, Rajesh K. Gupta 0001, Alexandru Nicolau
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2004 IDAP: a tool for high-level power estimation of custom array structures
abstract
While array structures are a significant source of power dissipation, there is a lack of accurate high-level power estimators that account for varying array circuit implementation styles. We present a methodology and a tool, the implementation-dependent array power (IDAP) estimator, that model power dissipation in SRAM-based arrays accurately based on a high-level description of the array. The models are parameterized by the array operations and various technology dependent parameters. The methodology is generic and the IDAP tool has been validated on industrial designs across a wide variety of array implementations in the e500 processor core (e500 is the Motorola processor core that is compliant with the PowerPC Book E architecture). For these industrial designs, IDAP generates high-level estimates for dynamic power dissipation that are accurate with an error margin of less than 22.2% of detailed (layout extracted) SPICE simulations. We apply the tool in three different scenarios: 1) identifying the subblocks that contribute to power significantly; 2) evaluating the effect of bitline-voltage swing on array power; and 3) evaluating the effect of memory bit-cell dimensions on array power.
Mahesh Mamidipaka, Kamal S. Khouri, Nikil Dutt, Magdy S. Abadir
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2004 Modeling and validation of pipeline specifications
abstract
Verification is one of the most complex and expensive tasks in the current Systems-on-Chip design process. Many existing approaches employ a bottom-up approach to pipeline validation, where the functionality of an existing pipelined processor is, in essence, reverse-engineered from its RT-level implementation. Our validation technique is complementary to these bottom-up approaches. Our approach leverages the system architect's knowledge about the behavior of the pipelined architecture, through architecture description language (ADL) constructs, and thus allows a powerful top-down approach to pipeline validation. The most important requirement in top-down validation process is to ensure that the specification (reference model) is golden. This paper addresses automatic validation of processor, memory, and coprocessor pipelines described in an ADL. We present a graph-based modeling that captures both structure and behavior of the architecture. Based on this model, we present algorithms to ensure that the static behavior of the pipeline is well formed by analyzing the structural aspects of the specification. We applied our methodology to verify specification of several realistic architectures from different architectural domains to demonstrate the usefulness of our approach.
Prabhat Mishra 0001, Nikil Dutt
ACM Trans. Embed. Comput. Syst.2
2004 Processor-memory coexploration using an architecture description language
abstract
Memory represents a major bottleneck in modern embedded systems in terms of cost, power, and performance. Traditionally, memory organizations for programmable embedded systems assume a fixed cache hierarchy. With the widening processor--memory gap, more aggressive memory technologies and organizations have appeared, allowing customization of a heterogeneous memory architecture tuned for specific target applications. However, such a processor--memory coexploration approach critically needs the ability to explicitly capture heterogeneous memory architectures. We present in this paper a language-based approach to explicitly capture the memory subsystem configuration, generate a memory-aware software toolkit, and perform coexploration of the processor--memory architectures. We present a set of experiments using our memory-aware architectural description language (ADL) to drive the exploration of the memory subsystem for the TI C6211 processor architecture, demonstrating cost, performance, and energy trade-offs.
Prabhat Mishra 0001, Mahesh Mamidipaka, Nikil Dutt
ACM Trans. Embed. Comput. Syst.3
2004 Coordinated parallelizing compiler optimizations and high-level synthesis
abstract
We present a high-level synthesis methodology that applies a coordinated set of coarse-grain and fine-grain parallelizing transformations. The transformations are applied both during a pre-synthesis phase and during scheduling, with the objective of optimizing the results of synthesis and reducing the impact of control flow constructs on the quality of results. We first apply a set of source level presynthesis transformations that include common sub-expression elimination (CSE), copy propagation, dead code elimination and loop-invariant code motion, along with more coarse-level code restructuring transformations such as loop unrolling. We then explore scheduling techniques that use a set of aggressive speculative code motions to maximally parallelize the design by re-ordering, speculating and sometimes even duplicating operations in the design. In particular, we present a new technique called "Dynamic CSE" that dynamically coordinates CSE and code motions such as speculation and conditional speculation during scheduling. We implemented our parallelizing high-level synthesis in the SPARK framework. This framework takes a behavioral description in ANSI-C as input and generates synthesizable register-transfer level VHDL. Our results from computationally expensive portions of three moderately complex design targets, namely, MPEG-1, MPEG-2 and the GIMP image processing tool, validate the utility of our approach to the behavioral synthesis of designs with complex control flows.
Rajesh K. Gupta 0001, Nikil Dutt, Alexandru Nicolau
ACM Trans. Design Autom. Electr. Syst.3
2003 Evaluating Memory Architectures for Media Applications on Coarse-Grained Recon.gurable Architectures
Jongeun Lee, Kiyoung Choi, Nikil Dutt
ASAP3
2003 Reducing code size for heterogeneous-connectivity-based VLIW DSPs through synthesis of instruction set extensions
abstract
VLIW DSP architectures exhibit heterogeneous connections between functional units and register files for speeding up special tasks. Such architectural characteristics can be effectively exploited through the use of complex instruction set extensions (ISEs). Although VLIWs are increasingly being used for DSP applications to achieve very high performance, such architectures are known to suffer from increased code size. This paper addresses how to generate ISEs that can result in significant code size reduction in VLIW DSPs without degrading performance. Unfortunately, contemporary techniques for instruction set synthesis fail to extract legal ISEs for heterogeneous-connectivity-based architectures. We propose a Heuristic-based algorithm to synthesize ISEs for a generalized heterogeneous-connectivity-based VLIW DSP architecture. We achieve an average code size reduction of 25% on the MiBench suite with no penalty in performance by applying our ISE generation algorithm on the TI TMS320C6xx, a representative VLIW DSP.
Partha Biswas, Nikil Dutt
CASES2
2003 Instruction set compiled simulation: a technique for fast and flexible instruction set simulation
abstract
Instruction set simulators are critical tools for the exploration and validation of new programmable architectures. Due to increasing complexity of the architectures and time-to-market pressure, performance is the most important feature of an instruction-set simulator. Interpretive simulators are flexible but slow, whereas compiled simulators deliver speed at the cost of flexibility. This paper presents a novel technique for generation of fast instruction set simulators that combines the benefit of both compiled and interpretive simulation. We achieve fast instruction accurate simulation through two mechanisms. First, we move the time consuming decoding process from run-time to compile time while maintaining the flexibility of the interpretive simulation. Second, we use a novel instruction abstraction technique to generate aggressively optimized decoded instructions that further improves simulation performance. Our instruction set compiled simulation (IS-CS) technique delivers upto 40% performance improvement over the best known published result that has the flexibility of interpretive simulation. We illustrate the applicability of the IS-CS technique using the ARM7 embedded processor.
Mehrdad Reshadi, Prabhat Mishra 0001, Nikil Dutt
DAC3
2003 Dynamic Conditional Branch Balancing during the High-Level Synthesis of Control-Intensive Designs
Nikil Dutt, Rajesh K. Gupta 0001, Alexandru Nicolau
DATE2
2003 On-chip Stack Based Memory Organization for Low Power Embedded Architectures
abstract
This paper presents an on-chip stack based memory organization that effectively reduces the energy dissipation in programmable embedded system architectures. Most embedded-systems use the notion of stack for implementation of function calls. However such stack data is stored in processor address space, typically in the main memory and accessed through caches. Our analysis of several benchmarks show that the callee saved registers and return addresses for function calls constitute a significant portion of the total memory accesses. We propose a separate stack-based memory organization to store these registers and return addresses. Our experimental results show that effective use of such stack-based memories yield significant reductions in system power/energy, while simultaneously improving the system performance. Application of our approach to the SPECint95 and MediaBench benchmark suites show up to 32.5% reduction in energy in L1 data caches, with marginal improvements in system performance.
Mahesh Mamidipaka, Nikil Dutt
DATE2
2003 IDAP: A Tool for High Level Power Estimation of Custom Array Structures
Mahesh Mamidipaka, Kamal S. Khouri, Nikil Dutt, Magdy S. Abadir
ICCAD3
2003 Interface Synthesis using Memory Mapping for an FPGA Platform
abstract
Several system-on-chip (SoC) platforms have recently emerged that use reconfigurable logic (FPGAs) as a programmable coprocessor to reduce the computational load on the main processor core. We present an interface synthesis approach that forms part of our hardware-software codesign methodology for such an FPGA-based platform. The approach is based on a novel memory mapping algorithm that maps data used by both the hardware and the software to shared memories on the reconfigurable fabric. The memory mapping algorithm couples with a high-level synthesis tool and uses scheduling information to map variables, arrays and complex data structures to the shared memories in a way that minimizes the number of registers and multiplexers used in the hardware interface. We also present three software schemes that enable the application software to communicate with this hardware interface. We demonstrate the utility of our approach and study the trade-offs involved using a case study of the codesign of a computationally expensive portion of the MPEG-1 multimedia application on to the Altera Nios platform.
Manev Luthra, Nikil Dutt, Rajesh K. Gupta 0001, Alexandru Nicolau
ICCD3
2003 Reducing Compilation Time Overhead in Compiled Simulators
abstract
Compiled simulation is a well known technique for improving the performance of instruction set simulators at the cost of compilation time. However the compilation time overhead makes such usage of compiler optimizations impractical especially for large applications. We propose a hybrid compiled simulation approach that is simple, generates an optimized decoder and has almost no compilation overhead comparing to static compiled simulation. Using two contemporary processor models- ARM7 and Sparc- we demonstrated that our technique can reduce the compilation time by 99% on the average, from several thousands of seconds to only tens of seconds.
Mehrdad Reshadi, Nikil Dutt
ICCD2
2003 Energy-efficient instruction set synthesis for application-specific processors
abstract
Several techniques have been proposed to enhance the energy-efficiency of ASIPs (Application-Specific Instruction set Processors). While those techniques can reduce the energy consumption with a minimal change in the instruction set (IS), they fail to exploit the opportunity of designing the entire IS from the energy-efficiency perspective. In this paper, we present an energy-efficient IS synthesis technique that can comprehensively reduce the energy-delay product (EDP) of ASIPs through optimal instruction encoding, considering both the instruction bitwidth and the dynamic instruction count. Experimental results with a typical embedded RISC processor show that our technique can generate application-specific IS's that are up to 40% more energy-efficient over the native IS for several application benchmarks.
Jongeun Lee, Kiyoung Choi, Nikil Dutt
ISLPED3
2003 An algorithm for mapping loops onto coarse-grained reconfigurable architectures
abstract
With the increasing demand for flexible yet highly efficient architecture platforms for media applications, there is a growing interest in the Coarse-grained Reconfigurable Architectures (CRAs). While many CRAs have demonstrated impressive performance improvement, the lack of compilation technology for such architectures causes a bottleneck in the current design process. In this paper, we present a novel mapping algorithm designed to support Reconfigurable ALU Array (RAA) architectures, that represent a significant class of CRAs. More specifically we present a core mapping algorithm that addresses the problem of placing and routing the operations of a loop body onto the ALU array, to be executed in a loop pipelined fashion. Experimental results using our mapping algorithm on a typical RAA show that our algorithm not only has very fast compilation time but can also generate quality mappings exhibiting high memory bandwidth utilization and low global interconnection requirements. Comparison with manual mapping also indicates that our algorithm can generate near-optimal mappings for several loops.
Jongeun Lee, Kiyoung Choi, Nikil Dutt
LCTES3
2003 Integrated power management for video streaming to mobile handheld devices
abstract
Optimizing user experience for streaming video applications on handheld devices is a significant research challenge. In this paper, we propose an integrated power management approach that unifies low level architectural optimizations (CPU, memory, register), OS power-saving mechanisms (Dynamic Voltage Scaling) and adaptive middleware techniques (admission control, optimal transcoding, network traffic regulation). Specifically, we identify interaction parameters between the different levels and optimize them to significantly reduce power consumption. With knowledge of device configurations, dynamic device parameters and changing system conditions, the middleware layer selects an appropriate video quality and fine tunes the architecture for optimized delivery of video. Our performance results indicate that architectural optimizations that are cognizant of user level parameters(e.g. transcoded video quality) can provide energy gains as high as 57.5% for the CPU and memory. Middleware adaptations to changing network noise levels can save as much as 70% of energy consumed by the wireless network interface. Furthermore, we demonstrate how such an integrated framework, that supports tight coupling of inter-level parameters can enhance user experience on a handheld substantially.
Shivajit Mohapatra, Radu Cornea, Nikil Dutt, Alexandru Nicolau, Nalini Venkatasubramanian
ACM Multimedia3
2003 Exploring Efficient Operating Points for Voltage Scaled Embedded Processor Cores
abstract
Portable and battery operated devices pose a unique design challenge in terms of performance requirements, low-power constraints, and short design cycles. Embedded soft cores, on the other hand, provide functional flexibility and guarantee rapid design and thus are gaining popularity in designing such portable and battery operated devices. To address the low power needs, dynamic voltage scaled (DVS) processors provide a new tradeoff dimension to the designer. This work proposes an application-specific design space exploration framework for selecting energy-efficient operating points in an embedded soft core. Specifically, we address the problem of selecting an appropriate number of operating voltage/frequency points and the distribution of these points along the valid voltage span of a processor, given the application that is to be executed on the processor. Furthermore, we provide a static intra-task scheduling technique that reduces energy consumption (4-20% in our experiments) even when the worst-case application execution time does not leave any slack for effective voltage scaling. We have experimentally verified our technologies on a large set of embedded benchmarks selected from MiBench, PowerStone, and MediaBench.
Marcio Buss, Tony Givargis, Nikil Dutt
RTSS3
2003 Access pattern-based memory and connectivity architecture exploration
abstract
Memory accesses represent a major bottleneck in embedded systems power and performance. Traditionally, designers tried to alleviate this problem by relying on a simple cache hierarchy, or a limited use of special purpose memory modules such as stream buffers. Although real-life applications contain a large number of memory references to a diverse set of data structures, a significant percentage of all memory accesses in the application are generated from a few memory instructions that exhibit predictable, well-known access patterns; this creates an opportunity for memory customization, targeting the needs of these access patterns. We present APEX, an approach that extracts, analyzes and clusters the most active access patterns in the application, and aggressively customizes the memory architecture to match the needs of the application. Moreover, though the memory modules are important, the rate at which the memory system can produce the data for the CPU is significantly impacted by the connectivity architecture between the memory subsystem and the CPU. Thus, it is critical to consider the connectivity architecture early in the design flow, in conjunction with the memory architecture. We couple the exploration of memory modules together with their connectivity, to evaluate a wide range of cost, performance, and energy connectivity architectures. We use a heuristic to prune the design space, guiding the exploration towards the most promising designs. We present experiments on a set of large real-life benchmarks, showing significant performance improvements for varied cost and power characteristics, allowing the designer to evaluate customized memory and connectivity configurations for embedded systems.
Peter Grun, Nikil Dutt, Alexandru Nicolau
ACM Trans. Embed. Comput. Syst.2
2003 RTGEN-an algorithm for automatic generation of reservation tables from architectural descriptions
abstract
Reservation Tables (RTs) have long been used to detect conflicts between operations that simultaneously access the same architectural resource. Traditionally, these RTs have been specified explicitly by the designer. However, the increasing complexity of modern processors makes the manual specification of RTs cumbersome and error prone. Furthermore, manual specification of such conflict information is infeasible for supporting rapid architectural exploration. In this paper, we present an algorithm to automatically generate RTs from a high-level processor description with the goal of avoiding manual specification of RTs, resulting in more concise architectural specifications and also supporting faster turnaround time in design space exploration. We demonstrate the utility of our approach on a set of experiments using the TI C6201 very long instruction word digital signal processor and DLX processor architectures, and a suite of multimedia and scientific applications.
Peter Grun, Ashok Halambi, Nikil Dutt, Alexandru Nicolau
IEEE Trans. Very Large Scale Integr. Syst.3
2003 Adaptive low-power address encoding techniques using self-organizing lists
abstract
Off-chip bus transitions are a major source of power dissipation for embedded systems. In this paper, new adaptive encoding schemes are proposed that significantly reduce transition activity on data and multiplexed address buses. These adaptive techniques are based on self-organizing lists to achieve reduction in transition activity by exploiting the spatial and temporal locality of the addresses. Also the proposed techniques do not require any extra bit lines and have minimal delay overhead. The techniques are evaluated for efficiency using a wide variety of application programs including SPEC 95 benchmark set. Unlike previous approaches that focus on instruction address buses, experiments demonstrate significant reduction in transition activity of up to 54% in data address buses and up to 59% in multiplexed address buses. The average reductions are twice those obtained using current schemes on a data address bus and more than twice those obtained on a multiplexed address bus.
Mahesh Mamidipaka, Daniel S. Hirschberg, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.3
2002 Coordinated transformations for high-level synthesis of high performance microprocessor blocks
abstract
High performance microprocessor designs are partially characterized by functional blocks consisting of a large number of operations that are packed into very few cycles (often single-cycle) with little or no resource constraints but tight bounds on the cycle time. Extreme parallelization, conditional and speculative execution of operations is essential to meet the processor performance goals. However, this is a tedious task for which classical high-level synthesis (HLS) formulations are inadequate and thus rarely used. In this paper, we present a new methodology for application of HLS targeted to such microprocessor functional blocks that can potentially speed up the design space exploration for microprocessor designs. Our methodology consists of a coordinated set of source-level and fine-grain parallelizing compiler transformations that targets these behavioral descriptions, specifically loop constructs in them and enables efficient chaining of operations and high-level synthesis of the functional blocks. As a case study in understanding the complexity and challenges in the use of HLS, we walk the reader through the detailed design of an instruction length decoder drawn from the Pentium-family of processors. The chief contribution of this paper is formulation of a domain-specific methodology for application of high-level synthesis techniques to a domain that rarely, if ever, finds use for it.
Nicolae Savoiu, Nikil Dutt, Rajesh K. Gupta 0001, Alexandru Nicolau, Timothy Kam, Michael Kishinevsky, Shai Rotem
DAC3
2002 Profile-Based Dynamic Voltage Scheduling Using Program Checkpoints
abstract
Dynamic voltage scaling (DVS) is a known effective mechanism for reducing CPU energy consumption without significant performance degradation. While a lot of work has been done on inter-task scheduling algorithms to implement DVS under operating system control, new research challenges exist in intra-task DVS techniques under software and compiler control. In this paper we introduce a novel intra-task DVS technique under compiler control using program checkpoints. Checkpoints are generated at compile time and indicate places in the code where the processor speed and voltage should be re-calculated. Checkpoints also carry user-defined time constraints. Our technique handles multiple intra-task performance deadlines and modulates power consumption according to a run-time power budget. We experimented with two heuristics for adjusting the clock frequency and voltage. For the particular benchmark studied, one heuristic yielded 63% more energy savings than the other. With the best of the heuristics we designed, our technique resulted in 82% energy savings over the execution of the program without employing DVS.
Ilya Issenin, Radu Cornea, Rajesh K. Gupta 0001, Nikil Dutt, Alexander V. Veidenbaum, Alexandru Nicolau
DATE5
2002 Memory System Connectivity Exploration
abstract
In programmable embedded systems, the memory subsystem represents a major cost, performance and power bottleneck. To optimize the system for such different goals, the designer would like to perform Design Space Exploration, evaluating different memory modules from a memory IP library, and selecting the most promising designs. However while the memory modules are important, the rate at which the memory system can produce the data for the CPU is significantly impacted by the connectivity architecture between the memory subsystem and the CPU. Thus, it is critical go consider the connectivity architecture early in the design flow, in conjunction with the memory architecture. We present a connectivity architecture exploration approach, evaluating a wide range of cost, performance, and energy connectivity architectures. When coupled with our memory modules exploration approach, we can significantly improve the system behavior We present experiments on a set of large real-life benchmarks, showing significant performance improvements for varied cost and power characteristics, allowing the designer to tailor the performance, cost and power of the programmable embedded system.
Peter Grun, Nikil Dutt, Alexandru Nicolau
DATE2
2002 An Efficient Compiler Technique for Code Size Reduction Using Reduced Bit-Width ISAs
abstract
For many embedded applications, program code size is a critical design factor. One promising approach for reducing code size is to employ a "dual instruction set", where processor architectures support a normal (usually 32 bit) Instruction Set, and a narrow, space-efficient (usually 16 bit) Instruction Set with a limited set of opcodes and access to a limited set of registers. This future, however, requires compilers that can reduce code size by compiling for both Instruction Sets. Existing compiler techniques operate at the function-level granularity and are unable to make the trade-off between increased register pressure (resulting in more spills) and decreased code size. We present a profitability based compiler heuristic that operates at the instruction-level granularity and is able to effectively take advantage: of both Instruction Sets. We also demonstrate improved code size reduction, for the MIPS 32/16 bit ISA, using our technique. Our approach more than doubles the code size reduction achieved by existing compilers.
Ashok Halambi, Aviral Shrivastava, Partha Biswas, Nikil Dutt, Alexandru Nicolau
DATE4
2002 Automatic Verification of In-Order Execution In Microprocessors with Fragmented Pipelines and Multicycle Functional Units
abstract
As embedded systems continue to face increasingly higher performance requirements, deeply pipelined processor architectures are being employed to meet desired system performance. System architects critically need modeling techniques that allow exploration, evaluation, customization and validation of different processor pipeline configurations, tuned for a specific application domain. We propose a novel finite state machine (FSM) based modeling of pipelined processors and define a set of properties that can be used to verify the correctness of in-order execution in the presence of fragmented pipelines and multicycle functional units. Our approach leverages the system architect's knowledge about the behavior of the pipelined processor through architecture description language (ADL) constructs, and thus allows a powerful top-down approach to pipeline verification. We applied this methodology to the DLX processor to demonstrate the usefulness of our approach.
Prabhat Mishra 0001, Nikil Dutt, Alexandru Nicolau, Hiroyuki Tomiyama
DATE2
2002 Memory Architectures for Embedded Systems-On-Chip
Preeti Ranjan Panda, Nikil Dutt
HiPC2
2002 Efficient instruction encoding for automatic instruction set design of configurable ASIPs
abstract
Application-specific instructions can significantly improve the performance, energy, and code size of configurable processors. A common approach used in the design of such instructions is to convert application-specific operation patterns into new complex instructions. However, processors with a fixed instruction bitwidth cannot accommodate all the potentially interesting operation patterns, due to the limited code space afforded by the fixed instruction bitwidth. We present a novel instruction set synthesis technique that employs an efficient instruction encoding method to achieve maximal performance improvement. We build a library of complex instructions with various encoding alternatives and select the best set of complex instructions while satisfying the instruction bitwidth constraint. We formulate the problem using integer linear programming and also present an effective heuristic algorithm. Experimental results using our technique generate instruction sets that show improvements of up to 38% over the native instruction set for several realistic benchmark applications running on a typical embedded RISC processor.
Jongeun Lee, Kiyoung Choi, Nikil Dutt
ICCAD3
2001 New directions in compiler technology for embedded systems (embedded tutorial)
abstract
Traditionally, compiler technology has focused on the generation of code with the goal of improving performance for a variety of applications running on general-purpose processor architectures. In the embedded system space, compiler technology is faced with many new challenges, including: code generation for specialized architectural features, requireing a highly flexible degree of retargetability; memory-aware code generation that exploits the timing and structure of the embedded system's memory organization; optimizing software to meet both real-time and performance constraints; energy- and power-aware software generation, both from the context of energy minimization, as well as power modulation; code size minimization for memory-constrained embedded systems; coarse-grain transformations for tightly-coupled, memory-constrained multi-processor architectures; and interaction with the operating system for active management of embedded system resources. This paper discusses new directions for compiler technology, surveys some of the current research efforts and illustrates proposed solutions to selected issues.
Nikil Dutt, Alexandru Nicolau, Hiroyuki Tomiyama, Ashok Halambi
ASP-DAC1
2001 Speculation Techniques for High Level Synthesis of Control Intensive Designs
abstract
The quality of synthesis results for most high level synthesis approaches is strongly affected by the choice of control flow (through conditions and loops) in the input description. In this paper, we explore the effectiveness of various types of code motions, such as moving operations across conditionals, out of conditionals (speculation) and into conditionals (reverse speculation), and how they can be effectively directed by heuristics so as to lead to improved synthesis results in terms of fewer execution cycles and fewer number of states in the finite state machine controller. We also study the effects of the code motions on the area and latency of the final synthesized netlist. Based on speculative code motions, we present a novel way to perform early condition execution that leads to significant improvements in highly control-intensive designs. Overall, reductions of up to 38 \% in execution cycles are obtained with all the code motions enabled.
Nicolae Savoiu, Nikil Dutt, Rajesh K. Gupta 0001, Alexandru Nicolau
DAC4
2001 Access pattern based local memory customization for low power embedded systems
abstract
Memory accesses represent a major bottleneck in embedded systems power and performance. Traditionally, the local memory relied on a large cache to store all the variables in the application. However, especially in large real-life applications, different types of data exhibit divergent types of locality and access patterns, with diverse locality and bandwidth needs. Traditional caches had to compromise between the different types of locality required by the access patterns, and trade-off performance against bandwidth requirement. Instead, our approach customizes the local memory architecture matching the diverse access patterns and locality types present in the application, to reduce the main memory bandwidth requirement, and significantly improve power consumption, without sacrificing performance. Our approach generated an average 30% memory power reduction without degrading performance on a set of large multimedia/general purpose applications and scientific kernels, over the best traditional cache configuration of similar size, demonstrating the utility of our algorithm.
Peter Grun, Nikil Dutt, Alexandru Nicolau
DATE2
2001 Low power address encoding using self-organizing lists
abstract
Off-chip bus transitions are a major source of power dissipation for embedded systems. In this paper, new adaptive encoding schemes are proposed that significantly reduce transition activity on data and multiplexed address buses, that do not add redundancy in space or time and which have minimal delay overhead. These adaptive techniques are based on self-organising lists to achieve reduction in transition activity by exploiting the spatial and temporal locality of the addresses. Unlike previous approaches that focus on instruction address buses, experiments demonstrate significant reduction in transition activity of up to 54% in data address buses and up to 59% in multiplexed address buses. The average reductions are twice those obtained using current schemes on a data address bus and more than twice those obtained on a multiplexed address bus. 1.
Mahesh Mamidipaka, Daniel S. Hirschberg, Nikil Dutt
ISLPED3
2001 V-SAT: A visual specification and analysis tool for system-on-chip exploration
Asheesh Khare, Ashok Halambi, Nicolae Savoiu, Peter Grun, Nikil Dutt, Alexandru Nicolau
J. Syst. Archit.5
2001 Data and memory optimization techniques for embedded systems
abstract
We present a survey of the state-of-the-art techniques used in performing data and memory-related optimizations in embedded systems. The optimizations are targeted directly or indirectly at the memory subsystem, and impact one or more out of three important cost metrics: area, performance, and power dissipation of the resulting implementation. We first examine architecture-independent optimizations in the form of code transoformations. We next cover a broad spectrum of optimization techniques that address memory architectures at varying levels of granularity, ranging from register files to on-chip memory, data caches, and dynamic memory (DRAM). We end with memory addressing related issues.
Preeti Ranjan Panda, Francky Catthoor, Nikil Dutt, Koen Danckaert, Erik Brockmeyer, Chidamber Kulkarni, Arnout Vandecappelle, Per Gunnar Kjeldsberg
ACM Trans. Design Autom. Electr. Syst.3
2000 Memory aware compilation through accurate timing extraction
abstract
Memory delays represent a major bottleneck in embedded systems performance. Newer memory modules exhibiting efficient access modes (e.g., page-, burst-mode) partly alleviate this bottleneck. However, such features can not be efficiently exploited in processor-based embedded systems without memory-aware compiler support. We describe a memory-aware compiler approach that exploits such efficient memory access modes by extracting accurate timing information, allowing the compiler's scheduler to perform global code reordering to better hide the latency of memory operations. Our memory-aware compiler scheduled several benchmarks on the TI C6201 processor architecture interfaced with a 2-bank synchronous DRAM and generated average improvements of 24% over the best possible schedule using a traditional (memory-transparent) optimizing compiler, demonstrating the utility of our memory-aware compilation approach.
Peter Grun, Nikil Dutt, Alexandru Nicolau
DAC2
2000 How to Solve the Current Memory Access and Data Transfer Bottlenecks: At the Processor Architecture or at the Compiler Level?
abstract
Current processor architectures, both in the programmable and custom case, become more and more dominated by the data access bottlenecks in the cache, system bus and main memory subsystems. In order to provide sufficiently high data throughput in the emerging era of highly parallel processors where many arithmetic resources can work concurrently, novel solutions for the memory access and data transfer will have to be introduced. The crucial question we want to address is where one can expect these novel solutions to reside: will they be mainly innovative processor architecture ideas, or novel approaches in the application compiler/synthesis technology, or a mix.
Francky Catthoor, Nikil Dutt, Christoforos E. Kozyrakis
DATE2
2000 Architecture Exploration of Parameterizable EPIC SOC Architectures
abstract
Design Space Exploration (DSE) of programmable systems-on-chip (SOC) incorporating parameterizable processor cores is difficult due to the complex and intrinsically nonstructured interactions between different architectural features of the processor (such as wide parallelism, and deep pipelines), the compiler and the application. Changing different processor features implies generating detailed operation conflict information - represented as Reservation Tables (RTs). If done manually, it can be a very tedious and error prone task, especially for deep pipelines, with complex resource sharing and large nonstructured instruction sets. In this paper we use RTGEN, an approach for automatic generation of RTs, to drive rapid architectural exploration of a large number of designs. We present exploration experiments on a large set of VLIW-like EPIC architectures, for varying port sharing, number of functional units, multicycling units, and with varied latency configurations. Our experiments uncovered several non-intuitive architecture design points, giving the system-level designer further flexibility in exploration of programmable SOC architectures.
Ashok Halambi, Radu Cornea, Peter Grun, Nikil Dutt, Alexandru Nicolau
DATE4
2000 MIST: An Algorithm for Memory Miss Traffic Management
abstract
Cache misses represent a major bottleneck in embedded systems performance. Traditionally, compilers optimistically treated all memory accesses as cache hits, relying on the memory controller to account for longer miss delays. However, the memory controller has only a local view of the program, and is not able to efficiently hide the latency of these memory operations. Our compiler technique actively manages cache misses, and performs global miss traffic optimizations, to better hide the latency of the memory operations. Our memory-aware compiler scheduled several benchmarks on the TIC6211 processor architecture with a direct mapped cache, and generated an average of 61.6% improvement over the best schedule of the traditional (memory-transparent) optimizing compiler, demonstrating the utility of our miss traffic optimization approach.
Peter Grun, Nikil Dutt, Alexandru Nicolau
ICCAD2
2000 System and Architecture-Level Power Reduction for Microprocessor-Based Communication and Multi-Media Applications
abstract
Current microprocessor architectures become more and more dominated by the data access bottlenecks in the cache, system bus and main memory subsystems. These also have a major influence on the system (board-level) power consumption. In practice this means lower energy consumption for a given throughput requirement. In the booming domain of (largely embedded) cost-sensitive communication and multi-media applications, more and more implementations make use of microprocessor based platforms for flexibility reasons. However, in order to provide sufficiently high data throughput at reasonable power consumption for these demanding applications, novel solutions for the memory access and data transfer will have to be introduced. These will have to be situated both at the processor architecture and the algorithm/compiler level. The question we want to address in this paper is what would these solutions look like. We show that they will be based on processor architecture optimizations, on novel approaches in the application of compiler technology, and on exploiting the interface between the system hardware and software.
Lode Nachtergaele, Vivek Tiwari, Nikil Dutt
ICCAD3
2000 High-level library mapping for memories
abstract
We present high-level library mapping, a technique that synthesizes a source memory module from a library of target memory modules. In this paper, we define the problem of high-level library mapping for memories, identify and solve the three subproblems associated with this task, and finally combine these solutions into a suite of two memory mapping algorithms. Experimental results on a number of memory-intensive designs demonstrate that our memory mapping approach generates a wide variety of cost-effective designs, often counter-intuitive ones, based on a user-given cost function, the target library, and the mapping algorithm used.
Pradip K. Jha, Nikil Dutt
ACM Trans. Design Autom. Electr. Syst.2
2000 On-chip vs. off-chip memory: the data partitioning problem in embedded processor-based systems
abstract
Efficient utilization of on-chip memory space is extremely important in modern embedded system applications based on processor cores. In addition to a data cache that interfaces with slower off-chip memory, a fast on-chip SRAM, called Scratch-Pad memory, is often used in several applications, so that critical data can be stored there with a guaranteed fast access time. We present a technique for efficiently exploiting on-chip Scratch-Pad memory by partitioning the application's scalar and arrayed variables into off-chip DRAM and on-chip Scratch-Pad SRAM, with the goal of minimizing the total execution time of embedded applications. We also present extensions of our proposed memory assignment strategy to handle context switching between multiple programs, as well as a generalized memory hierarchy. Our experiments on code kernels from typical applications show that our technique results in significant performance improvements.
Preeti Ranjan Panda, Nikil Dutt, Alexandru Nicolau
ACM Trans. Design Autom. Electr. Syst.2
2000 Guest editorial 11th international symposium on system-level synthesis and design (ISSS'98)
Allen C.-H. Wu, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.2
1999 EXPRESSION: A Language for Architecture Exploration through Compiler/Simulator Retargetability
abstract
We describe EXPRESSION, a language supporting architectural design space exploration for embedded systems-on-chip (SOC) and automatic generation of a retargetable compiler/simulator toolkit. Key features of our language-driven design methodology include: a mixed behavioral/structural representation supporting a natural specification of the architecture, explicit specification of the memory, subsystem allowing novel memory organizations and hierarchies; clean syntax and ease of modification supporting architectural exploration; a single specification supporting consistency and completeness checking of the architecture; and efficient specification of architectural resource constraints allowing extraction of detailed reservation tables for compiler scheduling. We illustrate key features of EXPRESSION through simple examples and demonstrate its efficacy in supporting exploration and automatic software toolkit generation for an embedded SOC codesign flow.
Ashok Halambi, Peter Grun, Vijay Ganesh 0001, Asheesh Khare, Nikil Dutt, Alexandru Nicolau
DATE5
1999 Design of a set-top box system on a chip (abstract)
Nikil Dutt, Eric M. Foster
ICCAD1
1999 On the rapid prototyping and design of a wireless communication system on a chip (abstract)
Nikil Dutt, Brian Kelley
ICCAD1
1999 Augmenting Loop Tiling with Data Alignment for Improved Cache Performance
abstract
Loop blocking (tiling) is a well-known compiler optimization that helps improve cache performance by dividing the loop iteration space into smaller blocks (tiles); reuse of array elements within each tile is maximized by ensuring that the working set for the tile fits into the data cache. Padding is a data alignment technique that involves the insertion of dummy elements into a data structure for improving cache performance. In this work, we present DAT, a technique that augments loop tiling with data alignment, achieving improved efficiency (by ensuring that the cache is never under-utilized) as well as improved flexibility (by eliminating self-interference cache conflicts independent of the tile size). This results in a more stable and better cache performance than existing approaches, in addition to maximizing cache utilization, eliminating self-interference, and minimizing cross-interference conflicts. Further, while all previous efforts are targeted at programs characterized by the reuse of a single array, we also address the issue of minimizing conflict misses when several tiled arrays are involved. To validate our technique, we ran extensive experiments using both simulations as well as actual measurements on SUN Sparc5 and Sparc10 workstations. The results on benchmarks exhibiting varying memory access patterns demonstrate the effectiveness of our technique through consistently high hit ratios and improved performance across varying problem sizes.
Preeti Ranjan Panda, Hiroshi Nakamura, Nikil Dutt, Alexandru Nicolau
IEEE Trans. Computers3
1999 Local memory exploration and optimization in embedded systems
abstract
Embedded processor-based systems allow for the tailoring of the on-chip memory architecture based on application specific requirements. We present an analytical strategy for exploring the on-chip memory architecture for a given application, based on a memory performance estimation scheme. The analytical technique has the important advantage of enabling a fast evaluation of candidate memory architectures in the early stages of system design. Many digital signal-processing applications involve array accesses and loop nests that can benefit from such an exploration. Our experiments demonstrate that our estimations closely follow the actual simulated performance at significantly reduced run times.
Preeti Ranjan Panda, Nikil Dutt, Alexandru Nicolau
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1999 Low-power memory mapping through reducing address bus activity
abstract
Arrays in behavioral specifications that are too large to fit into on-chip registers are usually mapped to off-chip memories during behavioral synthesis. We address the problem of system power reduction through transition count minimization on the memory address bus when these arrays are accessed from memory. We exploit regularity and spatial locality in the memory accesses and determine the mapping of behavioral array references to physical memory locations to minimize address bus transitions. We describe array mapping strategies for two important memory configurations: all behavioral arrays mapped to a single off-chip memory and arrays mapped into multiple memory modules drawn from a library. For the single memory configuration, we describe a heuristic for selecting a memory mapping scheme to achieve low power for each behavioral array. For mapping into a library of multiple memory modules, we formulate the problem as three logical-to-physical memory mapping subtasks and present experiments demonstrating the transition count reductions based on our approach. Our experiments on several image processing benchmarks show power savings of up to 63% through reduced transition activity on the memory address bus in the single memory case. We also observe a further transition count reduction by a factor of 1.5-6.7 over a straightforward mapping scheme in the multiple memories configuration.
Preeti Ranjan Panda, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.2
1998 Data Cache Sizing for Embedded Processor Applications
abstract
We present a technique for determining the best data cache size required for a given memory-intensive application. A careful memory and cache line assignment strategy based on the analysis of the array access patterns effects a significant reduction in the required data cache size, with no negative impact on the performance, thereby freeing vital on-chip silicon area for other hardware resources. Experiments on several benchmark kernels performed on LSI Logic's CW4001 embedded processor simulator confirm the soundness of our cache sizing and memory assignment strategy and the accuracy of our analytical predictions.
Preeti Ranjan Panda, Nikil Dutt, Alexandru Nicolau
DATE2
1998 Embedded memories in system design - from technology to systems architecture
abstract
No abstract available.
Soren Hein, Vijay Nagasamy, Bernhard Rohfleisch, Christoforos E. Kozyrakis, Nikil Dutt, Francky Catthoor
ICCAD5
1998 Incorporating DRAM access modes into high-level synthesis
abstract
Memory-intensive behaviors often contain large arrays that are synthesized into off-chip memories. With the increasing gap between on-chip and off-chip memory access delays, it is imperative to exploit the efficient access mode features of modern-day memories (e.g., page-mode DRAM's) in order to alleviate the memory bandwidth bottleneck. Although recent research efforts in high-level synthesis (HLS) have addressed the issue of memory-based synthesis, current techniques are unable to exploit efficiently the special access modes of these off-chip memories, resulting in significantly inferior performance using these memory library parts. Our work addresses this issue by (a) modeling realistic off-chip memory access modes for HLS, (b) presenting algorithms to infer applicability of HLS with these memory access modes, and (c) transforming input behavior to provide further memory access optimizations during HLS. We demonstrate the utility of our approach using a suite of memory-intensive benchmarks with a realistic DRAM library module. Experimental results show a significant performance improvement (more than 40%) as a result of our optimization techniques.
Preeti Ranjan Panda, Nikil Dutt, Alexandru Nicolau
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1997 Exploiting off-chip memory access modes in high-level synthesis
abstract
Memory-intensive behaviors often contain large arrays that are synthesized into off-chip memories. With the increasing gap between on-chip and off-chip memory access delays, it is imperative to exploit the efficient access mode features of modern-day memories (e.g. page-mode DRAMs) in order to alleviate the memory bandwidth bottleneck. Our work addresses this issue by: (a) modeling realistic off-chip memory access modes for High-level Synthesis (HLS), (b) presenting algorithms to infer applicability of HLS with these memory access modes, and (c) transforming input behavior to provide further memory access optimizations during HLS. We demonstrate the utility of our approach using a suite of memory-intensive benchmarks with a realistic DRAM library module. Experimental results show a significant performance improvement (more than 40%) as a result of our optimization techniques.
Preeti Ranjan Panda, Nikil Dutt, Alexandru Nicolau
ICCAD2
1997 A Data Alignment Technique for Improving Cache Performance
abstract
We address the problem of improving the data cache performance of numerical applications-specifically, those with blocked (or tiled) loops. We present DAT, a data alignment technique utilizing array-padding, to improve program performance through minimizing cache conflict misses. We describe algorithms for selecting tile sizes for maximizing data cache utilization, and computing pad sizes for eliminating self-interference conflicts in the chosen tile. We also present a generalization of the technique to handle applications with several tiled arrays. Our experimental results comparing our technique with previous published approaches on machines with different cache configurations show consistently good performance on several benchmark programs, for a variety of problem sizes.
Preeti Ranjan Panda, Hiroshi Nakamura, Nikil Dutt, Alexandru Nicolau
ICCD3
1997 A unified lower bound estimation technique for high-level synthesis
abstract
The importance of effective lower bound estimation (LBE) techniques is well established in high-level synthesis (HLS) since it allows more efficient exploration of the design space while providing other HLS tools with the capability of predicting the effect of specific tools on the design space. Much of the previous work has focused on LBE techniques that use very simple cost models which primarily focus on the functional unit resources. With the push toward submicron technologies, simple models that use functional unit resources alone are not accurate enough to allow effective design space exploration since the effects of storage and interconnect can indeed dominate the cost function. In this paper, we present an integrated approach aimed at predicting lower bounds on hardware resources needed to implement a behavioral description within a given amount of time. Our area cost model accounts for storage (register) and interconnect resources (buses) in addition to functional resources. Our timing model uses a finer granularity that permits the modeling of functional unit, register, and interconnect delays. Our approach is integrated because we consider the dependencies between the different types of resources as well as the ordering in which the resources are allocated. We tested our technique for functional unit, storage, and interconnect requirements on several high-level synthesis benchmarks, and observed near-optimal results. We believe that our comprehensive LBE approach can lead to better quality HLS solutions in less time, and we demonstrate this approach in our paper.
Seong Yong Ohm, Fadi J. Kurdahi, Nikil Dutt
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1997 Memory data organization for improved cache performance in embedded processor applications
abstract
Code generation for embedded processors opens up the possibility for several performance optimization techniques that have been ignored by traditional compilers due to compilation time constraints. We present techniques that take into account the parameters of the data caches for organizing scalar and array variables declared in embedded code into memory, with the objective of improving data cache performance. We present techniques for clustering variables to minimize compulsory cache misses, and for solving the memory assignment problem to minimize conflict cache misses. Our experiments with benchmark code kernels from DSP and other domains on the CW4001 embedded processor from LSI Logic indicate significant improvements in data cache performance by the application of our memory organization technique.
Preeti Ranjan Panda, Nikil Dutt, Alexandru Nicolau
ACM Trans. Design Autom. Electr. Syst.2
1996 Low-power mapping of behavioral arrays to multiple memories
abstract
Large data arrays in behavioral specifications are usually mapped to off-chip memories during system synthesis. We address the problem of system power reduction through transition count minimization on the address bus during memory accesses, when mapping behavioral arrays to multiple memory modules drawn from a library. We formulate the problem as three logical-to-physical memory mapping subtask, provide algorithms for each subtask, and present experiments that demonstrate the transition count reductions based on our approach. Our experiments show a transition count reduction by a factor of 1.5-6.7 over a straightforward mapping scheme.
Preeti Ranjan Panda, Nikil Dutt
ISLPED2
1996 Elimination of redundant memory traffic in high-level synthesis
abstract
This paper presents a new transformation for the scheduling of memory-access operations in high-level synthesis. This transformation is suited to memory-intensive applications with synthesized designs containing a secondary store accessed by explicit instructions. Such memory-intensive behaviors are commonly observed in video compression, image convolution, hydrodynamics and mechatronics. Our transformation removes load and store instructions which become redundant or unnecessary during the transformation of loops. The advantage of this reduction is the decrease of secondary memory bandwidth demands. This technique is implemented in our Percolation-Based Scheduler which we used to conduct experiments on a suite of memory-intensive benchmarks. Our results demonstrate a significant reduction in the number of memory operations and an increase in performance on these benchmarks.
David J. Kolson, Alexandru Nicolau, Nikil Dutt
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1996 Optimal register assignment to loops for embedded code generation
abstract
One of the challenging tasks in code generation for embedded systems is register assignment. When more live variables than registers exist, some variables will necessarily be accessed from data memory. Because loops are typically executed many times and are often time-critical, good register assignment in loops is exceedingly important as accessing data memory can degrade performance. The issue of finding an optimal register assignment to loops has been open for some time. In this article, we present a technique for optimal (i.e., spill minimizing) register assignment to loops. First we present a technique for register assignment to architecture styles that are characterized by a consolidated register file. Then we extend the technique to include architecture styles that are characterized by distributed memories and/or a combination of general- and special-purpose registers. Experimental results demonstrate that although the optimal algorithm may be computationally prohibitive, heuristic versions obtain results with performance better than that of an existing graph coloring approach.
David J. Kolson, Alexandru Nicolau, Nikil Dutt, Ken Kennedy
ACM Trans. Design Autom. Electr. Syst.3
1996 High-level library mapping for arithmetic components
abstract
We describe high-level library mapping (HLLM), a technique that permits reuse of complex RT-level databook components (specifically ALUs). HLLM can be used to couple existing databook libraries, module generators and custom-designed components with the output of architectural or behavioral synthesis. In this paper, we define the problem of high-level library mapping, present some algorithmic formulations for HLLM of ALUs, and demonstrate the versatility of our approach on a variety of libraries. We also compare HLLM against the traditional mapping approach using logic synthesis. Our experiments show that HLLM for ALUs outperforms logic synthesis in area, delay, and runtime, indicating that HLLM is a promising approach for reuse of datapath components in architectural design and high-level synthesis.
Pradip K. Jha, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.2
1995 Reclocking for high-level synthesis
abstract
No abstract available.
Pradip K. Jha, Nikil Dutt, Sri Parameswaran
ASP-DAC2
1994 Design Reuse: Fact or Fiction? (Panel)
abstract
Design reuse is expected to be a key enabler for system designs in the 90's. Proponents claim that design reuse increases competitiveness through better quality, improved predictability and better productivity. On the other hand, skeptics say that design reuse is not realizable due to several barriers including rapid changes in technology, lack of standardized libraries and the presence of human, as opposed to technical barriers. Is design reuse a reality? How much is reused in practice? Where is it most applicable? Do concrete metrics exist for reusability? The panel will assess the track record of design reuse and discuss its status and future.
Nikil Dutt, David Agnew, Raúl Camposano, Antun Domic, Manfred Wiesel, Hiroto Yasuura
DAC1
1994 Minimization of Memory Traffic in High-Level Synthesis
abstract
In this paper we present a new transformation for the scheduling of memory accessing operations in High-Level Synthesis. This transformation is suited to memory-intensive applications with synthesized designs containing a secondary store accessed by explicit instructions. Such memory-intensive behaviors are commonly observed in video compression, image convolution, hydro-dynamics and mechatronics. Our transformation removes load instructions which become redundant during the transformation of loops. The advantage of this reduction is the decrease of secondary memory bandwidth demands. Our experiments on benchmarks from several application areas show that a significant reduction in the number of memory loads is obtainable.
David J. Kolson, Alexandru Nicolau, Nikil Dutt
DAC3
1994 Integrating program transformations in the memory-based synthesis of image and video algorithms
David J. Kolson, Alexandru Nicolau, Nikil Dutt
ICCAD3
1994 Comprehensive lower bound estimation from behavioral descriptions
abstract
In this paper, we present a comprehensive technique for lower bound estimation (LBE) of resources from behavioral descriptions. Previous work has focused on LBE techniques that use very simple cost models which primarily focus on the functional unit resources. Our cost model accounts for storage resources in additionto functionalresources. Our timing model uses a finer granularity that permits the modeling of functional unit, register and interconnect delays. We tested our LBE technique for both functional unit and storage requirements on several high-level synthesis benchmarks and observed near-optimal results. 1
Seong Yong Ohm, Fadi J. Kurdahi, Nikil Dutt
ICCAD3
1994 Partitioning of Variables for Multiple-Register-File VLIW Architectures
abstract
Recent trends in microprocessor design heavily rely on large register files with large I/O bandwidths for sustaining performance; a possible solution to relieve this bottleneck is the adoption of multiple register files. In this paper we show how the problem of assigning variables to multiple register banks can be reduced to that of a hypergraph coloring and, also, propose a technique to perform this coloring; this technique is applied to the problem of variable partitioning for rnultipltregister- file VLIW architectures.
Andrea Capitanio, Nikil Dutt, Alexandru Nicolau
ICPP (1)2
1993 High-Level Synthesis of Scalable Architectures for IIR Filters using Multichip Modules
abstract
We present a new technique for the high-level synthesis of scalable^1 MCM-based architectures implementing infinite-impulse response(IIR) filters. Our technique is based on the regular schedules, a class of parallel schedules for computing mth-order IIR filters. The simplicity of the regular schedules facilitates characterization of their inter-processor communications, which is generally difficult to express for parallel algorithms. The characterization of inter-processor communications of the regular schedules enables us to generate instruction-level behavior of the design that can be easily mapped onto MCM-based architectures. We illustrate this mapping of the regular schedules onto an MCM-based architecture by designing a special-purpose processor for the fifth-order elliptic wave filter. Our design yields a scalable performance measured in the filter's sample rate, which is not known to have been achieved by previously published designs. This work differs significantly from "traditional" high-level synthesis techniques in its emphasis on synthesizing scalable, high-performance multichip designs.
Haigeng Wang, Nikil Dutt, Alexandru Nicolau, Kai-Yeung Siu
DAC2
1993 A language for designer controlled behavioral synthesis
Nikil Dutt
Integr.1
1993 Rapid estimation for parameterized components in high-level synthesis
abstract
An important benefit of high-level synthesis is rapid design space exploration through examination of different design alternatives. However, such design space exploration is not feasible without fast and accurate area and delay estimates of the synthesized designs. These estimates must factor in physical design effects and technology-specific information in order to achieve accuracy. High-level synthesis tools often use abstract, parameterized component generators for describing the synthesized RT design, and thus need to be supported by fast and accurate estimators for these parameterized RT-components. Ideally, one would like to obtain the actual area and delay attributes of each component by constructing (or generating) the designs. However, such constructive methods require excessive run times, prohibiting on-line integration with the tasks of scheduling and allocation. This paper describes a fast (constant-time) method for estimating the area and delay of regular-structured generic RT components that are tuned to a particular technology library. The estimation models are generated using a least-square approximation on a set of sample data points from selected component implementations. The authors performed an extensive set of experiments to validate the estimation technique on combinational as well as sequential RT component generators. The results show a prediction of the area and delay to within 10% of the actual values. These models have also been integrated with a high-level synthesis system to permit on-line estimation of a component's area and delay.>
Pradip K. Jha, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.2
1992 Equivalent design representations and transformations for interactive scheduling
abstract
It is pointed out that high-level synthesis (HLS) requires more designer interaction to better meet the needs of experienced designers. However, attempts to create a highly interactive synthesis process are hampered by incompatibility of various representations used during synthesis. To overcome this problem, equivalent representations are needed, as well as equivalence-preserving synthesis transformations. The structured finite state machine (SFSM) design model for scheduled behavior is presented, its equivalence to the control data flow graph (CDFG) model is shown, and primitive behavior-preserving transformations for scheduling are defined. This model and these transformations have been integrated into the BIF interactive environment to permit manual rescheduling of a design.>
Roger P. Ang, Nikil Dutt
ICCAD2
1992 Partitioned register files for VLIWs: a preliminary analysis of tradeoffs
Andrea Capitanio, Nikil Dutt, Alexandru Nicolau
MICRO2
1991 Bridging High-Level Synthesis to RTL Technology Libraries
abstract
The output of high-level synthesis typically consists of a netlist of generic RTL components and a state sequencing table.While module generators and logic synthesis tools can be used to map RTL components into standard cells or layout geometries, they cannot provide technology mapping into the data book libraries of functional RTL cells used commonly throughout the industrial design community.In this paper, we introduce an approach to implementing generic RTL components with technology-specific RTL library cells.This approach addresses the criticism of designers who feel that high-level synthesis tools should be used in conjunction with existing RTL data books.We describe how GENUS, a library of generic RTL components, is organized for use in high-level synthesis and how DTAS, a functional synthesis system, is used to map GENUS components into RTL library cells.
Nikil Dutt, James R. Kipps
DAC1
1990 An Intermediate Representation for Behavioral Synthesis
abstract
This paper describes an intermediate representation for behavioral and structural designs that is based on annotated state tables. It facilitates user control of the synthesis process by allowing specification of partially design structures, and a mixture of behavior, structure and user specified bindings between the abstract behavior and the structure. The format's general model allows the capture of synchronous and asynchronous behavior, and permits hierarchical descriptions with concurrency. The format is easily translated to VHDL for simulation at each stage of the design process. It therefore complements a good simulation language (VHDL) by providing an excellent input path for behavioral and register-transfer synthesis. The format's simple and uniform syntax allows it to be used both as an intermediate exchange format for various behavioral synthesis tools, and as a graphical tabular interface for the user, thereby allowing a natural medium for automatic or manual refinement of the design.
Nikil Dutt, Tedd Hadley, Daniel Gajski
DAC1
1989 Designer Controlled Behavioral Synthesis
abstract
This paper describes features of EXEL, a graphic language that gives the designer control over the behavioral synthesis process. Control is achieved by allowing the designer to partially specify the structural design into which the description is going to be compiled, or by binding desired variables and operators to particular components or connections, and binding desired operations to particular states of the final design. EXEL's compiler runs on SUN-3 workstations and is written in C and SUNVIEW.
Nikil Dutt, Daniel Gajski
DAC1