Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ann Gordon-Ross

dblp:47/5909 · DBLP profile ↗
← Back
73ranked-venue papers
12as first author
0since 2021 · last 2019
0000-0001-8865-8381ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 62 · 12 first-authorSoftware engineering, systems software and programming languages · 9 · 2 first-authorComputer networks · 5Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Processor architecture and microarchitecture · 23% Embedded and real-time systems · 22% Memory systems · 20%
Computer networks
4 papers
Edge and fog computing · 37% Cellular and mobile networks · 37% Internet of things and sensor networks · 22%

Topics — the 27 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Embedded and real-time systems › embedded hardware platform
multicore embedded systems
0.322014
Multi-Core Embedded Wireless Sensor Networks: Architecture and Applications · IEEE Trans. Parallel Distributed Syst. 2014
High-Performance Energy-Efficient Multicore Embedded Computing · IEEE Trans. Parallel Distributed Syst. 2012
Edge and fog computing
iot edge computing
0.312018
Microprocessor Optimizations for the Internet of Things: A Survey · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Memory systems
cache
0.222013
A Cache Tuning Heuristic for Multicore Architectures · IEEE Trans. Computers 2013
A Self-Tuning Configurable Cache · DAC 2007
Processor architecture and microarchitecture
multicore design
0.222014
Multi-Core Embedded Wireless Sensor Networks: Architecture and Applications · IEEE Trans. Parallel Distributed Syst. 2014
High-Performance Energy-Efficient Multicore Embedded Computing · IEEE Trans. Parallel Distributed Syst. 2012
Embedded and real-time systems › wireless communication
wireless sensor networks
0.212014
Multi-Core Embedded Wireless Sensor Networks: Architecture and Applications · IEEE Trans. Parallel Distributed Syst. 2014
Energy-efficient computing › power management › memory power management
cache energy reduction
0.222013
A Cache Tuning Heuristic for Multicore Architectures · IEEE Trans. Computers 2013
A Self-Tuning Configurable Cache · DAC 2007
Performance modeling and evaluation › simulation
cache simulation
0.212013
T-SPaCS - A Two-Level Single-Pass Cache Simulation Methodology · IEEE Trans. Computers 2013
Memory systems › cache management
cache tuning
0.212013
A Cache Tuning Heuristic for Multicore Architectures · IEEE Trans. Computers 2013
Performance modeling and evaluation
simulation
0.212013
T-SPaCS - A Two-Level Single-Pass Cache Simulation Methodology · IEEE Trans. Computers 2013
Internet of things and sensor networks
wireless sensor network
0.112012
An MDP-Based Dynamic Optimization Methodology for Wireless Sensor Networks · IEEE Trans. Parallel Distributed Syst. 2012
Electronic design automation
hardware/software co-design
0.112012
High-Performance Energy-Efficient Multicore Embedded Computing · IEEE Trans. Parallel Distributed Syst. 2012
Cellular and mobile networks › low-latency communication
connection setup latency
0.112010
SIP-Based IMS Signaling Analysis for WiMax-3G Interworking Architectures · IEEE Trans. Mob. Comput. 2010
Cellular and mobile networks › mobile networks › mobile network architecture › mobile core network
IP multimedia subsystem
0.112010
SIP-Based IMS Signaling Analysis for WiMax-3G Interworking Architectures · IEEE Trans. Mob. Comput. 2010
Performance modeling and evaluation
workload characterization
0.112018
Microprocessor Optimizations for the Internet of Things: A Survey · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Reconfigurable computing and FPGAs
FPGA implementation
0.112016
MACS: A Highly Customizable Low-Latency Communication Architecture · IEEE Trans. Parallel Distributed Syst. 2016
Memory systems › cache › cache organization
configurable cache
0.112007
A Self-Tuning Configurable Cache · DAC 2007
Memory systems
cache design
0.112006
Configurable cache subsetting for fast cache tuning · DAC 2006
Electronic design automation
design space exploration
0.112006
Configurable cache subsetting for fast cache tuning · DAC 2006
Internet of things and sensor networks › wireless sensor network
in-network processing
0.112014
Multi-Core Embedded Wireless Sensor Networks: Architecture and Applications · IEEE Trans. Parallel Distributed Syst. 2014
Processor architecture and microarchitecture
dynamic optimization
0.112005
Frequent Loop Detection Using Efficient Nonintrusive On-Chip Hardware · IEEE Trans. Computers 2005
Embedded and real-time systems
embedded processor
0.112005
Frequent Loop Detection Using Efficient Nonintrusive On-Chip Hardware · IEEE Trans. Computers 2005
Embedded and real-time systems › runtime monitoring
non-intrusive profiling
0.112005
Frequent Loop Detection Using Efficient Nonintrusive On-Chip Hardware · IEEE Trans. Computers 2005
Memory systems › memory hierarchy
cache hierarchy
0.012013
T-SPaCS - A Two-Level Single-Pass Cache Simulation Methodology · IEEE Trans. Computers 2013
Processor architecture and microarchitecture
chip multiprocessor
0.012013
A Cache Tuning Heuristic for Multicore Architectures · IEEE Trans. Computers 2013
Energy-efficient computing
energy management
0.012012
An MDP-Based Dynamic Optimization Methodology for Wireless Sensor Networks · IEEE Trans. Parallel Distributed Syst. 2012
Internet architecture and protocols › signaling protocol
SIP
0.012010
SIP-Based IMS Signaling Analysis for WiMax-3G Interworking Architectures · IEEE Trans. Mob. Comput. 2010
Energy-efficient computing
power management
0.012007
A Self-Tuning Configurable Cache · DAC 2007

Methods — techniques the papers use, named apart from their topics

workload characterization · 0.7microarchitectural survey · 0.7performance comparison · 0.4parallelization · 0.4trace simulation · 0.2path resolution algorithm · 0.2trace-driven simulation · 0.2single-pass simulation · 0.2hardware cache tuner · 0.2design space search heuristic · 0.2markov decision process · 0.1signaling delay analysis · 0.1
YearPublicationVenuePosition
2019 Dynamic Scheduling on Heterogeneous Multicores
abstract
Heterogeneous multicore systems help meet design goals by using disparate hardware components that are suitable for different application requirements/design goals. The individual cores may also have different tunable hardware parameters for additional specialization. However, this complicates scheduling since to reap the benefits of specialization, applications should be scheduled to the core that offers the best configuration based on the application's requirements and design goals. This scheduling decision could be made by exploring the design space to evaluate different configurations to determine the best configuration, or by executing the application in a base configuration to gather execution statistics to predict the best configuration. However, given increasingly complex systems, these methods may be infeasible given extremely large design spaces or difficulty in choosing a representative base configuration. In this paper, we present a dynamic scheduling methodology that uses predictive methods to schedule applications to best configurations for reduced energy consumption for a system with configurable caches. We use an artificial neural network (ANN) to train our predictive model using hardware counters. The trained ANN can then be used to predict the best core and a tuning heuristic explores the design space to determine the best configuration on non-best cores. If the best core is busy, our scheduler considers alternative idle cores or the application is stalled depending on which decision is energy advantageous. Our experiments show that system energy can be reduced by 28% on average as compared to a fixed-core system where all cores offer the same configuration.
Ayobami S. Edun, Ruben Vazquez, Ann Gordon-Ross, Greg Stitt
DATE3
2019 Accelerating Scientific Discovery with SCAIGATE Science Gateway
abstract
The demand for computational accelerators (GPUs, FPGAs, ASICs, etc.) is growing due to the widening variety of datacenter applications fueled by recent scientific breakthroughs that leverage artificial intelligence (AI). As much as these applications (e.g., cosmology, physics, etc.) have continued to witness record-breaking accuracy in predictive capabilities due to AI widespread influence, the infrastructure and workflow to take these applications out of research labs into production and business use-cases continues to lag. To address these important infrastructural challenges, we present SCAIGATE, a prototype science gateway with a simplified workflow aimed at facilitating model building/validation workflows in large-scale scientific applications.
David Ojika, Bhavesh Patel, Ann Gordon-Ross, Herman Lam
eScience4
2019 Energy Prediction for Cache Tuning in Embedded Systems
abstract
Modern embedded systems are longer tasked at operating a single application or function and are increasingly required to operate more like general purpose desktop computers. Conforming to modern usage demands is extremely challenging given an embedded system's stringent design constraints, such as power, energy, and performance. Adherence to these constraints can be achieved by specializing/tuning the underlying system to application-specific execution requirements and characteristics by tuning a system's configurable parameters to meet these requirements given design constraints. Configurable parameters include architectural voltage, frequency, cache size, line size, and associativity, etc. However, given the complexity of modern systems, exploring these large design spaces is infeasible when the number of configurable parameters and valid parameter values increases beyond a trivial amount. In this paper, we propose using machine learning in lieu of traditional design space exploration techniques. In this work, we evaluate the potential for using an artificial neural network (ANN)-based prediction module for energy prediction. Since the cache hierarchy has a large impact on total energy consumption, without loss of generality, we study a configurable cache hierarchy with configurable cache size, associativity, and line size. We design and train an energy prediction module to infer the best cache configuration for an application based on the application's execution characteristics. Our approach requires only a single profiling run of the application to collect these characteristics. Our energy prediction module then predicts the energy consumption for all the configurations in the cache design space based on these characteristics, and outputs the configuration with the lowest energy consumption, thus essentially performing exhaustive design space exploration with a single execution. Our results show that our prediction module predicts the best instruction and data cache configurations for the majority of the applications, yielding an average energy degradation of less than 2% for both the instruction and data caches as compared to the optimal configuration determined by exhaustive design space exploration.
Ruben Vazquez, Ann Gordon-Ross, Greg Stitt
ICCD2
2018 Microprocessor Optimizations for the Internet of Things: A Survey
abstract
The Internet of Things (IoT) refers to a pervasive presence of interconnected and uniquely identifiable physical devices. These devices' goal is to gather data and drive actions in order to improve productivity, and ultimately reduce or eliminate reliance on human intervention for data acquisition, interpretation, and use. The proliferation of these connected low-power devices will result in a data explosion that will significantly increase data transmission costs with respect to energy consumption and latency. Edge computing reduces these costs by performing computations at the edge nodes, prior to data transmission, to interpret and/or utilize the data. While much research has focused on the IoT's connected nature and communication challenges, the challenges of IoT embedded computing with respect to device microprocessors has received much less attention. This paper explores IoT applications' execution characteristics from a microarchitectural perspective and the microarchitectural characteristics that will enable efficient and effective edge computing. To tractably represent a wide variety of next-generation IoT applications, we present a broad IoT application classification methodology based on application functions, to enable quicker workload characterizations for IoT microprocessors. We then survey and discuss potential microarchitectural optimizations and computing paradigms that will enable the design of right-provisioned microprocessors that are efficient, configurable, extensible, and scalable. This paper provides a foundation for the analysis and design of a diverse set of microprocessor architectures for next-generation IoT devices.
Tosiron Adegbija, Anita Rogacs, Chandrakant Patel, Ann Gordon-Ross
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2018 A comparison-free sorting algorithm on CPUs and GPUs
Saleh Abdel-Hafeez, Ann Gordon-Ross, Samer Abubaker
J. Supercomput.2
2018 PhLock: A Cache Energy Saving Technique Using Phase-Based Cache Locking
abstract
Caches are commonly used to bridge the processor-memory performance gap in embedded systems. Since embedded systems typically have stringent design constraints imposed by physical size, battery capacity, and real-time deadlines much research focuses on cache optimizations, such as improved performance and/or reduced energy consumption. Cache locking is a popular cache optimization that loads and retains/locks selected memory contents from an executing application into the cache to increase the cache's predictability. Previous work has shown that cache locking also has the potential to improve cache energy consumption. In this paper, we introduce phase-based cache locking, PhLock, which leverages an application's varying runtime characteristics to dynamically select the locked memory contents to optimize cache energy consumption. Using a variety of applications from the SPEC2006 and MiBench benchmark suites, experimental results show that PhLock is promising for reducing both the instruction and data caches' energy consumption. As compared to a nonlocking cache, PhLock reduced the instruction and data cache energy consumption by an average of 5% and 39%, respectively, for SPEC2006 applications, and by 75% and 14%, respectively, for MiBench benchmarks.
Tosiron Adegbija, Ann Gordon-Ross
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Optimizing FPGA Performance, Power, and Dependability with Linear Programming
abstract
Field-programmable gate arrays (FPGA) are an increasingly attractive alternative to traditional microprocessor-based computing architectures in extreme-computing domains, such as aerospace and supercomputing. FPGAs offer several resource types that offer different tradeoffs between speed, power, and area, which make FPGAs highly flexible for varying application computational requirements. However, since an application’s computational operations can map to different resource types, a major challenge in leveraging resource-diverse FPGAs is determining the optimal distribution of these operations across the device’s available resources for varying FPGA devices, resulting in an extremely large design space. In order to facilitate fast design-space exploration, this article presents a method based on linear programming (LP) that determines the optimal operation distribution for a particular device and application with respect to performance, power, or dependability metrics. Our LP method is an effective tool for exploring early designs by quickly analyzing thousands of FPGAs to determine the best FPGA devices and operation distributions, which significantly reduces design time. We demonstrate our LP method’s effectiveness with two case studies involving dot-product and distance-calculation kernels on a range of Virtex-5 FPGAs. Results show that our LP method selects optimal distributions of operations to within an average of 4% of actual values.
Nicholas Wulf, Alan D. George, Ann Gordon-Ross
ACM Trans. Reconfigurable Technol. Syst.3
2017 An Efficient O(N) Comparison-Free Sorting Algorithm
abstract
In this paper, we propose a novel sorting algorithm that sorts input data integer elements on-the-fly without any comparison operations between the data-comparison-free sorting. We present a complete hardware structure, associated timing diagrams, and a formal mathematical proof, which show an overall sorting time, in terms of clock cycles, that is linearly proportional to the number of inputs, giving a speed complexity on the order of O(N). Our hardware-based sorting algorithm precludes the need for SRAM-based memory or complex circuitry, such as pipelining structures, but rather uses simple registers to hold the binary elements and the elements' associated number of occurrences in the input set, and uses matrix-mapping operations to perform the sorting process. Thus, the total transistor count complexity is on the order of O(N). We evaluate an application-specified integrated circuit design of our sorting algorithm for a sample sorting of N = 1024 elements of size K = 10-bit using 90-nm Taiwan Semiconductor Manufacturing Company (TSMC) technology with a 1 V power supply. Results verify that our sorting requires approximately 4-6 μs to sort the 1024 elements with a clock cycle time of 0.5 GHz, consumes 1.6 mW of power, and has a total transistor count of less than 750 000.
Saleh Abdel-Hafeez, Ann Gordon-Ross
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Configuration prefetching and reuse for preemptive hardware multitasking on partially reconfigurable FPGAs
Aurelio Morales-Villanueva, Ann Gordon-Ross
DATE3
2016 Quality of Service-Aware, Scalable Cache Tuning Algorithm in Consumer-based Embedded Devices
abstract
To meet energy and quality of service (QoS) constraints in consumer-based embedded devices (CEDs), configurable caches can be tuned to a best configuration that consumes the least amount of energy while adhering to QoS expectations. However, due to disparate consumer QoS expectations and a myriad of unknown, third-party CED applications, tuning caches in CEDs is very challenging. In this paper, we propose a quality of service-aware, scalable tuning algorithm for configurable caches, which requires no a priori knowledge of applications or design-time efforts.
Mohamad Hammam Alsafrjalani, Ann Gordon-Ross
ACM Great Lakes Symposium on VLSI2
2016 MACS: A Highly Customizable Low-Latency Communication Architecture
abstract
Networks-on-chips (NoCs) are an increasingly popular communication infrastructure in single chip VLSI design for enhancing parallelism and system scalability. Processing elements (PEs) connect to a communication topology via NoC switches, which are responsible for runtime establishment and management of inter-PE communication channels. Since NoC switch design directly affects overall system performance and exploited communication parallelism, much previous work focused on efficient NoC switch design. In this paper, we present MACS-a highly parametric NoC switch architecture that provides reduced data transfer latency, increased designer flexibility, and scalability as compared to previous architectures by combining and enhancing several NoC design strategies. MACS enhances inter-PE communication using a circuit switching technique with minimal adaptive routing and a simple and fair path resolution algorithm to maximize bandwidth utilization. We evaluate area and performance of an FPGA implementation of MACS, and, show that compared to previous work, MACS offers a 2× to 7× decrease in average channel setup latency, a 1.7× to 2× reduction in area requirements, similar average packet latency, up to a 6× increase in the network saturation point, and up to a 1.4× increase in bandwidth utilization. Additionally, we illustrate MACS's low average channel setup latency using six network traffic patterns and eight parallel JPEG decompression core trace simulations.
Ann Gordon-Ross
IEEE Trans. Parallel Distributed Syst.2
2016 A Framework for Evaluating and Optimizing FPGA-Based SoCs for Aerospace Computing
abstract
On-board processing systems are often deployed in harsh aerospace environments and must therefore adhere to stringent constraints such as low power, small size, and high dependability in the presence of faults. Field-programmable gate arrays (FPGAs) are often an attractive option for designers seeking low-power, high-performance devices. However, unlike nonreconfigurable devices, radiation effects can alter an FPGA’s functionality instead of just the device’s data, requiring designers to consider fault-tolerant strategies to mitigate these effects. In this article, we present a framework to ease these system design challenges and aid designers in considering a broad range of devices and fault-tolerant strategies for on-board processing, highlighting the most promising options and tradeoffs early in the design process. This article focuses on the power, dependability, and lifetime evaluation metrics, which our framework calculates and leverages to evaluate the effectiveness of varying system-on-chip (SoC) designs. Finally, we use our framework to evaluate SoC designs for a case study on a hyperspectral-imaging (HSI) mission to demonstrate our framework’s ability to identify efficient and effective SoC designs.
Nicholas Wulf, Alan D. George, Ann Gordon-Ross
ACM Trans. Reconfigurable Technol. Syst.3
2015 Phase-based Cache Locking for Embedded Systems
abstract
Since caches are commonly used in embedded systems, which typically have stringent design constraints imposed by physical size, battery capacity, real-time deadlines, etc., much research focuses on cache optimizations, such as improved performance and/or reduced energy consumption. Cache locking is a popular cache optimization that loads and retains/locks selected memory contents from an executing application into the cache to increase the cache's predictability. Previous work has shown that cache locking also has the potential to improve cache performance and energy consumption. In this paper, we introduce phase-based cache locking, which leverages an application's varying runtime characteristics to dynamically select the locked memory contents to optimize cache performance and energy consumption. Experimental results show that our phase-based cache locking methodology can improve the data cache's miss rates and energy consumption by an average of 24% and 20%, respectively.
Tosiron Adegbija, Ann Gordon-Ross
ACM Great Lakes Symposium on VLSI2
2015 Modeling and Analysis of Fault Detection and Fault Tolerance in Wireless Sensor Networks
abstract
Technological advancements in communications and embedded systems have led to the proliferation of Wireless Sensor Networks (WSNs) in a wide variety of application domains. These application domains include but are not limited to mission-critical (e.g., security, defense, space, satellite) or safety-related (e.g., health care, active volcano monitoring) systems. One commonality across all WSN application domains is the need to meet application requirements (e.g., lifetime, reliability). Many application domains require that sensor nodes be deployed in harsh environments, such as on the ocean floor or in an active volcano, making these nodes more prone to failures. Sensor node failures can be catastrophic for critical or safety-related systems. This article models and analyzes fault detection and fault tolerance in WSNs. To determine the effectiveness and accuracy of fault detection algorithms, we simulate these algorithms using ns-2. We investigate the synergy between fault detection and fault tolerance and use the fault detection algorithms’ accuracies in our modeling of Fault-Tolerant (FT) WSNs. We develop Markov models for characterizing WSN reliability and Mean Time to Failure (MTTF) to facilitate WSN application-specific design. Results obtained from our FT modeling reveal that an FT WSN composed of duplex sensor nodes can result in as high as a 100% MTTF increase and approximately a 350% improvement in reliability over a Non-Fault-Tolerant (NFT) WSN. The article also highlights future research directions for the design and deployment of reliable and trustworthy WSNs.
Arslan Munir, Joseph Antoon, Ann Gordon-Ross
ACM Trans. Embed. Comput. Syst.3
2014 Dynamic Scheduling for Reduced Energy in Configuration-Subsetted Heterogeneous Multicore Systems
abstract
Heterogeneous and configurable multicore systems provide hardware specialization to meet disparate application hardware requirements. However, effective multicore system specialization can require a priori knowledge of the applications, application profiling information, and/or dynamic hardware tuning to schedule and execute applications on the most energy efficient cores. Furthermore, even though highly disparate core heterogeneity and/or highly configurable parameters with numerous potential parameter values result in more fine-grained specialization and higher energy savings potential, these large design spaces are challenging to efficiently explore. To address these challenges, we propose a novel configuration-subsisted heterogeneous and configurable multicore system, wherein each core offers a small subset of the design space, and propose a novel scheduling and tuning (SaT) algorithm to efficiently exploit the energy savings potential of this system. Our proposed architecture and algorithm require no a priori application knowledge or profiling, and incurs minimal runtime overhead. Results reveal energy savings potential and insights on energy tradeoffs in heterogeneous, configurable systems.
Mohamad Hammam Alsafrjalani, Ann Gordon-Ross
EUC2
2014 Minimum Effort Design Space Subsetting for Configurable Caches
abstract
Configurable caches can significantly reduce energy consumption by adapting the system's cache configuration to the applications' specific requirements to meet system design and optimization goals. However, large configuration design spaces require prohibitive design space exploration time (e.g., due to lengthy design space analyses, simulations, and/or evaluations) to determine the best configuration given these requirements and goals. To significantly reduce design space exploration time, we evaluate a design space subsetting method that removes energy-redundant configurations (i.e., configurations that provide similar energy savings as other configurations), thus significantly reducing the design space while still providing high-quality, energy-saving configurations. Prior work verified design space subsetting's efficacy, however, prior work required extensive design-time effort and complete a priori knowledge of the system's anticipated applications. In this work, we alleviate these limitations and significantly broaden the usability of design space subsetting. Results show that complete a priori knowledge of the anticipated applications is not necessary, and only a small set of applications representative of the anticipated applications' general domains (or applications with similar requirements) is sufficient to provide energy savings within 5.6% of the complete, unsubsetted design space.
Mohamad Hammam Alsafrjalani, Ann Gordon-Ross, Pablo Viana
EUC2
2014 Thermal-aware phase-based tuning of embedded systems
abstract
Due to embedded systems' stringent design constraints, much prior work focused on optimizing energy consumption and/or performance. However, since embedded systems have fewer cooling options, rising temperature, and thus temperature optimization, is an emergent concern. We present thermal-aware phase-based tuning--TaPT--that determines Pareto optimal configurations for fine-grained execution time, energy, and temperature tradeoffs. Results show that TaPT reduces execution time, energy, and temperature by as much as 5%, 30%, and 25%, respectively, while adhering to designer-specified design constraints.
Tosiron Adegbija, Ann Gordon-Ross
ACM Great Lakes Symposium on VLSI2
2014 Analysis of cache tuner architectural layouts for multicore embedded systems
abstract
Due to the memory hierarchy's large contribution to a microprocessor's total power, cache tuning is an ideal method for optimizing overall power consumption in embedded systems. Since most embedded systems are power and area constrained, the hardware and/or software that orchestrate cache tuning - the cache tuner - must not impose significant power and area overhead. Furthermore, as embedded systems increasingly trend towards multicore, inter-core data sharing, communication, and synchronization impose additional cache tuner design complexity, necessitating cross-core cache tuning coordination. In order to minimize cache tuner overhead, cache tuner design must consider these overheads and scalability. Whereas prior work proposes low-overhead cache tuners, scalability to multicore systems requires additional considerations. In this work, we present a low-overhead, scalable cache tuner and extensively evaluate various cache tuner design tradeoffs with respect to power and area for constrained multicore embedded systems. Based on our analysis, we formulate valuable insights and designer-assisted guidelines for modeling scalable and efficient cache tuners that best achieve optimization goals while maintaining power and area constraints.
Tosiron Adegbija, Ann Gordon-Ross, Marisha Rawlins
IPCCC2
2014 A queueing theoretic approach for performance evaluation of low-power multi-core embedded systems
Arslan Munir, Ann Gordon-Ross, Sanjay Ranka, Farinaz Koushanfar
J. Parallel Distributed Comput.2
2014 Multi-Core Embedded Wireless Sensor Networks: Architecture and Applications
abstract
Technological advancements in the silicon industry, as predicted by Moore's law, have enabled integration of billions of transistors on a single chip. To exploit this high transistor density for high performance, embedded systems are undergoing a transition from single-core to multi-core. Although a majority of embedded wireless sensor networks (EWSNs) consist of single-core embedded sensor nodes, multi-core embedded sensor nodes are envisioned to burgeon in selected application domains that require complex in-network processing of the sensed data. In this paper, we propose an architecture for heterogeneous hierarchical multi-core embedded wireless sensor networks (MCEWSNs) as well as an architecture for multi-core embedded sensor nodes used in MCEWSNs. We elaborate several compute-intensive tasks performed by sensor networks and application domains that would especially benefit from multi-core embedded sensor nodes. This paper also investigates the feasibility of two multi-core architectural paradigms-symmetric multiprocessors (SMPs) and tiled many-core architectures (TMAs)-for MCEWSNs. We compare and analyze the performance of an SMP (an Intel-based SMP) and a TMA (Tilera's TILEPro64) based on a parallelized information fusion application for various performance metrics (e.g., runtime, speedup, efficiency, cost, and performance per watt). Results reveal that TMAs exploit data locality effectively and are more suitable for MCEWSN applications that require integer manipulation of sensor data, such as information fusion, and have little or no communication between the parallelized tasks. To demonstrate the practical relevance of MCEWSNs, this paper also discusses several state-of-the-art multi-core embedded sensor node prototypes developed in academia and industry. We further discuss research challenges and future research directions for MCEWSNs.
Arslan Munir, Ann Gordon-Ross, Sanjay Ranka
IEEE Trans. Parallel Distributed Syst.2
2013 PRML: A Modeling Language for Rapid Design Exploration of Partially Reconfigurable FPGAs
abstract
Leveraging partial reconfiguration (PR) can improve system flexibility, cost, and performance/power/area tradeoffs over non-PR functionally-equivalent systems, however, realizing these benefits is challenging, time-consuming, and PR must be considered early during application design to reduce design exploration time and improve system quality. To facilitate realizing these benefits, we present an application design framework and an abstract modeling language for PR (PRML). By applying extensive PRML modeling guidelines to a complex arithmetic core, we show PRML's potential for efficient PR capability analysis, enabling designers to determine Pareto optimal systems during application formulation based on designer-specified area and performance metrics.
Ann Gordon-Ross
FCCM2
2013 On-chip Context Save and Restore of Hardware Tasks on Partially Reconfigurable FPGAs
abstract
Partial reconfiguration (PR) of field-programmable gate arrays (FPGAs) enables hardware tasks to time multiplex PR regions (PRRs) by isolating reconfiguration to only the reconfigured PRR, which avoids halting the entire FPGA's execution. Time multiplexing PRRs requires support for unloading/loading tasks and for resuming a task's execution state. In order to resume a task's execution state, the execution state (context) must be saved when the task is unloaded so that the execution state can be restored when the task resumes- context save (CS) and context restore (CR), respectively. In this paper, we present a software-based, on-chip context save and restore (CSR) for PR-capable FPGAs. As compared to prior work, our CSR is autonomous (i.e., does not require any external host support), does not require custom on-chip hardware, is portable across any system design, and does not require tool flow modifications or special tools. Experimental results extensively evaluate the CSR execution time based on PRR size, enabling designers to trade off PRR granularity for CSR execution time based on application requirements.
Aurelio Morales-Villanueva, Ann Gordon-Ross
FCCM2
2013 Exploiting dynamic phase distance mapping for phase-based tuning of embedded systems
abstract
Phase-based tuning increases optimization potential by configuring system parameters for application execution phases. Previous work proposed phase distance mapping (PDM), which relied on extensive a priori analysis of executing applications to dynamically estimate the best configuration using the correlation between phases. We propose DynaPDM, a new dynamic phase distance mapping methodology that eliminates a priori designer effort, dynamically analyzes phases, and determines the best configurations, yielding average energy delay product savings of 28%-an 8% improvement on PDM-and configurations within 1% of the optimal.
Tosiron Adegbija, Ann Gordon-Ross
ICCD2
2013 A Cache Tuning Heuristic for Multicore Architectures
abstract
Since multicore architectures are becoming more popular, recent multicore optimizations focus on energy consumption. In this paper, we focus on reducing the energy consumption in the data and instruction cache hierarchies in a multicore system. First, we present a level one data cache tuning heuristic for a heterogeneous multicore system, which classifies applications based on data sharing and cache behavior and uses this classification to guide cache tuning and reduce the number of cores that need to be tuned. Results reveal average energy savings of 25 percent for 2, 4, 8, and 16-core systems while searching only 1 percent of the design space. Next, we present a level one instruction cache tuning heuristic that reduces energy consumption in the instruction cache hierarchy by an average of 53 percent for 2, 4, 8, and 16-core systems, while searching less than 1 percent of the design space. Finally, we develop a custom, global hardware cache tuner for a dual-core system and show that our cache tuner has low area, energy, and power overheads.
Marisha Rawlins, Ann Gordon-Ross
IEEE Trans. Computers2
2013 T-SPaCS - A Two-Level Single-Pass Cache Simulation Methodology
abstract
The cache hierarchy's large contribution to total microprocessor system power makes caches a good optimization candidate. To facilitate a fast design-time cache optimization process, we propose a single-pass trace-driven cache simulation methodology-T-SPaCS-for a two-level exclusive cache hierarchy. Direct adaptation of conventional trace-driven cache simulation to two-level caches requires significant storage and simulation time as numerous stacks record cache access patterns for each level one and level two cache combination and each stack is repeatedly processed. T-SPaCS significantly reduces storage space and simulation time using a set of stacks that only record the complete cache access pattern. Thereby, T-SPaCS simulates all cache configurations for both the level one and level two caches simultaneously in a single pass. Experimental results show that T-SPaCS is 21.02X faster on average than sequential simulation for instruction caches and 33.34X faster for data caches. A simplified, but minimally lossy version of T-SPaCS (simplified-T-SPaCS) increases the average simulation speedup to 30.15X for instruction caches and 41.31X for data caches. We leverage T-SPaCS and simplified-T-SPaCS for determining the lowest energy cache configuration to quantify the effects of lossiness and observe that T-SPaCS and simplified-T-SPaCS still find the lowest energy cache configuration as compared to exact simulation.
Wei Zang, Ann Gordon-Ross
IEEE Trans. Computers2
2013 Dynamic profiling and fuzzy-logic-based optimization of sensor network platforms
abstract
The commercialization of sensor-based platforms is facilitating the realization of numerous sensor network applications with diverse application requirements. However, sensor network platforms are becoming increasingly complex to design and optimize due to the multitude of interdependent parameters that must be considered. To further complicate matters, application experts oftentimes are not trained engineers, but rather biologists, teachers, or agriculturists who wish to utilize the sensor-based platforms for various domain-specific tasks. To assist both platform developers and application experts, we present a centralized dynamic profiling and optimization platform for sensor-based systems that enables application experts to rapidly optimize a sensor network for a particular application without requiring extensive knowledge of, and experience with, the underlying physical hardware platform. In this article, we present an optimization framework that allows developers to characterize application requirements through high-level design metrics and fuzzy-logic-based optimization. We further analyze the benefits of utilizing dynamic profiling information to eliminate the guesswork of creating a “good” benchmark, present several reoptimization evaluation algorithms used to detect if re-optimization is necessary, and highlight the benefits of the proposed dynamic optimization framework compared to static optimization alternatives.
Adrian Lizarraga, Roman L. Lysecky, Susan Lysecky, Ann Gordon-Ross
ACM Trans. Embed. Comput. Syst.4
2013 Adaptive loop caching using lightweight runtime control flow analysis
abstract
Loop caches provide an effective method for decreasing memory hierarchy energy consumption by storing frequently executed code (critical regions) in a more energy efficient structure than the level one cache. However, due to code structure restrictions or costly design time pre-analysis efforts, previous loop cache designs are not suitable for all applications and system scenarios. We present an adaptive loop cache that is amenable to a wider range of system scenarios, which can provide an additional 20% average instruction cache energy savings (with individual benchmark energy savings as high as 69%) compared to the next best loop cache, the preloaded loop cache.
Marisha Rawlins, Ann Gordon-Ross
ACM Trans. Embed. Comput. Syst.2
2013 High-performance optimizations on tiled many-core embedded systems: a matrix multiplication case study
Arslan Munir, Farinaz Koushanfar, Ann Gordon-Ross, Sanjay Ranka
J. Supercomput.3
2013 Scalable Digital CMOS Comparator Using a Parallel Prefix Tree
abstract
We present a new comparator design featuring wide-range and high-speed operation using only conventional digital CMOS cells. Our comparator exploits a novel scalable parallel prefix structure that leverages the comparison outcome of the most significant bit, proceeding bitwise toward the least significant bit only when the compared bits are equal. This method reduces dynamic power dissipation by eliminating unnecessary transitions in a parallel prefix structure that generates the N-bit comparison result after (log4N)+(log16N)+4 CMOS gate delays. Our comparator is composed of locally interconnected CMOS gates with a maximum fan-in and fan-out of five and four, respectively, independent of the comparator bitwidth. The main advantages of our design are high speed and power efficiency, maintained over a wide range. Additionally, our design uses a regular reconfigurable VLSI topology, which allows analytical derivation of the input-output delay as a function of bitwidth. HSPICE simulation for a 64-b comparator shows a worst case input-output delay of 0.86 ns and a maximum power dissipation of 7.7 mW using 0.15- μm TSMC technology at 1 GHz.
Saleh Abdel-Hafeez, Ann Gordon-Ross, Behrooz Parhami
IEEE Trans. Very Large Scale Integr. Syst.2
2012 An application classification guided cache tuning heuristic for multi-core architectures
abstract
Since multi-core architectures are becoming more popular, recent multi-core optimizations focus on energy consumption. We present a level one data cache tuning heuristic for a heterogeneous multi-core system, which classifies applications based on data sharing and cache behavior, and uses this classification to guide cache tuning and reduce the number of cores that need to be tuned. Results reveal average energy savings of 25% for 2-, 4-, 8-, and 16-core systems while searching only 1% of the design space.
Marisha Rawlins, Ann Gordon-Ross
ASP-DAC2
2012 Online algorithms for wireless sensor networks dynamic optimization
abstract
Technological advancements in wireless communications and embedded systems have led to the proliferation of wireless sensor network (WSN) applications, each with varying application requirements (i.e., lifetime, throughput, reliability, etc.). Sensor node tunable parameters enable WSN designers to specialize/tune a sensor node to meet application requirements, but however, parameter tuning is a challenging process that requires designer expertise to consider sensor node complexities and changing environmental stimuli. In this paper, we develop lightweight, online optimization algorithms for sensor node parameter tuning, which enables dynamic optimizations to meet application requirements and adapt to changing environmental stimuli. Results reveal that our online optimizations quickly converge to a near optimal solution using minimal computational and storage resources, and are thus amenable for implementation on resource and energy-constrained sensor nodes.
Arslan Munir, Ann Gordon-Ross, Susan Lysecky, Roman L. Lysecky
CCNC2
2012 Dynamic phase-based tuning for embedded systems using phase distance mapping
abstract
Phase-based tuning specializes a system's tunable parameters to the varying runtime requirements of an application's different phases of execution to meet optimization goals. Since the design space for tunable systems can be very large, one of the major challenges in phase-based tuning is determining the best configuration for each phase without incurring significant tuning overhead (e.g., energy and/or performance) during design space exploration. In this paper, we propose phase distance mapping, which directly determines the best configuration for a phase, thereby eliminating design space exploration. Phase distance mapping applies the correlation between a known phase's characteristics and best configuration to determine a new phase's best configuration based on the new phase's characteristics. Experimental results verify that our phase distance mapping approach determines configurations within 3% of the optimal configurations on average and yields an energy delay product savings of 26% on average.
Tosiron Adegbija, Ann Gordon-Ross, Arslan Munir
ICCD2
2012 Parallelized benchmark-driven performance evaluation of SMPs and tiled multi-core architectures for embedded systems
abstract
With Moore's law supplying billions of transistors on-chip, embedded systems are undergoing a transition from single-core to multi-core to exploit this high transistor density for high performance. However, there exists a plethora of multi-core architectures and the suitability of these multi-core architectures for different embedded domains (e.g., distributed, real-time, reliability-constrained) requires investigation. Despite the diversity of embedded domains, one of the critical applications in many embedded domains (especially distributed embedded domains) is information fusion. Furthermore, many other applications consist of various kernels, such as Gaussian elimination (used in network coding), that dominate the execution time. In this paper, we evaluate two embedded systems multi-core architectural paradigms: symmetric multiprocessors (SMPs) and tiled multi-core architectures (TMAs). We base our evaluation on a parallelized information fusion application and benchmarks that are used as building blocks in applications for SMPs and TMAs. We compare and analyze the performance of an Intel-based SMP and Tilera's TILEPro64 TMA based on our parallelized benchmarks for the following performance metrics: runtime, speedup, efficiency, cost, scalability, and performance per watt. Results reveal that TMAs are more suitable for applications requiring integer manipulation of data with little communication between the parallelized tasks (e.g., information fusion) whereas SMPs are more suitable for applications with floating point computations and a large amount of communication between processor cores.
Arslan Munir, Ann Gordon-Ross, Sanjay Ranka
IPCCC2
2012 A single-pass cache simulation methodology for two-level unified caches
abstract
Cache tuning is the process of determining the optimal cache configuration given an application's requirements for reducing energy consumption and improving performance. As embedded systems trend towards unified second-level caches for improved performance, the need for fast cache tuning methodologies for multi-level cache hierarchies is becoming more critical. In this paper, we present U-SPaCS, a single-pass cache simulation methodology for design-time tuning of two-level cache hierarchies with a unified second-level cache. To afford fast simulation time, U-SPaCS maintains unique cache block addresses in a set of stacks, which enables simulation of all cache configurations for the level one instruction and data caches, and level two unified cache simultaneously in a single pass of an application's time-ordered instruction and data access trace. Experiments show that U-SPaCS can accurately determine the miss rates for a configurable cache design space consisting of 2,187 cache configurations with a 41X speedup in average simulation time as compared to the most widely-used trace-driven cache simulation.
Wei Zang, Ann Gordon-Ross
ISPASS2
2012 Combining code reordering and cache configuration
abstract
The instruction cache is a popular optimization target due to the cache's high impact on system performance and power and because of the cache's predictable temporal and spatial locality. This article is an in depth study on the interaction of code reordering (a long-known technique) and cache configuration (a relatively new technique). Experimental results show that code reordering coupled with cache configuration reveals additional energy savings as high as 10--15% for several benchmarks with reduced cache area as high as 48%. To exploit these additional benefits, we architect and evaluate several design exploration heuristics for combining these two methods.
Ann Gordon-Ross, Frank Vahid, Nikil Dutt
ACM Trans. Embed. Comput. Syst.1
2012 Dynamic Cache Reconfiguration for Soft Real-Time Systems
abstract
In recent years, efficient dynamic reconfiguration techniques have been widely employed for system optimization. Dynamic cache reconfiguration is a promising approach for reducing energy consumption as well as for improving overall system performance. It is a major challenge to introduce cache reconfiguration into real-time multitasking systems, since dynamic analysis may adversely affect tasks with timing constraints. This article presents a novel approach for implementing cache reconfiguration in soft real-time systems by efficiently leveraging static analysis during runtime to minimize energy while maintaining the same service level. To the best of our knowledge, this is the first attempt to integrate dynamic cache reconfiguration in real-time scheduling techniques. Our experimental results using a wide variety of applications have demonstrated that our approach can significantly reduce the cache energy consumption in soft real-time systems (up to 74%).
Weixun Wang, Prabhat Mishra 0001, Ann Gordon-Ross
ACM Trans. Embed. Comput. Syst.3
2012 An MDP-Based Dynamic Optimization Methodology for Wireless Sensor Networks
abstract
Wireless sensor networks (WSNs) are distributed systems that have proliferated across diverse application domains (e.g., security/defense, health care, etc.). One commonality across all WSN domains is the need to meet application requirements (i.e., lifetime, responsiveness, etc.) through domain specific sensor node design. Techniques such as sensor node parameter tuning enable WSN designers to specialize tunable parameters (i.e., processor voltage and frequency, sensing frequency, etc.) to meet these application requirements. However, given WSN domain diversity, varying environmental situations (stimuli), and sensor node complexity, sensor node parameter tuning is a very challenging task. In this paper, we propose an automated Markov Decision Process (MDP)-based methodology to prescribe optimal sensor node operation (selection of values for tunable parameters such as processor voltage, processor frequency, and sensing frequency) to meet application requirements and adapt to changing environmental stimuli. Numerical results confirm the optimality of our proposed methodology and reveal that our methodology more closely meets application requirements compared to other feasible policies.
Arslan Munir, Ann Gordon-Ross
IEEE Trans. Parallel Distributed Syst.2
2012 High-Performance Energy-Efficient Multicore Embedded Computing
abstract
With Moore's law supplying billions of transistors on-chip, embedded systems are undergoing a transition from single-core to multicore to exploit this high-transistor density for high performance. Embedded systems differ from traditional high-performance supercomputers in that power is a first-order constraint for embedded systems; whereas, performance is the major benchmark for supercomputers. The increase in on-chip transistor density exacerbates power/thermal issues in embedded systems, which necessitates novel hardware/software power/thermal management techniques to meet the ever-increasing high-performance embedded computing demands in an energy-efficient manner. This paper outlines typical requirements of embedded applications and discusses state-of-the-art hardware/software high-performance energy-efficient embedded computing (HPEEC) techniques that help meeting these requirements. We also discuss modern multicore processors that leverage these HPEEC techniques to deliver high performance per watt. Finally, we present design challenges and future research directions for HPEEC system development.
Arslan Munir, Sanjay Ranka, Ann Gordon-Ross
IEEE Trans. Parallel Distributed Syst.3
2012 Reconfigurable Fault Tolerance: A Comprehensive Framework for Reliable and Adaptive FPGA-Based Space Computing
abstract
Commercial SRAM-based, field-programmable gate arrays (FPGAs) have the potential to provide space applications with the necessary performance to meet next-generation mission requirements. However, mitigating an FPGA’s susceptibility to single-event upset (SEU) radiation is challenging. Triple-modular redundancy (TMR) techniques are traditionally used to mitigate radiation effects, but TMR incurs substantial overheads such as increased area and power requirements. In order to reduce these overheads while still providing sufficient radiation mitigation, we propose a reconfigurable fault tolerance (RFT) framework that enables system designers to dynamically adjust a system’s level of redundancy and fault mitigation based on the varying radiation incurred at different orbital positions. This framework includes an adaptive hardware architecture that leverages FPGA reconfigurable techniques to enable significant processing to be performed efficiently and reliably when environmental factors permit. To accurately estimate upset rates, we propose an upset rate modeling tool that captures time-varying radiation effects for arbitrary satellite orbits using a collection of existing, publically available tools and models. We perform fault-injection testing on a prototype RFT platform to validate the RFT architecture and RFT performability models. We combine our RFT hardware architecture and the modeled upset rates using phased-mission Markov modeling to estimate performability gains achievable using our framework for two case-study orbits.
Adam Jacobs, Grzegorz Cieslewski, Alan D. George, Ann Gordon-Ross, Herman Lam
ACM Trans. Reconfigurable Technol. Syst.4
2011 An integrated development toolset and implementation methodology for partially reconfigurable system-on-chips
abstract
Partial reconfiguration (PR) enhances traditional FPGA-based system-on-chips (SoCs) by providing additional benefits such as reduced area and increased functionality as compared to non-PR SoCs. However, since leveraging these additional benefits requires specific designer expertise and increased development time, PR has not yet gained widespread usage. In this paper, we present an integrated development toolset that automates the implementation of PR SoCs on FPGA devices and leverage this tool in a rapid design space exploration case study.
Abelardo Jara-Berrocal, Ann Gordon-Ross
ASAP2
2011 On the interplay of loop caching, code compression, and cache configuration
abstract
Even though much previous work explores varying instruction cache optimization techniques individually, little work explores the combined effects of these techniques (i.e., do they complement or obviate each other). In this paper we explore the interaction of three optimizations: loop caching, cache tuning, and code compression. Results show that loop caching increases energy savings by as much as 26% compared to cache tuning alone and reduces decompression energy by as much as 73%.
Marisha Rawlins, Ann Gordon-Ross
ASP-DAC2
2011 T-SPaCS - A two-level single-pass cache simulation methodology
abstract
The cache hierarchy's large contribution to total microprocessor system power makes caches a good optimization candidate. We propose a single-pass trace-driven cache simulation methodology - T-SPaCS - for a two-level exclusive instruction cache hierarchy. Instead of storing and simulating numerous stacks repeatedly as in direct adaptation of a conventional trace-driven cache simulation to two level caches, T-SPaCS simulates both the level one and level two caches simultaneously using one stack. Experimental results show T-SPaCS efficiently and accurately determines the optimal cache configuration (lowest energy).
Wei Zang, Ann Gordon-Ross
ASP-DAC2
2011 Hardware module reuse and runtime assembly for dynamic management of reconfigurable resources
abstract
Partial reconfiguration (PR) enhances traditional FPGA-based systems-on-a-chip (SoCs) by providing benefits such as reduced area requirements and increased system flexibility. In multi-application PR SoCs, a dynamic resource manager (DRM) must efficiently orchestrate PR hardware resource management (access to and sharing of PR resources) in order to minimize the percentage of wasted/unused PR resources and reconfiguration time overhead. In this paper, we present DRM software that leverages two techniques, hardware module reuse and dynamic inter-module communication, to reduce wasted/unused PR hardware resources by 13% and reduce reconfiguration time by 33% as compared to a DRM without these techniques.
Abelardo Jara-Berrocal, Ann Gordon-Ross
FPT2
2011 Formulation-level design space exploration for partially reconfigurable FPGAs
abstract
Exploiting the benefits afforded by runtime partial reconfiguration (PR) on modern field-programmable gate arrays (FPGAs)requires PR-capable applications and associated PR-architectures, both of which are challenging tasks due to competing implementation metrics(e.g., area, power, operating frequency, etc.) and results in unmanageable design spaces. PR design space exploration (DSE) techniques and tools assist designers in efficiently and effectively exploring this design space. This paper presents the first, to the best of our knowledge, formulation-level PR DSE tool - FoRSE. FoRSE leverages the application's PR-architecture and mathematical FPGA device models and vendor-specified PR technology to generate Pareto-optimal sets of PR-floorplans and devices based on designer-designated implementation metrics. FoRSE can prune an application's implementation design space by three to four orders of magnitude in approximately 15 seconds.
Ann Gordon-Ross
FPT2
2011 Partially reconfigurable system-on-chips for adaptive fault tolerance
abstract
Due to the runtime flexibility of modern dynamically reconfigurable SRAM-based FPGAs, FPGA devices have become an attractive platform for developing system-on-chips (SoCs) for space applications (space SoCs). However, since the FPGA's SRAM is highly susceptible to space radiation, system reliability is a primary concern for space SoCs. To maintain system reliability and mitigate space radiation effects, space SoCs must be designed with redundant copies of system functionality. Space SoCs must contain enough redundancy to ensure system reliability for the highest anticipated radiation level, which imposes a large area overhead. However, since radiation levels vary based on the system's orbital position, the system does not always require the highest level of redundancy. Space SoCs that can adapt the system redundancy based on the current radiation level can achieve more effective device utilization. In this paper, we present a flexible, FPGA-based, adaptive SoC for space system development. Our space SoC leverages partial reconfiguration to dynamically adapt the system's level of redundancy according to varying radiation levels. We present a software algorithm to manage the system's adaptability, implement the SoC on a Xilinx Virtex-5 device, and evaluate the SoC's resource utilization using the International Space Station's orbit.
Shaon Yousuf, Adam Jacobs, Ann Gordon-Ross
FPT3
2011 Markov Modeling of Fault-Tolerant Wireless Sensor Networks
abstract
Technological advancements in communications and embedded systems have led to the proliferation of wireless sensor networks (WSNs) in a wide variety of application domains. One commonality across all WSN application domains is the need to meet application requirements (e.g., lifetime, reliability, etc.). Many application domains require that sensor nodes be deployed in harsh environments (e.g., ocean floor, active volcanoes), making these sensor nodes more prone to failures. Unfortunately, sensor node failures can be catastrophic for critical or safety related systems. To improve reliability in such systems, we propose a fault-tolerant sensor node model for applications with high reliability requirements. We develop Markov models for characterizing WSN reliability and MTTF (Mean Time to Failure) to facilitate WSN application-specific design. Results show that our proposed fault-tolerant model can result in as high as a 100% MTTF increase and approximately a 350% improvement in reliability over a non-fault-tolerant WSN. Results also highlight the significance of a robust fault detection algorithm to leverage the benefits of fault-tolerant WSNs.
Arslan Munir, Ann Gordon-Ross
ICCCN2
2011 A queueing theoretic approach for performance evaluation of low-power multi-core embedded systems
abstract
With Moore's law supplying billions of transistors on-chip, embedded systems are undergoing a transition from single-core to multi-core to exploit this high transistor density for high performance. However, the optimal layout of these multiple cores along with the memory subsystem (caches and main memory) to satisfy power, area, and often stringent real-time constraints is a challenging design endeavor. The short time-to-market constraint of embedded systems exacerbates this design challenge and necessitates the architectural modeling of embedded systems to reduce the time-to-market by expediting target applications to device/architecture mapping. In this paper, we present a queueing theoretic approach for modeling multi-core embedded systems that provides a quick and inexpensive performance evaluation both in terms of time and resources as compared to the development of multi-core simulators and running benchmarks on these simulators. We also calculate chip area and power consumption for different multi-core embedded architectures with a varying number of processor cores and cache configurations to provide a comparative analysis of multicore embedded architectures in terms of performance, area, and power consumption. Our performance and power results indicate that multi-core embedded system architectures that leverage shared last-level caches (LLCs) provide the best LLC performance per watt but may introduce main memory response time and throughput bottlenecks for high cache miss rates, whereas architectures leveraging a hybrid of private and shared LLCs alleviate main memory bottlenecks at the expense of reduced performance per watt.
Arslan Munir, Ann Gordon-Ross, Sanjay Ranka
ICCD2
2011 CPACT - The conditional parameter adjustment cache tuner for dual-core architectures
abstract
Cache tuning reveals substantial energy savings for single-core architectures, but has yet to be explored for multi-core architectures. In this paper we explore level one (L1) data cache tuning in a heterogeneous dual-core system where each data cache can have a different configuration. We show that L1 data cache tuning in a dual-core system achieves 25% average energy savings, which is comparable to single-core data cache tuning. We present the dual-core tuning heuristic CPACT, which finds cache configurations within 1% of the optimal configuration while searching only 1% of the design space. Finally, we provide valuable insights on core-interactions and data coherence revealed when tuning the multithreaded SPLASH-2 benchmarks.
Marisha Rawlins, Ann Gordon-Ross
ICCD2
2011 A Digital CMOS Parallel Counter Architecture Based on State Look-Ahead Logic
abstract
We present a high-speed wide-range parallel counter that achieves high operating frequencies through a novel pipeline partitioning methodology (a counting path and state look-ahead path), using only three simple repeated CMOS-logic module types: an initial module generates anticipated counting states for higher significant bit modules through the state look-ahead path, simple D-type flip-flops, and 2-bit counters. The state look-ahead path prepares the counting path's next counter state prior to the clock edge such that the clock edge triggers all modules simultaneously, thus concurrently updating the count state with a uniform delay at all counting path modules/stages with respect to the clock edge. The structure is scalable to arbitrary N-bit counter widths (2-to-2N range) using only the three module types and no fan-in or fan-out increase. The counter's delay is comprised of the initial module access time (a simple 2-bit counting stage), one three-input and-gate delay, and a D-type flip-flop setup-hold time. We implemented our proposed counter using a 0.15-μ m TSMC digital cell library and verified maximum operating speeds of 2 and 1.8 GHz for 8- and 17-bit counters, respectively. Finally, the area of a sample 8-bit counter was 78 125 μ m2(510 transistors) and consumed 13.89 mW at 2 GHz.
Saleh Abdel-Hafeez, Ann Gordon-Ross
IEEE Trans. Very Large Scale Integr. Syst.2
2010 VAPRES: A Virtual Architecture for Partially Reconfigurable Embedded Systems
abstract
Due to the runtime flexibility offered by field programmable gate arrays (FPGAs), FPGAs are popular devices for stream processing systems, since many stream processing applications require runtime adaptability (i.e. throughput, data transformations, etc.). FPGAs can offer this adaptability through runtime assembly of stream processing systems that are decomposed into hardware modules. Runtime hardware module assembly consists of dynamic hardware module replacement and hardware module communication reconfiguration. In this paper, we architect a flexible base embedded system amenable to runtime assembly of stream processing systems using custom communication architecture with dynamic streaming channel establishment between hardware modules. We present a hardware module swapping methodology that replaces hardware modules without stream processing interruption. Finally, we formulate two design flows, system and application construction, to provide system and application designer assistance.
Abelardo Jara-Berrocal, Ann Gordon-Ross
DATE2
2010 Lightweight runtime control flow analysis for adaptive loop caching
abstract
Loop caches provide an effective method for decreasing memory hierarchy energy consumption by storing frequently executed code in a more energy efficient structure than the level one cache. However, due to code structure restrictions and/or costly design time pre-analysis efforts, previous loop cache designs are not suitable for all applications and system scenarios. In this paper, we present an adaptive loop cache that is amenable to a wide range of system scenarios, providing an additional 20% average instruction memory hierarchy energy savings (with individual benchmark energy savings as high as 69%) compared to the best previous loop cache design.
Marisha Rawlins, Ann Gordon-Ross
ACM Great Lakes Symposium on VLSI2
2010 A lightweight dynamic optimization methodology for wireless sensor networks
abstract
Technological advancements in embedded systems due to Moore's law have lead to the proliferation of wireless sensor networks (WSNs) in different application domains (e.g. defense, health care, surveillance systems) with different application requirements (e.g. lifetime, reliability). Many commercial-off-the-shelf (COTS) sensor nodes can be specialized to meet these requirements using tunable parameters (e.g. voltage, frequency) to specialize the operating state. Since a sensor node's performance depends greatly on environmental stimuli, dynamic optimizations enable sensor nodes to automatically determine their operating state in-situ. However, dynamic optimization methodology development given a large design space and resource constraints (memory and computational) is a very challenging task. In this paper, we propose a lightweight dynamic optimization methodology that intelligently selects initial tunable parameter values to produce a high-quality initial operating state in one-shot for time-critical or highly constrained applications. Further operating state improvements are made using an efficient greedy exploration algorithm, achieving optimal or near-optimal operating states while exploring only 0.04% of the design space on average.
Arslan Munir, Ann Gordon-Ross, Susan Lysecky, Roman L. Lysecky
WiMob2
2010 SIP-Based IMS Signaling Analysis for WiMax-3G Interworking Architectures
abstract
The third-generation partnership project (3GPP) and 3GPP2 have standardized the IP multimedia subsystem (IMS) to provide ubiquitous and access network-independent IP-based services for next-generation networks via merging cellular networks and the Internet. The application layer Session Initiation Protocol (SIP), standardized by 3GPP and 3GPP2 for IMS, is responsible for IMS session establishment, management, and transformation. The IEEE 802.16 worldwide interoperability for microwave access (WiMax) promises to provide high data rate broadband wireless access services. In this paper, we propose two novel interworking architectures to integrate WiMax and third-generation (3G) networks. Moreover, we analyze the SIP-based IMS registration and session setup signaling delay for 3G and WiMax networks with specific reference to their interworking architectures. Finally, we explore the effects of different WiMax-3G interworking architectures on the IMS registration and session setup signaling delay.
Arslan Munir, Ann Gordon-Ross
IEEE Trans. Mob. Comput.2
2009 Bitstream relocation with local clock domains for partially reconfigurable FPGAs
abstract
Partial Reconfiguration (PR) of FPGAs presents many opportunities for application design flexibility, enabling tasks to dynamically swap in and out of the FPGA without entire system interruption. However, mapping a task to any available PR region (PRR) requires a unique partial bitstream for each PRR. This replication can introduce significant overheads in terms of bitstream storage and communication requirements. Previous research in partial bitstream relocation can alleviate these overheads by transforming a single partial bitstream to map to any available PRR. However, careful steps are necessary to ensure proper functionality of relocated partial bitstreams and may result in clock routing inefficiencies. These routing inefficiencies can be alleviated by using regional clock resources introduced in the Virtex-4 FPGAs to implement local clock domains. PRRs can internally drive local clock domains, enabling each PRR to vary its clock frequency with respect to a single global clock signal, as opposed to sending multiple global clock signals (one for each desired clock frequency) to each PRR. We introduce this novel local clock domain (LCD) concept, which provides enhanced PR design flexibility. However, integration of LCDs and partial bitstream relocation introduces new challenges. In this paper, we identify motivating application domains for this integration, analyze integration benefits, and provide a detailed integration methodology.
Adam Flynn, Ann Gordon-Ross, Alan D. George
DATE2
2009 SCORES: A scalable and parametric streams-based communication architecture for modular reconfigurable systems
abstract
Parallel architectures have become an increasingly popular method in which to achieve high performance with low power consumption. In order to leverage these benefits, applications are decomposed into multiple computational modules (tasks) that collectively operate and communicate in parallel. In this paper, we present a scalable and highly parametric streams-based communication architecture for inter-module communication for FPGA-based systems - SCORES. This communication architecture improves on previous methods by providing increased application specialization and heterogeneous module clock frequencies, as well as providing a means for low latency communication and data throughput guarantees.
Abelardo Jara-Berrocal, Ann Gordon-Ross
DATE2
2009 Exploiting Partially Reconfigurable FPGAs for Situation-Based Reconfiguration in Wireless Sensor Networks
abstract
Wireless sensor networks (WSNs) are typically composed of very small, battery-operated devices (sensor nodes) containing simple microprocessors with few computational resources. However, the rapidly increasing popularity of WSNs has placed increased computational demands upon these systems, due to increasingly complex operating environments and enhanced data-sensing technology. Whereas introducing more powerful microprocessors into sensor nodes addresses these demands, sensor nodes do not contain sufficient energy reserves to support these microprocessors. In this paper, we present a partially reconfigurable FPGA-based architecture and methodology to provide increased WSN flexibility and computational resources, resulting in superior power consumption and performance compared to a microprocessor capable of satisfying similar demands.
Rafael García, Ann Gordon-Ross, Alan D. George
FCCM2
2009 Macs: A Minimal Adaptive routing circuit-switched architecture for scalable and parametric NoCs
abstract
Networks-on-chips (NoCs) are an emerging communication topology paradigm in single chip VLSI design, enhancing parallelism and system scalability. Processing units (PUs) connect to the communication topology via routers, which are responsible for runtime establishment and management of inter-PU communication channels. Router design directly affects overall system performance and exploited parallelism. In this paper, we present a highly parametric NoC architecture, MACS, providing increased system speed, designer flexibility, and scalability as compared to previous methods. In addition, MACS enhances inter-PU communication using a circuit-switching technique with dedicated, high frequency communication channels. Compared to previous work, MACS offers a 5x increase in operating frequency and a 2x reduction in area overhead.
Ann Gordon-Ross
FPL2
2009 Fast Configurable-Cache Tuning With a Unified Second-Level Cache
abstract
Tuning a configurable cache subsystem to an application can greatly reduce memory hierarchy energy consumption. Previous tuning methods use a level one configurable cache only, or a second level with separate instruction and data configurable caches. We instead use a commercially-common unified second level cache, a seemingly minor difference that actually expands the configuration space from 500 to about 20 000. We develop additive way tuning for tuning a cache subsystem with this large space, yielding 61% energy savings and 9% performance improvements over a nonconfigurable cache, greatly outperforming an extension of a previous method.
Ann Gordon-Ross, Frank Vahid, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.1
2008 Phase-based cache reconfiguration for a highly-configurable two-level cache hierarchy
abstract
Phase-based tuning methodologies specialize system parameters for each application phase of execution. Parameters are varied during execution, as opposed to remaining fixed as in an application-based tuning methodology. Prior work and logic suggests phase-based tuning may provide significant savings over application-based tuning. We investigate this hypothesis using a detailed cache model and tune a highly-configurable cache on a per-phase basis compared to tuning once per application, and found phase-based tuning to yield improvements of up to 37% in performance and 20% in energy over application-based tuning. Furthermore, we extend previous phase-based tuning of a configurable cache by significantly increasing configurability and show 14% energy improvement compared to previous methods. In addition, we quantify the overhead imposed due to cache reconfiguration.
Ann Gordon-Ross, Jeremy Lau, Brad Calder
ACM Great Lakes Symposium on VLSI1
2008 A table-based method for single-pass cache optimization
abstract
Due to the large contribution of the memory subsystem to total system power, the memory subsystem is highly amenable to customization for reduced power/energy and/or improved performance. Cache parameters such as total size, line size, and associativity can be specialized to the needs of an application for system optimization. In order to determine the best values for cache parameters, most methodologies utilize repetitious application execution to individually analyze each configuration explored. In this paper we propose a simplified yet efficient technique to accurately estimate the miss rate of many different cache configurations in just one single-pass of execution. The approach utilizes simple data structures in the form of a multi-layered table and elementary bitwise operations to capture the locality characteristics of an application's addressing behavior. The proposed technique intends to ease miss rate estimation and reduce cache exploration time.
Pablo Viana, Ann Gordon-Ross, Edna Barros, Frank Vahid
ACM Great Lakes Symposium on VLSI2
2008 A resource efficient content inspection system for next generation Smart NICs
abstract
The aggregate power consumption of the Internet is increasing at an alarming rate, due in part to the rapid increase in the number of connected edge devices such as desktop PCs. Despite being left idle 75% of the time, 90% of PCs have their power management features disabled. Consequently, much recent research has focused on reducing power consumption of Internet edge devices. One such method for reducing PC power consumption is by augmenting the network interface card (NIC) with enhanced processing capabilities. These capabilities pave the way for green computing by allowing the PC to transition to a low-power sleep state while the NIC responds to network traffic on behalf of the PC - a technique known as power proxying. However, such a Smart-NIC (SNIC) requires specialized low-power, resource-constrained processing, and architectural features in order to realize such capabilities. In this paper, we present a NIC-based packet content inspection system for power proxying and network intrusion detection. We use a novel partitioned TCAM technique that results in 87% energy savings and a 62% lower energy-delay product than existing non-partitioned router-based techniques, thus making our technique highly suitable for SNIC-based deployment.
Karthik Sabhanatarajan, Ann Gordon-Ross
ICCD2
2008 Real-time performance analysis of Adaptive Link Rate
abstract
High speed links are widely deployed in modern day computer networks to meet the ever growing needs for increasing data bandwidth. However, with the increase in the link rate, the power consumption of the network interfaces increases exponentially, compounding growing concerns about network power consumption. Fortunately, network traffic characteristics show that rapid link rates are not always required. During times of reduced network traffic, the Adaptive Link Rate (ALR) mechanism allows link rates to be reduced with little impact on network performance. Current research has focused on policies to control when and how to change link rates, and have shown promising energy savings. However, these works have been largely simulative, and have not addressed many of the challenges involved in implementation. In this paper, we develop a hardware prototype ALR system and address real-time challenges involved in realizing such an implementation. We also identify new considerations for control policy development given current technology capabilities as well as future projections.
Baoke Zhang, Karthik Sabhanatarajan, Ann Gordon-Ross, Alan D. George
LCN3
2007 A Self-Tuning Configurable Cache
abstract
The memory hierarchy of a system can consume up to 50% of microprocessor system power. Previous work has shown that tuning a configurable cache to a particular application can reduce memory subsystem energy by 62% on average. We introduce a self-tuning cache that performs transparent runtime cache tuning, thus relieving the application designer and/or compiler from predetermining an application's cache configuration. The self-tuning cache applies tuning at a determined tuning interval. A good interval balances tuning process energy overhead against the energy overhead of running in a sub-optimal cache configuration, which we show wastes much energy. We present a self-tuning cache that dynamically varies the tuning interval, resulting in average energy reduction of as much as 29%, falling within 13% of an oracle-based optimal method.
Ann Gordon-Ross, Frank Vahid
DAC1
2007 A one-shot configurable-cache tuner for improved energy and performance
abstract
We introduce a new non-intrusive on-chip cache-tuning hardware module capable of accurately predicting the best configuration of a configurable cache for an executing application. Previous dynamic cache tuning approaches change the cache configuration several times as part of the tuning search process, executing the application using inferior configurations and temporarily causing energy and performance overhead. The introduced tuner uses a different approach, which non-intrusively collects data on addresses issued by the microprocessor, analyzes that data to predict the best cache configuration, and then updates the cache to the new best configuration in "one-shot", without ever having to examine inferior configurations. The result is less energy and less performance overhead, meaning that cache tuning can be applied more frequently. We show through experiments that the one-shot cache tuner can reduce memory-access related energy for instructions by 35% and comes within 4% of a previous intrusive approach, and results in 4.6 times less energy overhead and a 7.7 times speedup in tuning time compared to a previous intrusive approach, at the main expense of 12% larger size
Ann Gordon-Ross, Pablo Viana, Frank Vahid, Walid A. Najjar, Edna Barros
DATE1
2006 Configurable cache subsetting for fast cache tuning
abstract
Numerous variations of configurable caches, having variable parameters like total size, line size, and associativity, have been proposed in commercial microprocessors in recent years. Tuning a configurable cache to a target application has been shown to reduce memory-access power by over 50%. However, searching the configuration space for the best configuration can require much time or power, even when using recent cache tuning heuristics. We sought to determine, for a particular domain of applications, the smallest subset of cache configurations that would still enable effective tuning. For a suite of 34 benchmarks and a cache with 18 possible configurations, we determine through an exhaustive search of all possible subsets, that only 3 or 4 candidate configurations are necessary to support tuning. We introduce a new heuristic, adapted from an efficient and effective heuristic developed for data mining, to quickly determine the best configurations for any sized subset, with near optimal results. We then consider a configurable cache with 17,640 possible configurations and improve our heuristic to include a pre-pruning step, yielding near optimal tuning results. We conclude that only 3 or 4 possible cache configurations are needed to offer a near optimal configuration for every benchmark in our suite - resulting in a 91% reduction in design space exploration time over a state-of-the-art cache tuning heuristic.
Pablo Viana, Ann Gordon-Ross, Eamonn J. Keogh, Edna Barros, Frank Vahid
DAC2
2005 A first look at the interplay of code reordering and configurable caches
abstract
The instruction cache is a popular target for optimizations of microprocessor-based systems because of the cache's high impact on system performance and power, and because of the cache's predictable temporal and spatial locality. Optimization techniques can be designed based on this predictability. We explore for the first time the interplay of two popular instruction cache optimization techniques: the long-known technique of code reordering and the relatively-new technique of cache configuration. We address the question of whether those two optimizations complement each other or if one optimization dominates the other. Through experiments using embedded system benchmarks, we show that cache configuration dominates a particular category of code reordering techniques with respect to optimizing performance and energy, obviating the need for reordering. We also examine the modern scenario of synthesized custom caches, and show that combining cache configuration with code reordering results in cache size reductions of 13% on average, and up to 89% in some benchmarks, beyond just cache configuration alone.
Ann Gordon-Ross, Frank Vahid, Nikil Dutt
ACM Great Lakes Symposium on VLSI1
2005 Fast configurable-cache tuning with a unified second-level cache
abstract
Tuning a configurable cache subsystem to an application can greatly reduce memory hierarchy energy consumption. Previous tuning methods use a level one configurable cache only, or a second level with separate instruction and data configurable caches. We instead use a commercially-common unified second level, a seemingly minor difference that actually expands the configuration space from 500 to about 20,000. We develop additive way tuning for tuning a cache subsystem with this large space, yielding 62% energy savings and 35% performance improvements over a non-configurable cache, greatly outperforming an extension of a previous method
Ann Gordon-Ross, Frank Vahid, Nikil Dutt
ISLPED1
2005 Frequent Loop Detection Using Efficient Nonintrusive On-Chip Hardware
abstract
Dynamic software optimization methods are becoming increasingly popular for improving software performance and power. The first step in dynamic optimization consists of detecting frequently executed code, or "critical regions." Most previous critical region detectors have been targeted to desktop processors. We introduce a critical region detector targeted to embedded processors, with the unique features of being very size and power efficient and being completely nonintrusive to the software's execution-features needed in timing-sensitive embedded systems. Our detector not only finds the critical regions, but also determines their relative frequencies, a potentially important feature for selecting among alternative dynamic optimization methods. Our detector uses a tiny cache-like structure coupled with a small amount of logic. We provide results of extensive explorations across 19 embedded system benchmarks. We show that highly accurate results can be achieved with only a 0.02 percent power overhead, acceptable size overhead; and zero runtime overhead. Our detector is currently being used as part of a dynamic hardware/software partitioning approach, but is applicable to a wide variety of situations.
Ann Gordon-Ross, Frank Vahid
IEEE Trans. Computers1
2004 Automatic Tuning of Two-Level Caches to Embedded Applications
abstract
The power consumed by the memory hierarchy of a microprocessor can contribute to as much as 50% of the total microprocessor system power, and is thus a good candidate for optimizations. We present an automated method for tuning two-level caches to embedded applications for reduced energy consumption. The method is applicable to both a simulation-based exploration environment and a hardware-based system prototyping environment. We introduce the two-level cache tuner, or TCaT - a heuristic for searching the huge solution space of possible configurations. The heuristic interlaces the exploration of the two cache levels and searches the various cache parameters in a specific order based on their impact on energy. We show the integrity of our heuristic across multiple memory configurations and even in the presence of hardware/software partitioning - a common optimization capable of achieving significant speedups and/or reduced energy consumption. We apply our exploration heuristic to a large set of embedded applications. Our experiments demonstrate the efficacy of our heuristic: on average the heuristic examines only 7% of the possible cache configurations, but results in cache sub-system energy savings of 53%, only 1% more than the optimal cache configuration. In addition, the configured cache achieves an average speedup of 30% over the base cache configuration due to tuning of cache line size to the application's needs.
Ann Gordon-Ross, Frank Vahid, Nikil Dutt
DATE1
2003 Frequent loop detection using efficient non-intrusive on-chip hardware
abstract
Dynamic software optimization methods are becoming increasingly popular for improving software performance and power. The first step in dynamic optimization consists of detecting frequently executed code, or "critical regions." Previous critical region detectors have been targeted to desktop processors. We introduce a critical region detector targeted to embedded processors, with the unique features of being very size and power efficient, and being completely non-intrusive to the software's execution - features needed in timing-sensitive embedded systems. Our detector not only finds the critical regions, but also determines their relative frequencies, a potentially important feature for selecting among alternative dynamic optimization methods. Our detector uses a tiny cache coupled with a small amount of logic. We provide results of extensive explorations across seventeen embedded system benchmarks. We show that highly accurate results can be achieved with only a 0.02% power overhead and acceptable size overhead. Our detector is currently being used as part of a dynamic hardware/software partitioning approach, but is applicable to a wide-variety of situations.
Ann Gordon-Ross, Frank Vahid
CASES1
2003 Tiny instruction caches for low power embedded systems
abstract
Instruction caches have traditionally been used to improve software performance. Recently, several tiny instruction cache designs, including filter caches and dynamic loop caches, have been proposed to instead reduce software power. We propose several new tiny instruction cache designs, including preloaded loop caches, and one-level and two-level hybrid dynamic/preloaded loop caches. We evaluate the existing and proposed designs on embedded system software benchmarks from both the Powerstone and MediaBench suites, on two different processor architectures, for a variety of different technologies. We show on average that filter caching achieves the best instruction fetch energy reductions of 60--80%, but at the cost of about 20% performance degradation, which could also affect overall energy savings. We show that dynamic loop caching gives good instruction fetch energy savings of about 30%, but that if a designer is able to profile a program, preloaded loop caching can more than double the savings. We describe automated methods for quickly determining the best loop cache configuration, methods useful in a core-based design flow.
Ann Gordon-Ross, Susan Cotterell, Frank Vahid
ACM Trans. Embed. Comput. Syst.1
2002 Dynamic Loop Caching Meets Preloaded Loop Caching - A Hybrid Approach
abstract
Dynamically-loaded tagless loop caching reduces instruction fetch power for embedded software with small loops, but only supports simple loops without taken branches. Preloaded tagless loop caching supports complex loops with branches and thus can reduce power further, but has a limit on the total number of instructions cached. We show that each does well on particular benchmarks, but neither is best across all of those benchmarks. We present a new hybrid loop cache that only preloads the complex loops, while dynamically loading other loops, thus achieving the strengths of each approach. We demonstrate better power savings than either previous approach alone.
Ann Gordon-Ross, Frank Vahid
ICCD1
2001 A self-optimizing embedded microprocessor using a loop table for low power
abstract
We describe an approach for a microprocessor to tune itself to its fixed application to reduce power in an embedded system. We define a basic architecture and methodology supporting a microprocessor self-optimizing mode. We also introduce a loop table as a tunable component, although self-optimization can be done for other tunable components too. We highlight experimental results illustrating good power reductions with no performance penalty. Keywords System-on-a-chip, self-optimizing architecture, embedded systems, parameterized architectures, cores, low-power, tuning, platforms.
Frank Vahid, Ann Gordon-Ross
ISLPED2