EDBT 2026 Demo / reviewers in the wild / expert
Ann Gordon-Ross
dblp:47/5909
· DBLP profile ↗
73ranked-venue papers
12as first author
0since 2021 · last 2019
0000-0001-8865-8381ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 62 · 12 first-authorSoftware engineering, systems software and programming languages · 9 · 2 first-authorComputer networks · 5Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
Processor architecture and microarchitecture · 23% Embedded and real-time systems · 22% Memory systems · 20% | |
| Computer networks
4 papers |
Edge and fog computing · 37% Cellular and mobile networks · 37% Internet of things and sensor networks · 22% |
Topics — the 27 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Embedded and real-time systems › embedded hardware platform
multicore embedded systems |
0.3 | 2 | 2014 | Multi-Core Embedded Wireless Sensor Networks: Architecture and Applications · IEEE Trans. Parallel Distributed Syst. 2014 High-Performance Energy-Efficient Multicore Embedded Computing · IEEE Trans. Parallel Distributed Syst. 2012 |
Edge and fog computing
iot edge computing |
0.3 | 1 | 2018 | Microprocessor Optimizations for the Internet of Things: A Survey · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Memory systems
cache |
0.2 | 2 | 2013 | A Cache Tuning Heuristic for Multicore Architectures · IEEE Trans. Computers 2013 A Self-Tuning Configurable Cache · DAC 2007 |
Processor architecture and microarchitecture
multicore design |
0.2 | 2 | 2014 | Multi-Core Embedded Wireless Sensor Networks: Architecture and Applications · IEEE Trans. Parallel Distributed Syst. 2014 High-Performance Energy-Efficient Multicore Embedded Computing · IEEE Trans. Parallel Distributed Syst. 2012 |
Embedded and real-time systems › wireless communication
wireless sensor networks |
0.2 | 1 | 2014 | Multi-Core Embedded Wireless Sensor Networks: Architecture and Applications · IEEE Trans. Parallel Distributed Syst. 2014 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.2 | 2 | 2013 | A Cache Tuning Heuristic for Multicore Architectures · IEEE Trans. Computers 2013 A Self-Tuning Configurable Cache · DAC 2007 |
Performance modeling and evaluation › simulation
cache simulation |
0.2 | 1 | 2013 | T-SPaCS - A Two-Level Single-Pass Cache Simulation Methodology · IEEE Trans. Computers 2013 |
Memory systems › cache management
cache tuning |
0.2 | 1 | 2013 | A Cache Tuning Heuristic for Multicore Architectures · IEEE Trans. Computers 2013 |
Performance modeling and evaluation
simulation |
0.2 | 1 | 2013 | T-SPaCS - A Two-Level Single-Pass Cache Simulation Methodology · IEEE Trans. Computers 2013 |
Internet of things and sensor networks
wireless sensor network |
0.1 | 1 | 2012 | An MDP-Based Dynamic Optimization Methodology for Wireless Sensor Networks · IEEE Trans. Parallel Distributed Syst. 2012 |
Electronic design automation
hardware/software co-design |
0.1 | 1 | 2012 | High-Performance Energy-Efficient Multicore Embedded Computing · IEEE Trans. Parallel Distributed Syst. 2012 |
Cellular and mobile networks › low-latency communication
connection setup latency |
0.1 | 1 | 2010 | SIP-Based IMS Signaling Analysis for WiMax-3G Interworking Architectures · IEEE Trans. Mob. Comput. 2010 |
Cellular and mobile networks › mobile networks › mobile network architecture › mobile core network
IP multimedia subsystem |
0.1 | 1 | 2010 | SIP-Based IMS Signaling Analysis for WiMax-3G Interworking Architectures · IEEE Trans. Mob. Comput. 2010 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2018 | Microprocessor Optimizations for the Internet of Things: A Survey · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Reconfigurable computing and FPGAs
FPGA implementation |
0.1 | 1 | 2016 | MACS: A Highly Customizable Low-Latency Communication Architecture · IEEE Trans. Parallel Distributed Syst. 2016 |
Memory systems › cache › cache organization
configurable cache |
0.1 | 1 | 2007 | A Self-Tuning Configurable Cache · DAC 2007 |
Memory systems
cache design |
0.1 | 1 | 2006 | Configurable cache subsetting for fast cache tuning · DAC 2006 |
Electronic design automation
design space exploration |
0.1 | 1 | 2006 | Configurable cache subsetting for fast cache tuning · DAC 2006 |
Internet of things and sensor networks › wireless sensor network
in-network processing |
0.1 | 1 | 2014 | Multi-Core Embedded Wireless Sensor Networks: Architecture and Applications · IEEE Trans. Parallel Distributed Syst. 2014 |
Processor architecture and microarchitecture
dynamic optimization |
0.1 | 1 | 2005 | Frequent Loop Detection Using Efficient Nonintrusive On-Chip Hardware · IEEE Trans. Computers 2005 |
Embedded and real-time systems
embedded processor |
0.1 | 1 | 2005 | Frequent Loop Detection Using Efficient Nonintrusive On-Chip Hardware · IEEE Trans. Computers 2005 |
Embedded and real-time systems › runtime monitoring
non-intrusive profiling |
0.1 | 1 | 2005 | Frequent Loop Detection Using Efficient Nonintrusive On-Chip Hardware · IEEE Trans. Computers 2005 |
Memory systems › memory hierarchy
cache hierarchy |
0.0 | 1 | 2013 | T-SPaCS - A Two-Level Single-Pass Cache Simulation Methodology · IEEE Trans. Computers 2013 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2013 | A Cache Tuning Heuristic for Multicore Architectures · IEEE Trans. Computers 2013 |
Energy-efficient computing
energy management |
0.0 | 1 | 2012 | An MDP-Based Dynamic Optimization Methodology for Wireless Sensor Networks · IEEE Trans. Parallel Distributed Syst. 2012 |
Internet architecture and protocols › signaling protocol
SIP |
0.0 | 1 | 2010 | SIP-Based IMS Signaling Analysis for WiMax-3G Interworking Architectures · IEEE Trans. Mob. Comput. 2010 |
Energy-efficient computing
power management |
0.0 | 1 | 2007 | A Self-Tuning Configurable Cache · DAC 2007 |
Methods — techniques the papers use, named apart from their topics
workload characterization · 0.7microarchitectural survey · 0.7performance comparison · 0.4parallelization · 0.4trace simulation · 0.2path resolution algorithm · 0.2trace-driven simulation · 0.2single-pass simulation · 0.2hardware cache tuner · 0.2design space search heuristic · 0.2markov decision process · 0.1signaling delay analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Dynamic Scheduling on Heterogeneous MulticoresabstractHeterogeneous multicore systems help meet design goals by using disparate hardware components that are suitable for different application requirements/design goals. The individual cores may also have different tunable hardware parameters for additional specialization. However, this complicates scheduling since to reap the benefits of specialization, applications should be scheduled to the core that offers the best configuration based on the application's requirements and design goals. This scheduling decision could be made by exploring the design space to evaluate different configurations to determine the best configuration, or by executing the application in a base configuration to gather execution statistics to predict the best configuration. However, given increasingly complex systems, these methods may be infeasible given extremely large design spaces or difficulty in choosing a representative base configuration. In this paper, we present a dynamic scheduling methodology that uses predictive methods to schedule applications to best configurations for reduced energy consumption for a system with configurable caches. We use an artificial neural network (ANN) to train our predictive model using hardware counters. The trained ANN can then be used to predict the best core and a tuning heuristic explores the design space to determine the best configuration on non-best cores. If the best core is busy, our scheduler considers alternative idle cores or the application is stalled depending on which decision is energy advantageous. Our experiments show that system energy can be reduced by 28% on average as compared to a fixed-core system where all cores offer the same configuration. Ayobami S. Edun, Ruben Vazquez, Ann Gordon-Ross, Greg Stitt |
DATE | 3 |
| 2019 | Accelerating Scientific Discovery with SCAIGATE Science GatewayabstractThe demand for computational accelerators (GPUs, FPGAs, ASICs, etc.) is growing due to the widening variety of datacenter applications fueled by recent scientific breakthroughs that leverage artificial intelligence (AI). As much as these applications (e.g., cosmology, physics, etc.) have continued to witness record-breaking accuracy in predictive capabilities due to AI widespread influence, the infrastructure and workflow to take these applications out of research labs into production and business use-cases continues to lag. To address these important infrastructural challenges, we present SCAIGATE, a prototype science gateway with a simplified workflow aimed at facilitating model building/validation workflows in large-scale scientific applications. David Ojika, Bhavesh Patel, Ann Gordon-Ross, Herman Lam |
eScience | 4 |
| 2019 | Energy Prediction for Cache Tuning in Embedded SystemsabstractModern embedded systems are longer tasked at operating a single application or function and are increasingly required to operate more like general purpose desktop computers. Conforming to modern usage demands is extremely challenging given an embedded system's stringent design constraints, such as power, energy, and performance. Adherence to these constraints can be achieved by specializing/tuning the underlying system to application-specific execution requirements and characteristics by tuning a system's configurable parameters to meet these requirements given design constraints. Configurable parameters include architectural voltage, frequency, cache size, line size, and associativity, etc. However, given the complexity of modern systems, exploring these large design spaces is infeasible when the number of configurable parameters and valid parameter values increases beyond a trivial amount. In this paper, we propose using machine learning in lieu of traditional design space exploration techniques. In this work, we evaluate the potential for using an artificial neural network (ANN)-based prediction module for energy prediction. Since the cache hierarchy has a large impact on total energy consumption, without loss of generality, we study a configurable cache hierarchy with configurable cache size, associativity, and line size. We design and train an energy prediction module to infer the best cache configuration for an application based on the application's execution characteristics. Our approach requires only a single profiling run of the application to collect these characteristics. Our energy prediction module then predicts the energy consumption for all the configurations in the cache design space based on these characteristics, and outputs the configuration with the lowest energy consumption, thus essentially performing exhaustive design space exploration with a single execution. Our results show that our prediction module predicts the best instruction and data cache configurations for the majority of the applications, yielding an average energy degradation of less than 2% for both the instruction and data caches as compared to the optimal configuration determined by exhaustive design space exploration. Ruben Vazquez, Ann Gordon-Ross, Greg Stitt |
ICCD | 2 |
| 2018 | Microprocessor Optimizations for the Internet of Things: A SurveyabstractThe Internet of Things (IoT) refers to a pervasive presence of interconnected and uniquely identifiable physical devices. These devices' goal is to gather data and drive actions in order to improve productivity, and ultimately reduce or eliminate reliance on human intervention for data acquisition, interpretation, and use. The proliferation of these connected low-power devices will result in a data explosion that will significantly increase data transmission costs with respect to energy consumption and latency. Edge computing reduces these costs by performing computations at the edge nodes, prior to data transmission, to interpret and/or utilize the data. While much research has focused on the IoT's connected nature and communication challenges, the challenges of IoT embedded computing with respect to device microprocessors has received much less attention. This paper explores IoT applications' execution characteristics from a microarchitectural perspective and the microarchitectural characteristics that will enable efficient and effective edge computing. To tractably represent a wide variety of next-generation IoT applications, we present a broad IoT application classification methodology based on application functions, to enable quicker workload characterizations for IoT microprocessors. We then survey and discuss potential microarchitectural optimizations and computing paradigms that will enable the design of right-provisioned microprocessors that are efficient, configurable, extensible, and scalable. This paper provides a foundation for the analysis and design of a diverse set of microprocessor architectures for next-generation IoT devices. Tosiron Adegbija, Anita Rogacs, Chandrakant Patel, Ann Gordon-Ross |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | A comparison-free sorting algorithm on CPUs and GPUs
Saleh Abdel-Hafeez, Ann Gordon-Ross, Samer Abubaker |
J. Supercomput. | 2 |
| 2018 | PhLock: A Cache Energy Saving Technique Using Phase-Based Cache LockingabstractCaches are commonly used to bridge the processor-memory performance gap in embedded systems. Since embedded systems typically have stringent design constraints imposed by physical size, battery capacity, and real-time deadlines much research focuses on cache optimizations, such as improved performance and/or reduced energy consumption. Cache locking is a popular cache optimization that loads and retains/locks selected memory contents from an executing application into the cache to increase the cache's predictability. Previous work has shown that cache locking also has the potential to improve cache energy consumption. In this paper, we introduce phase-based cache locking, PhLock, which leverages an application's varying runtime characteristics to dynamically select the locked memory contents to optimize cache energy consumption. Using a variety of applications from the SPEC2006 and MiBench benchmark suites, experimental results show that PhLock is promising for reducing both the instruction and data caches' energy consumption. As compared to a nonlocking cache, PhLock reduced the instruction and data cache energy consumption by an average of 5% and 39%, respectively, for SPEC2006 applications, and by 75% and 14%, respectively, for MiBench benchmarks. Tosiron Adegbija, Ann Gordon-Ross |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Optimizing FPGA Performance, Power, and Dependability with Linear ProgrammingabstractField-programmable gate arrays (FPGA) are an increasingly attractive alternative to traditional microprocessor-based computing architectures in extreme-computing domains, such as aerospace and supercomputing. FPGAs offer several resource types that offer different tradeoffs between speed, power, and area, which make FPGAs highly flexible for varying application computational requirements. However, since an application’s computational operations can map to different resource types, a major challenge in leveraging resource-diverse FPGAs is determining the optimal distribution of these operations across the device’s available resources for varying FPGA devices, resulting in an extremely large design space. In order to facilitate fast design-space exploration, this article presents a method based on linear programming (LP) that determines the optimal operation distribution for a particular device and application with respect to performance, power, or dependability metrics. Our LP method is an effective tool for exploring early designs by quickly analyzing thousands of FPGAs to determine the best FPGA devices and operation distributions, which significantly reduces design time. We demonstrate our LP method’s effectiveness with two case studies involving dot-product and distance-calculation kernels on a range of Virtex-5 FPGAs. Results show that our LP method selects optimal distributions of operations to within an average of 4% of actual values. Nicholas Wulf, Alan D. George, Ann Gordon-Ross |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2017 | An Efficient O(N) Comparison-Free Sorting AlgorithmabstractIn this paper, we propose a novel sorting algorithm that sorts input data integer elements on-the-fly without any comparison operations between the data-comparison-free sorting. We present a complete hardware structure, associated timing diagrams, and a formal mathematical proof, which show an overall sorting time, in terms of clock cycles, that is linearly proportional to the number of inputs, giving a speed complexity on the order of O(N). Our hardware-based sorting algorithm precludes the need for SRAM-based memory or complex circuitry, such as pipelining structures, but rather uses simple registers to hold the binary elements and the elements' associated number of occurrences in the input set, and uses matrix-mapping operations to perform the sorting process. Thus, the total transistor count complexity is on the order of O(N). We evaluate an application-specified integrated circuit design of our sorting algorithm for a sample sorting of N = 1024 elements of size K = 10-bit using 90-nm Taiwan Semiconductor Manufacturing Company (TSMC) technology with a 1 V power supply. Results verify that our sorting requires approximately 4-6 μs to sort the 1024 elements with a clock cycle time of 0.5 GHz, consumes 1.6 mW of power, and has a total transistor count of less than 750 000. Saleh Abdel-Hafeez, Ann Gordon-Ross |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Configuration prefetching and reuse for preemptive hardware multitasking on partially reconfigurable FPGAs
Aurelio Morales-Villanueva, Ann Gordon-Ross |
DATE | 3 |
| 2016 | Quality of Service-Aware, Scalable Cache Tuning Algorithm in Consumer-based Embedded DevicesabstractTo meet energy and quality of service (QoS) constraints in consumer-based embedded devices (CEDs), configurable caches can be tuned to a best configuration that consumes the least amount of energy while adhering to QoS expectations. However, due to disparate consumer QoS expectations and a myriad of unknown, third-party CED applications, tuning caches in CEDs is very challenging. In this paper, we propose a quality of service-aware, scalable tuning algorithm for configurable caches, which requires no a priori knowledge of applications or design-time efforts. Mohamad Hammam Alsafrjalani, Ann Gordon-Ross |
ACM Great Lakes Symposium on VLSI | 2 |
| 2016 | MACS: A Highly Customizable Low-Latency Communication ArchitectureabstractNetworks-on-chips (NoCs) are an increasingly popular communication infrastructure in single chip VLSI design for enhancing parallelism and system scalability. Processing elements (PEs) connect to a communication topology via NoC switches, which are responsible for runtime establishment and management of inter-PE communication channels. Since NoC switch design directly affects overall system performance and exploited communication parallelism, much previous work focused on efficient NoC switch design. In this paper, we present MACS-a highly parametric NoC switch architecture that provides reduced data transfer latency, increased designer flexibility, and scalability as compared to previous architectures by combining and enhancing several NoC design strategies. MACS enhances inter-PE communication using a circuit switching technique with minimal adaptive routing and a simple and fair path resolution algorithm to maximize bandwidth utilization. We evaluate area and performance of an FPGA implementation of MACS, and, show that compared to previous work, MACS offers a 2× to 7× decrease in average channel setup latency, a 1.7× to 2× reduction in area requirements, similar average packet latency, up to a 6× increase in the network saturation point, and up to a 1.4× increase in bandwidth utilization. Additionally, we illustrate MACS's low average channel setup latency using six network traffic patterns and eight parallel JPEG decompression core trace simulations. Ann Gordon-Ross |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | A Framework for Evaluating and Optimizing FPGA-Based SoCs for Aerospace ComputingabstractOn-board processing systems are often deployed in harsh aerospace environments and must therefore adhere to stringent constraints such as low power, small size, and high dependability in the presence of faults. Field-programmable gate arrays (FPGAs) are often an attractive option for designers seeking low-power, high-performance devices. However, unlike nonreconfigurable devices, radiation effects can alter an FPGA’s functionality instead of just the device’s data, requiring designers to consider fault-tolerant strategies to mitigate these effects. In this article, we present a framework to ease these system design challenges and aid designers in considering a broad range of devices and fault-tolerant strategies for on-board processing, highlighting the most promising options and tradeoffs early in the design process. This article focuses on the power, dependability, and lifetime evaluation metrics, which our framework calculates and leverages to evaluate the effectiveness of varying system-on-chip (SoC) designs. Finally, we use our framework to evaluate SoC designs for a case study on a hyperspectral-imaging (HSI) mission to demonstrate our framework’s ability to identify efficient and effective SoC designs. Nicholas Wulf, Alan D. George, Ann Gordon-Ross |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2015 | Phase-based Cache Locking for Embedded SystemsabstractSince caches are commonly used in embedded systems, which typically have stringent design constraints imposed by physical size, battery capacity, real-time deadlines, etc., much research focuses on cache optimizations, such as improved performance and/or reduced energy consumption. Cache locking is a popular cache optimization that loads and retains/locks selected memory contents from an executing application into the cache to increase the cache's predictability. Previous work has shown that cache locking also has the potential to improve cache performance and energy consumption. In this paper, we introduce phase-based cache locking, which leverages an application's varying runtime characteristics to dynamically select the locked memory contents to optimize cache performance and energy consumption. Experimental results show that our phase-based cache locking methodology can improve the data cache's miss rates and energy consumption by an average of 24% and 20%, respectively. Tosiron Adegbija, Ann Gordon-Ross |
ACM Great Lakes Symposium on VLSI | 2 |
| 2015 | Modeling and Analysis of Fault Detection and Fault Tolerance in Wireless Sensor NetworksabstractTechnological advancements in communications and embedded systems have led to the proliferation of Wireless Sensor Networks (WSNs) in a wide variety of application domains. These application domains include but are not limited to mission-critical (e.g., security, defense, space, satellite) or safety-related (e.g., health care, active volcano monitoring) systems. One commonality across all WSN application domains is the need to meet application requirements (e.g., lifetime, reliability). Many application domains require that sensor nodes be deployed in harsh environments, such as on the ocean floor or in an active volcano, making these nodes more prone to failures. Sensor node failures can be catastrophic for critical or safety-related systems. This article models and analyzes fault detection and fault tolerance in WSNs. To determine the effectiveness and accuracy of fault detection algorithms, we simulate these algorithms using ns-2. We investigate the synergy between fault detection and fault tolerance and use the fault detection algorithms’ accuracies in our modeling of Fault-Tolerant (FT) WSNs. We develop Markov models for characterizing WSN reliability and Mean Time to Failure (MTTF) to facilitate WSN application-specific design. Results obtained from our FT modeling reveal that an FT WSN composed of duplex sensor nodes can result in as high as a 100% MTTF increase and approximately a 350% improvement in reliability over a Non-Fault-Tolerant (NFT) WSN. The article also highlights future research directions for the design and deployment of reliable and trustworthy WSNs. Arslan Munir, Joseph Antoon, Ann Gordon-Ross |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2014 | Dynamic Scheduling for Reduced Energy in Configuration-Subsetted Heterogeneous Multicore SystemsabstractHeterogeneous and configurable multicore systems provide hardware specialization to meet disparate application hardware requirements. However, effective multicore system specialization can require a priori knowledge of the applications, application profiling information, and/or dynamic hardware tuning to schedule and execute applications on the most energy efficient cores. Furthermore, even though highly disparate core heterogeneity and/or highly configurable parameters with numerous potential parameter values result in more fine-grained specialization and higher energy savings potential, these large design spaces are challenging to efficiently explore. To address these challenges, we propose a novel configuration-subsisted heterogeneous and configurable multicore system, wherein each core offers a small subset of the design space, and propose a novel scheduling and tuning (SaT) algorithm to efficiently exploit the energy savings potential of this system. Our proposed architecture and algorithm require no a priori application knowledge or profiling, and incurs minimal runtime overhead. Results reveal energy savings potential and insights on energy tradeoffs in heterogeneous, configurable systems. Mohamad Hammam Alsafrjalani, Ann Gordon-Ross |
EUC | 2 |
| 2014 | Minimum Effort Design Space Subsetting for Configurable CachesabstractConfigurable caches can significantly reduce energy consumption by adapting the system's cache configuration to the applications' specific requirements to meet system design and optimization goals. However, large configuration design spaces require prohibitive design space exploration time (e.g., due to lengthy design space analyses, simulations, and/or evaluations) to determine the best configuration given these requirements and goals. To significantly reduce design space exploration time, we evaluate a design space subsetting method that removes energy-redundant configurations (i.e., configurations that provide similar energy savings as other configurations), thus significantly reducing the design space while still providing high-quality, energy-saving configurations. Prior work verified design space subsetting's efficacy, however, prior work required extensive design-time effort and complete a priori knowledge of the system's anticipated applications. In this work, we alleviate these limitations and significantly broaden the usability of design space subsetting. Results show that complete a priori knowledge of the anticipated applications is not necessary, and only a small set of applications representative of the anticipated applications' general domains (or applications with similar requirements) is sufficient to provide energy savings within 5.6% of the complete, unsubsetted design space. Mohamad Hammam Alsafrjalani, Ann Gordon-Ross, Pablo Viana |
EUC | 2 |
| 2014 | Thermal-aware phase-based tuning of embedded systemsabstractDue to embedded systems' stringent design constraints, much prior work focused on optimizing energy consumption and/or performance. However, since embedded systems have fewer cooling options, rising temperature, and thus temperature optimization, is an emergent concern. We present thermal-aware phase-based tuning--TaPT--that determines Pareto optimal configurations for fine-grained execution time, energy, and temperature tradeoffs. Results show that TaPT reduces execution time, energy, and temperature by as much as 5%, 30%, and 25%, respectively, while adhering to designer-specified design constraints. Tosiron Adegbija, Ann Gordon-Ross |
ACM Great Lakes Symposium on VLSI | 2 |
| 2014 | Analysis of cache tuner architectural layouts for multicore embedded systemsabstractDue to the memory hierarchy's large contribution to a microprocessor's total power, cache tuning is an ideal method for optimizing overall power consumption in embedded systems. Since most embedded systems are power and area constrained, the hardware and/or software that orchestrate cache tuning - the cache tuner - must not impose significant power and area overhead. Furthermore, as embedded systems increasingly trend towards multicore, inter-core data sharing, communication, and synchronization impose additional cache tuner design complexity, necessitating cross-core cache tuning coordination. In order to minimize cache tuner overhead, cache tuner design must consider these overheads and scalability. Whereas prior work proposes low-overhead cache tuners, scalability to multicore systems requires additional considerations. In this work, we present a low-overhead, scalable cache tuner and extensively evaluate various cache tuner design tradeoffs with respect to power and area for constrained multicore embedded systems. Based on our analysis, we formulate valuable insights and designer-assisted guidelines for modeling scalable and efficient cache tuners that best achieve optimization goals while maintaining power and area constraints. Tosiron Adegbija, Ann Gordon-Ross, Marisha Rawlins |
IPCCC | 2 |
| 2014 | A queueing theoretic approach for performance evaluation of low-power multi-core embedded systems
Arslan Munir, Ann Gordon-Ross, Sanjay Ranka, Farinaz Koushanfar |
J. Parallel Distributed Comput. | 2 |
| 2014 | Multi-Core Embedded Wireless Sensor Networks: Architecture and ApplicationsabstractTechnological advancements in the silicon industry, as predicted by Moore's law, have enabled integration of billions of transistors on a single chip. To exploit this high transistor density for high performance, embedded systems are undergoing a transition from single-core to multi-core. Although a majority of embedded wireless sensor networks (EWSNs) consist of single-core embedded sensor nodes, multi-core embedded sensor nodes are envisioned to burgeon in selected application domains that require complex in-network processing of the sensed data. In this paper, we propose an architecture for heterogeneous hierarchical multi-core embedded wireless sensor networks (MCEWSNs) as well as an architecture for multi-core embedded sensor nodes used in MCEWSNs. We elaborate several compute-intensive tasks performed by sensor networks and application domains that would especially benefit from multi-core embedded sensor nodes. This paper also investigates the feasibility of two multi-core architectural paradigms-symmetric multiprocessors (SMPs) and tiled many-core architectures (TMAs)-for MCEWSNs. We compare and analyze the performance of an SMP (an Intel-based SMP) and a TMA (Tilera's TILEPro64) based on a parallelized information fusion application for various performance metrics (e.g., runtime, speedup, efficiency, cost, and performance per watt). Results reveal that TMAs exploit data locality effectively and are more suitable for MCEWSN applications that require integer manipulation of sensor data, such as information fusion, and have little or no communication between the parallelized tasks. To demonstrate the practical relevance of MCEWSNs, this paper also discusses several state-of-the-art multi-core embedded sensor node prototypes developed in academia and industry. We further discuss research challenges and future research directions for MCEWSNs. Arslan Munir, Ann Gordon-Ross, Sanjay Ranka |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2013 | PRML: A Modeling Language for Rapid Design Exploration of Partially Reconfigurable FPGAsabstractLeveraging partial reconfiguration (PR) can improve system flexibility, cost, and performance/power/area tradeoffs over non-PR functionally-equivalent systems, however, realizing these benefits is challenging, time-consuming, and PR must be considered early during application design to reduce design exploration time and improve system quality. To facilitate realizing these benefits, we present an application design framework and an abstract modeling language for PR (PRML). By applying extensive PRML modeling guidelines to a complex arithmetic core, we show PRML's potential for efficient PR capability analysis, enabling designers to determine Pareto optimal systems during application formulation based on designer-specified area and performance metrics. Ann Gordon-Ross |
FCCM | 2 |
| 2013 | On-chip Context Save and Restore of Hardware Tasks on Partially Reconfigurable FPGAsabstractPartial reconfiguration (PR) of field-programmable gate arrays (FPGAs) enables hardware tasks to time multiplex PR regions (PRRs) by isolating reconfiguration to only the reconfigured PRR, which avoids halting the entire FPGA's execution. Time multiplexing PRRs requires support for unloading/loading tasks and for resuming a task's execution state. In order to resume a task's execution state, the execution state (context) must be saved when the task is unloaded so that the execution state can be restored when the task resumes- context save (CS) and context restore (CR), respectively. In this paper, we present a software-based, on-chip context save and restore (CSR) for PR-capable FPGAs. As compared to prior work, our CSR is autonomous (i.e., does not require any external host support), does not require custom on-chip hardware, is portable across any system design, and does not require tool flow modifications or special tools. Experimental results extensively evaluate the CSR execution time based on PRR size, enabling designers to trade off PRR granularity for CSR execution time based on application requirements. Aurelio Morales-Villanueva, Ann Gordon-Ross |
FCCM | 2 |
| 2013 | Exploiting dynamic phase distance mapping for phase-based tuning of embedded systemsabstractPhase-based tuning increases optimization potential by configuring system parameters for application execution phases. Previous work proposed phase distance mapping (PDM), which relied on extensive a priori analysis of executing applications to dynamically estimate the best configuration using the correlation between phases. We propose DynaPDM, a new dynamic phase distance mapping methodology that eliminates a priori designer effort, dynamically analyzes phases, and determines the best configurations, yielding average energy delay product savings of 28%-an 8% improvement on PDM-and configurations within 1% of the optimal. Tosiron Adegbija, Ann Gordon-Ross |
ICCD | 2 |
| 2013 | A Cache Tuning Heuristic for Multicore ArchitecturesabstractSince multicore architectures are becoming more popular, recent multicore optimizations focus on energy consumption. In this paper, we focus on reducing the energy consumption in the data and instruction cache hierarchies in a multicore system. First, we present a level one data cache tuning heuristic for a heterogeneous multicore system, which classifies applications based on data sharing and cache behavior and uses this classification to guide cache tuning and reduce the number of cores that need to be tuned. Results reveal average energy savings of 25 percent for 2, 4, 8, and 16-core systems while searching only 1 percent of the design space. Next, we present a level one instruction cache tuning heuristic that reduces energy consumption in the instruction cache hierarchy by an average of 53 percent for 2, 4, 8, and 16-core systems, while searching less than 1 percent of the design space. Finally, we develop a custom, global hardware cache tuner for a dual-core system and show that our cache tuner has low area, energy, and power overheads. Marisha Rawlins, Ann Gordon-Ross |
IEEE Trans. Computers | 2 |
| 2013 | T-SPaCS - A Two-Level Single-Pass Cache Simulation MethodologyabstractThe cache hierarchy's large contribution to total microprocessor system power makes caches a good optimization candidate. To facilitate a fast design-time cache optimization process, we propose a single-pass trace-driven cache simulation methodology-T-SPaCS-for a two-level exclusive cache hierarchy. Direct adaptation of conventional trace-driven cache simulation to two-level caches requires significant storage and simulation time as numerous stacks record cache access patterns for each level one and level two cache combination and each stack is repeatedly processed. T-SPaCS significantly reduces storage space and simulation time using a set of stacks that only record the complete cache access pattern. Thereby, T-SPaCS simulates all cache configurations for both the level one and level two caches simultaneously in a single pass. Experimental results show that T-SPaCS is 21.02X faster on average than sequential simulation for instruction caches and 33.34X faster for data caches. A simplified, but minimally lossy version of T-SPaCS (simplified-T-SPaCS) increases the average simulation speedup to 30.15X for instruction caches and 41.31X for data caches. We leverage T-SPaCS and simplified-T-SPaCS for determining the lowest energy cache configuration to quantify the effects of lossiness and observe that T-SPaCS and simplified-T-SPaCS still find the lowest energy cache configuration as compared to exact simulation. Wei Zang, Ann Gordon-Ross |
IEEE Trans. Computers | 2 |
| 2013 | Dynamic profiling and fuzzy-logic-based optimization of sensor network platformsabstractThe commercialization of sensor-based platforms is facilitating the realization of numerous sensor network applications with diverse application requirements. However, sensor network platforms are becoming increasingly complex to design and optimize due to the multitude of interdependent parameters that must be considered. To further complicate matters, application experts oftentimes are not trained engineers, but rather biologists, teachers, or agriculturists who wish to utilize the sensor-based platforms for various domain-specific tasks. To assist both platform developers and application experts, we present a centralized dynamic profiling and optimization platform for sensor-based systems that enables application experts to rapidly optimize a sensor network for a particular application without requiring extensive knowledge of, and experience with, the underlying physical hardware platform. In this article, we present an optimization framework that allows developers to characterize application requirements through high-level design metrics and fuzzy-logic-based optimization. We further analyze the benefits of utilizing dynamic profiling information to eliminate the guesswork of creating a “good” benchmark, present several reoptimization evaluation algorithms used to detect if re-optimization is necessary, and highlight the benefits of the proposed dynamic optimization framework compared to static optimization alternatives. Adrian Lizarraga, Roman L. Lysecky, Susan Lysecky, Ann Gordon-Ross |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2013 | Adaptive loop caching using lightweight runtime control flow analysisabstractLoop caches provide an effective method for decreasing memory hierarchy energy consumption by storing frequently executed code (critical regions) in a more energy efficient structure than the level one cache. However, due to code structure restrictions or costly design time pre-analysis efforts, previous loop cache designs are not suitable for all applications and system scenarios. We present an adaptive loop cache that is amenable to a wider range of system scenarios, which can provide an additional 20% average instruction cache energy savings (with individual benchmark energy savings as high as 69%) compared to the next best loop cache, the preloaded loop cache. Marisha Rawlins, Ann Gordon-Ross |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2013 | High-performance optimizations on tiled many-core embedded systems: a matrix multiplication case study
Arslan Munir, Farinaz Koushanfar, Ann Gordon-Ross, Sanjay Ranka |
J. Supercomput. | 3 |
| 2013 | Scalable Digital CMOS Comparator Using a Parallel Prefix TreeabstractWe present a new comparator design featuring wide-range and high-speed operation using only conventional digital CMOS cells. Our comparator exploits a novel scalable parallel prefix structure that leverages the comparison outcome of the most significant bit, proceeding bitwise toward the least significant bit only when the compared bits are equal. This method reduces dynamic power dissipation by eliminating unnecessary transitions in a parallel prefix structure that generates the N-bit comparison result after (log4N)+(log16N)+4 CMOS gate delays. Our comparator is composed of locally interconnected CMOS gates with a maximum fan-in and fan-out of five and four, respectively, independent of the comparator bitwidth. The main advantages of our design are high speed and power efficiency, maintained over a wide range. Additionally, our design uses a regular reconfigurable VLSI topology, which allows analytical derivation of the input-output delay as a function of bitwidth. HSPICE simulation for a 64-b comparator shows a worst case input-output delay of 0.86 ns and a maximum power dissipation of 7.7 mW using 0.15- μm TSMC technology at 1 GHz. Saleh Abdel-Hafeez, Ann Gordon-Ross, Behrooz Parhami |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | An application classification guided cache tuning heuristic for multi-core architecturesabstractSince multi-core architectures are becoming more popular, recent multi-core optimizations focus on energy consumption. We present a level one data cache tuning heuristic for a heterogeneous multi-core system, which classifies applications based on data sharing and cache behavior, and uses this classification to guide cache tuning and reduce the number of cores that need to be tuned. Results reveal average energy savings of 25% for 2-, 4-, 8-, and 16-core systems while searching only 1% of the design space. Marisha Rawlins, Ann Gordon-Ross |
ASP-DAC | 2 |
| 2012 | Online algorithms for wireless sensor networks dynamic optimizationabstractTechnological advancements in wireless communications and embedded systems have led to the proliferation of wireless sensor network (WSN) applications, each with varying application requirements (i.e., lifetime, throughput, reliability, etc.). Sensor node tunable parameters enable WSN designers to specialize/tune a sensor node to meet application requirements, but however, parameter tuning is a challenging process that requires designer expertise to consider sensor node complexities and changing environmental stimuli. In this paper, we develop lightweight, online optimization algorithms for sensor node parameter tuning, which enables dynamic optimizations to meet application requirements and adapt to changing environmental stimuli. Results reveal that our online optimizations quickly converge to a near optimal solution using minimal computational and storage resources, and are thus amenable for implementation on resource and energy-constrained sensor nodes. Arslan Munir, Ann Gordon-Ross, Susan Lysecky, Roman L. Lysecky |
CCNC | 2 |
| 2012 | Dynamic phase-based tuning for embedded systems using phase distance mappingabstractPhase-based tuning specializes a system's tunable parameters to the varying runtime requirements of an application's different phases of execution to meet optimization goals. Since the design space for tunable systems can be very large, one of the major challenges in phase-based tuning is determining the best configuration for each phase without incurring significant tuning overhead (e.g., energy and/or performance) during design space exploration. In this paper, we propose phase distance mapping, which directly determines the best configuration for a phase, thereby eliminating design space exploration. Phase distance mapping applies the correlation between a known phase's characteristics and best configuration to determine a new phase's best configuration based on the new phase's characteristics. Experimental results verify that our phase distance mapping approach determines configurations within 3% of the optimal configurations on average and yields an energy delay product savings of 26% on average. Tosiron Adegbija, Ann Gordon-Ross, Arslan Munir |
ICCD | 2 |
| 2012 | Parallelized benchmark-driven performance evaluation of SMPs and tiled multi-core architectures for embedded systemsabstractWith Moore's law supplying billions of transistors on-chip, embedded systems are undergoing a transition from single-core to multi-core to exploit this high transistor density for high performance. However, there exists a plethora of multi-core architectures and the suitability of these multi-core architectures for different embedded domains (e.g., distributed, real-time, reliability-constrained) requires investigation. Despite the diversity of embedded domains, one of the critical applications in many embedded domains (especially distributed embedded domains) is information fusion. Furthermore, many other applications consist of various kernels, such as Gaussian elimination (used in network coding), that dominate the execution time. In this paper, we evaluate two embedded systems multi-core architectural paradigms: symmetric multiprocessors (SMPs) and tiled multi-core architectures (TMAs). We base our evaluation on a parallelized information fusion application and benchmarks that are used as building blocks in applications for SMPs and TMAs. We compare and analyze the performance of an Intel-based SMP and Tilera's TILEPro64 TMA based on our parallelized benchmarks for the following performance metrics: runtime, speedup, efficiency, cost, scalability, and performance per watt. Results reveal that TMAs are more suitable for applications requiring integer manipulation of data with little communication between the parallelized tasks (e.g., information fusion) whereas SMPs are more suitable for applications with floating point computations and a large amount of communication between processor cores. Arslan Munir, Ann Gordon-Ross, Sanjay Ranka |
IPCCC | 2 |
| 2012 | A single-pass cache simulation methodology for two-level unified cachesabstractCache tuning is the process of determining the optimal cache configuration given an application's requirements for reducing energy consumption and improving performance. As embedded systems trend towards unified second-level caches for improved performance, the need for fast cache tuning methodologies for multi-level cache hierarchies is becoming more critical. In this paper, we present U-SPaCS, a single-pass cache simulation methodology for design-time tuning of two-level cache hierarchies with a unified second-level cache. To afford fast simulation time, U-SPaCS maintains unique cache block addresses in a set of stacks, which enables simulation of all cache configurations for the level one instruction and data caches, and level two unified cache simultaneously in a single pass of an application's time-ordered instruction and data access trace. Experiments show that U-SPaCS can accurately determine the miss rates for a configurable cache design space consisting of 2,187 cache configurations with a 41X speedup in average simulation time as compared to the most widely-used trace-driven cache simulation. Wei Zang, Ann Gordon-Ross |
ISPASS | 2 |
| 2012 | Combining code reordering and cache configurationabstractThe instruction cache is a popular optimization target due to the cache's high impact on system performance and power and because of the cache's predictable temporal and spatial locality. This article is an in depth study on the interaction of code reordering (a long-known technique) and cache configuration (a relatively new technique). Experimental results show that code reordering coupled with cache configuration reveals additional energy savings as high as 10--15% for several benchmarks with reduced cache area as high as 48%. To exploit these additional benefits, we architect and evaluate several design exploration heuristics for combining these two methods. Ann Gordon-Ross, Frank Vahid, Nikil Dutt |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2012 | Dynamic Cache Reconfiguration for Soft Real-Time SystemsabstractIn recent years, efficient dynamic reconfiguration techniques have been widely employed for system optimization. Dynamic cache reconfiguration is a promising approach for reducing energy consumption as well as for improving overall system performance. It is a major challenge to introduce cache reconfiguration into real-time multitasking systems, since dynamic analysis may adversely affect tasks with timing constraints. This article presents a novel approach for implementing cache reconfiguration in soft real-time systems by efficiently leveraging static analysis during runtime to minimize energy while maintaining the same service level. To the best of our knowledge, this is the first attempt to integrate dynamic cache reconfiguration in real-time scheduling techniques. Our experimental results using a wide variety of applications have demonstrated that our approach can significantly reduce the cache energy consumption in soft real-time systems (up to 74%). Weixun Wang, Prabhat Mishra 0001, Ann Gordon-Ross |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2012 | An MDP-Based Dynamic Optimization Methodology for Wireless Sensor NetworksabstractWireless sensor networks (WSNs) are distributed systems that have proliferated across diverse application domains (e.g., security/defense, health care, etc.). One commonality across all WSN domains is the need to meet application requirements (i.e., lifetime, responsiveness, etc.) through domain specific sensor node design. Techniques such as sensor node parameter tuning enable WSN designers to specialize tunable parameters (i.e., processor voltage and frequency, sensing frequency, etc.) to meet these application requirements. However, given WSN domain diversity, varying environmental situations (stimuli), and sensor node complexity, sensor node parameter tuning is a very challenging task. In this paper, we propose an automated Markov Decision Process (MDP)-based methodology to prescribe optimal sensor node operation (selection of values for tunable parameters such as processor voltage, processor frequency, and sensing frequency) to meet application requirements and adapt to changing environmental stimuli. Numerical results confirm the optimality of our proposed methodology and reveal that our methodology more closely meets application requirements compared to other feasible policies. Arslan Munir, Ann Gordon-Ross |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | High-Performance Energy-Efficient Multicore Embedded ComputingabstractWith Moore's law supplying billions of transistors on-chip, embedded systems are undergoing a transition from single-core to multicore to exploit this high-transistor density for high performance. Embedded systems differ from traditional high-performance supercomputers in that power is a first-order constraint for embedded systems; whereas, performance is the major benchmark for supercomputers. The increase in on-chip transistor density exacerbates power/thermal issues in embedded systems, which necessitates novel hardware/software power/thermal management techniques to meet the ever-increasing high-performance embedded computing demands in an energy-efficient manner. This paper outlines typical requirements of embedded applications and discusses state-of-the-art hardware/software high-performance energy-efficient embedded computing (HPEEC) techniques that help meeting these requirements. We also discuss modern multicore processors that leverage these HPEEC techniques to deliver high performance per watt. Finally, we present design challenges and future research directions for HPEEC system development. Arslan Munir, Sanjay Ranka, Ann Gordon-Ross |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2012 | Reconfigurable Fault Tolerance: A Comprehensive Framework for Reliable and Adaptive FPGA-Based Space ComputingabstractCommercial SRAM-based, field-programmable gate arrays (FPGAs) have the potential to provide space applications with the necessary performance to meet next-generation mission requirements. However, mitigating an FPGA’s susceptibility to single-event upset (SEU) radiation is challenging. Triple-modular redundancy (TMR) techniques are traditionally used to mitigate radiation effects, but TMR incurs substantial overheads such as increased area and power requirements. In order to reduce these overheads while still providing sufficient radiation mitigation, we propose a reconfigurable fault tolerance (RFT) framework that enables system designers to dynamically adjust a system’s level of redundancy and fault mitigation based on the varying radiation incurred at different orbital positions. This framework includes an adaptive hardware architecture that leverages FPGA reconfigurable techniques to enable significant processing to be performed efficiently and reliably when environmental factors permit. To accurately estimate upset rates, we propose an upset rate modeling tool that captures time-varying radiation effects for arbitrary satellite orbits using a collection of existing, publically available tools and models. We perform fault-injection testing on a prototype RFT platform to validate the RFT architecture and RFT performability models. We combine our RFT hardware architecture and the modeled upset rates using phased-mission Markov modeling to estimate performability gains achievable using our framework for two case-study orbits. Adam Jacobs, Grzegorz Cieslewski, Alan D. George, Ann Gordon-Ross, Herman Lam |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2011 | An integrated development toolset and implementation methodology for partially reconfigurable system-on-chipsabstractPartial reconfiguration (PR) enhances traditional FPGA-based system-on-chips (SoCs) by providing additional benefits such as reduced area and increased functionality as compared to non-PR SoCs. However, since leveraging these additional benefits requires specific designer expertise and increased development time, PR has not yet gained widespread usage. In this paper, we present an integrated development toolset that automates the implementation of PR SoCs on FPGA devices and leverage this tool in a rapid design space exploration case study. Abelardo Jara-Berrocal, Ann Gordon-Ross |
ASAP | 2 |
| 2011 | On the interplay of loop caching, code compression, and cache configurationabstractEven though much previous work explores varying instruction cache optimization techniques individually, little work explores the combined effects of these techniques (i.e., do they complement or obviate each other). In this paper we explore the interaction of three optimizations: loop caching, cache tuning, and code compression. Results show that loop caching increases energy savings by as much as 26% compared to cache tuning alone and reduces decompression energy by as much as 73%. Marisha Rawlins, Ann Gordon-Ross |
ASP-DAC | 2 |
| 2011 | T-SPaCS - A two-level single-pass cache simulation methodologyabstractThe cache hierarchy's large contribution to total microprocessor system power makes caches a good optimization candidate. We propose a single-pass trace-driven cache simulation methodology - T-SPaCS - for a two-level exclusive instruction cache hierarchy. Instead of storing and simulating numerous stacks repeatedly as in direct adaptation of a conventional trace-driven cache simulation to two level caches, T-SPaCS simulates both the level one and level two caches simultaneously using one stack. Experimental results show T-SPaCS efficiently and accurately determines the optimal cache configuration (lowest energy). Wei Zang, Ann Gordon-Ross |
ASP-DAC | 2 |
| 2011 | Hardware module reuse and runtime assembly for dynamic management of reconfigurable resourcesabstractPartial reconfiguration (PR) enhances traditional FPGA-based systems-on-a-chip (SoCs) by providing benefits such as reduced area requirements and increased system flexibility. In multi-application PR SoCs, a dynamic resource manager (DRM) must efficiently orchestrate PR hardware resource management (access to and sharing of PR resources) in order to minimize the percentage of wasted/unused PR resources and reconfiguration time overhead. In this paper, we present DRM software that leverages two techniques, hardware module reuse and dynamic inter-module communication, to reduce wasted/unused PR hardware resources by 13% and reduce reconfiguration time by 33% as compared to a DRM without these techniques. Abelardo Jara-Berrocal, Ann Gordon-Ross |
FPT | 2 |
| 2011 | Formulation-level design space exploration for partially reconfigurable FPGAsabstractExploiting the benefits afforded by runtime partial reconfiguration (PR) on modern field-programmable gate arrays (FPGAs)requires PR-capable applications and associated PR-architectures, both of which are challenging tasks due to competing implementation metrics(e.g., area, power, operating frequency, etc.) and results in unmanageable design spaces. PR design space exploration (DSE) techniques and tools assist designers in efficiently and effectively exploring this design space. This paper presents the first, to the best of our knowledge, formulation-level PR DSE tool - FoRSE. FoRSE leverages the application's PR-architecture and mathematical FPGA device models and vendor-specified PR technology to generate Pareto-optimal sets of PR-floorplans and devices based on designer-designated implementation metrics. FoRSE can prune an application's implementation design space by three to four orders of magnitude in approximately 15 seconds. Ann Gordon-Ross |
FPT | 2 |
| 2011 | Partially reconfigurable system-on-chips for adaptive fault toleranceabstractDue to the runtime flexibility of modern dynamically reconfigurable SRAM-based FPGAs, FPGA devices have become an attractive platform for developing system-on-chips (SoCs) for space applications (space SoCs). However, since the FPGA's SRAM is highly susceptible to space radiation, system reliability is a primary concern for space SoCs. To maintain system reliability and mitigate space radiation effects, space SoCs must be designed with redundant copies of system functionality. Space SoCs must contain enough redundancy to ensure system reliability for the highest anticipated radiation level, which imposes a large area overhead. However, since radiation levels vary based on the system's orbital position, the system does not always require the highest level of redundancy. Space SoCs that can adapt the system redundancy based on the current radiation level can achieve more effective device utilization. In this paper, we present a flexible, FPGA-based, adaptive SoC for space system development. Our space SoC leverages partial reconfiguration to dynamically adapt the system's level of redundancy according to varying radiation levels. We present a software algorithm to manage the system's adaptability, implement the SoC on a Xilinx Virtex-5 device, and evaluate the SoC's resource utilization using the International Space Station's orbit. Shaon Yousuf, Adam Jacobs, Ann Gordon-Ross |
FPT | 3 |
| 2011 | Markov Modeling of Fault-Tolerant Wireless Sensor NetworksabstractTechnological advancements in communications and embedded systems have led to the proliferation of wireless sensor networks (WSNs) in a wide variety of application domains. One commonality across all WSN application domains is the need to meet application requirements (e.g., lifetime, reliability, etc.). Many application domains require that sensor nodes be deployed in harsh environments (e.g., ocean floor, active volcanoes), making these sensor nodes more prone to failures. Unfortunately, sensor node failures can be catastrophic for critical or safety related systems. To improve reliability in such systems, we propose a fault-tolerant sensor node model for applications with high reliability requirements. We develop Markov models for characterizing WSN reliability and MTTF (Mean Time to Failure) to facilitate WSN application-specific design. Results show that our proposed fault-tolerant model can result in as high as a 100% MTTF increase and approximately a 350% improvement in reliability over a non-fault-tolerant WSN. Results also highlight the significance of a robust fault detection algorithm to leverage the benefits of fault-tolerant WSNs. Arslan Munir, Ann Gordon-Ross |
ICCCN | 2 |
| 2011 | A queueing theoretic approach for performance evaluation of low-power multi-core embedded systemsabstractWith Moore's law supplying billions of transistors on-chip, embedded systems are undergoing a transition from single-core to multi-core to exploit this high transistor density for high performance. However, the optimal layout of these multiple cores along with the memory subsystem (caches and main memory) to satisfy power, area, and often stringent real-time constraints is a challenging design endeavor. The short time-to-market constraint of embedded systems exacerbates this design challenge and necessitates the architectural modeling of embedded systems to reduce the time-to-market by expediting target applications to device/architecture mapping. In this paper, we present a queueing theoretic approach for modeling multi-core embedded systems that provides a quick and inexpensive performance evaluation both in terms of time and resources as compared to the development of multi-core simulators and running benchmarks on these simulators. We also calculate chip area and power consumption for different multi-core embedded architectures with a varying number of processor cores and cache configurations to provide a comparative analysis of multicore embedded architectures in terms of performance, area, and power consumption. Our performance and power results indicate that multi-core embedded system architectures that leverage shared last-level caches (LLCs) provide the best LLC performance per watt but may introduce main memory response time and throughput bottlenecks for high cache miss rates, whereas architectures leveraging a hybrid of private and shared LLCs alleviate main memory bottlenecks at the expense of reduced performance per watt. Arslan Munir, Ann Gordon-Ross, Sanjay Ranka |
ICCD | 2 |
| 2011 | CPACT - The conditional parameter adjustment cache tuner for dual-core architecturesabstractCache tuning reveals substantial energy savings for single-core architectures, but has yet to be explored for multi-core architectures. In this paper we explore level one (L1) data cache tuning in a heterogeneous dual-core system where each data cache can have a different configuration. We show that L1 data cache tuning in a dual-core system achieves 25% average energy savings, which is comparable to single-core data cache tuning. We present the dual-core tuning heuristic CPACT, which finds cache configurations within 1% of the optimal configuration while searching only 1% of the design space. Finally, we provide valuable insights on core-interactions and data coherence revealed when tuning the multithreaded SPLASH-2 benchmarks. Marisha Rawlins, Ann Gordon-Ross |
ICCD | 2 |
| 2011 | A Digital CMOS Parallel Counter Architecture Based on State Look-Ahead LogicabstractWe present a high-speed wide-range parallel counter that achieves high operating frequencies through a novel pipeline partitioning methodology (a counting path and state look-ahead path), using only three simple repeated CMOS-logic module types: an initial module generates anticipated counting states for higher significant bit modules through the state look-ahead path, simple D-type flip-flops, and 2-bit counters. The state look-ahead path prepares the counting path's next counter state prior to the clock edge such that the clock edge triggers all modules simultaneously, thus concurrently updating the count state with a uniform delay at all counting path modules/stages with respect to the clock edge. The structure is scalable to arbitrary N-bit counter widths (2-to-2N range) using only the three module types and no fan-in or fan-out increase. The counter's delay is comprised of the initial module access time (a simple 2-bit counting stage), one three-input and-gate delay, and a D-type flip-flop setup-hold time. We implemented our proposed counter using a 0.15-μ m TSMC digital cell library and verified maximum operating speeds of 2 and 1.8 GHz for 8- and 17-bit counters, respectively. Finally, the area of a sample 8-bit counter was 78 125 μ m2(510 transistors) and consumed 13.89 mW at 2 GHz. Saleh Abdel-Hafeez, Ann Gordon-Ross |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | VAPRES: A Virtual Architecture for Partially Reconfigurable Embedded SystemsabstractDue to the runtime flexibility offered by field programmable gate arrays (FPGAs), FPGAs are popular devices for stream processing systems, since many stream processing applications require runtime adaptability (i.e. throughput, data transformations, etc.). FPGAs can offer this adaptability through runtime assembly of stream processing systems that are decomposed into hardware modules. Runtime hardware module assembly consists of dynamic hardware module replacement and hardware module communication reconfiguration. In this paper, we architect a flexible base embedded system amenable to runtime assembly of stream processing systems using custom communication architecture with dynamic streaming channel establishment between hardware modules. We present a hardware module swapping methodology that replaces hardware modules without stream processing interruption. Finally, we formulate two design flows, system and application construction, to provide system and application designer assistance. Abelardo Jara-Berrocal, Ann Gordon-Ross |
DATE | 2 |
| 2010 | Lightweight runtime control flow analysis for adaptive loop cachingabstractLoop caches provide an effective method for decreasing memory hierarchy energy consumption by storing frequently executed code in a more energy efficient structure than the level one cache. However, due to code structure restrictions and/or costly design time pre-analysis efforts, previous loop cache designs are not suitable for all applications and system scenarios. In this paper, we present an adaptive loop cache that is amenable to a wide range of system scenarios, providing an additional 20% average instruction memory hierarchy energy savings (with individual benchmark energy savings as high as 69%) compared to the best previous loop cache design. Marisha Rawlins, Ann Gordon-Ross |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | A lightweight dynamic optimization methodology for wireless sensor networksabstractTechnological advancements in embedded systems due to Moore's law have lead to the proliferation of wireless sensor networks (WSNs) in different application domains (e.g. defense, health care, surveillance systems) with different application requirements (e.g. lifetime, reliability). Many commercial-off-the-shelf (COTS) sensor nodes can be specialized to meet these requirements using tunable parameters (e.g. voltage, frequency) to specialize the operating state. Since a sensor node's performance depends greatly on environmental stimuli, dynamic optimizations enable sensor nodes to automatically determine their operating state in-situ. However, dynamic optimization methodology development given a large design space and resource constraints (memory and computational) is a very challenging task. In this paper, we propose a lightweight dynamic optimization methodology that intelligently selects initial tunable parameter values to produce a high-quality initial operating state in one-shot for time-critical or highly constrained applications. Further operating state improvements are made using an efficient greedy exploration algorithm, achieving optimal or near-optimal operating states while exploring only 0.04% of the design space on average. Arslan Munir, Ann Gordon-Ross, Susan Lysecky, Roman L. Lysecky |
WiMob | 2 |
| 2010 | SIP-Based IMS Signaling Analysis for WiMax-3G Interworking ArchitecturesabstractThe third-generation partnership project (3GPP) and 3GPP2 have standardized the IP multimedia subsystem (IMS) to provide ubiquitous and access network-independent IP-based services for next-generation networks via merging cellular networks and the Internet. The application layer Session Initiation Protocol (SIP), standardized by 3GPP and 3GPP2 for IMS, is responsible for IMS session establishment, management, and transformation. The IEEE 802.16 worldwide interoperability for microwave access (WiMax) promises to provide high data rate broadband wireless access services. In this paper, we propose two novel interworking architectures to integrate WiMax and third-generation (3G) networks. Moreover, we analyze the SIP-based IMS registration and session setup signaling delay for 3G and WiMax networks with specific reference to their interworking architectures. Finally, we explore the effects of different WiMax-3G interworking architectures on the IMS registration and session setup signaling delay. Arslan Munir, Ann Gordon-Ross |
IEEE Trans. Mob. Comput. | 2 |
| 2009 | Bitstream relocation with local clock domains for partially reconfigurable FPGAsabstractPartial Reconfiguration (PR) of FPGAs presents many opportunities for application design flexibility, enabling tasks to dynamically swap in and out of the FPGA without entire system interruption. However, mapping a task to any available PR region (PRR) requires a unique partial bitstream for each PRR. This replication can introduce significant overheads in terms of bitstream storage and communication requirements. Previous research in partial bitstream relocation can alleviate these overheads by transforming a single partial bitstream to map to any available PRR. However, careful steps are necessary to ensure proper functionality of relocated partial bitstreams and may result in clock routing inefficiencies. These routing inefficiencies can be alleviated by using regional clock resources introduced in the Virtex-4 FPGAs to implement local clock domains. PRRs can internally drive local clock domains, enabling each PRR to vary its clock frequency with respect to a single global clock signal, as opposed to sending multiple global clock signals (one for each desired clock frequency) to each PRR. We introduce this novel local clock domain (LCD) concept, which provides enhanced PR design flexibility. However, integration of LCDs and partial bitstream relocation introduces new challenges. In this paper, we identify motivating application domains for this integration, analyze integration benefits, and provide a detailed integration methodology. Adam Flynn, Ann Gordon-Ross, Alan D. George |
DATE | 2 |
| 2009 | SCORES: A scalable and parametric streams-based communication architecture for modular reconfigurable systemsabstractParallel architectures have become an increasingly popular method in which to achieve high performance with low power consumption. In order to leverage these benefits, applications are decomposed into multiple computational modules (tasks) that collectively operate and communicate in parallel. In this paper, we present a scalable and highly parametric streams-based communication architecture for inter-module communication for FPGA-based systems - SCORES. This communication architecture improves on previous methods by providing increased application specialization and heterogeneous module clock frequencies, as well as providing a means for low latency communication and data throughput guarantees. Abelardo Jara-Berrocal, Ann Gordon-Ross |
DATE | 2 |
| 2009 | Exploiting Partially Reconfigurable FPGAs for Situation-Based Reconfiguration in Wireless Sensor NetworksabstractWireless sensor networks (WSNs) are typically composed of very small, battery-operated devices (sensor nodes) containing simple microprocessors with few computational resources. However, the rapidly increasing popularity of WSNs has placed increased computational demands upon these systems, due to increasingly complex operating environments and enhanced data-sensing technology. Whereas introducing more powerful microprocessors into sensor nodes addresses these demands, sensor nodes do not contain sufficient energy reserves to support these microprocessors. In this paper, we present a partially reconfigurable FPGA-based architecture and methodology to provide increased WSN flexibility and computational resources, resulting in superior power consumption and performance compared to a microprocessor capable of satisfying similar demands. Rafael García, Ann Gordon-Ross, Alan D. George |
FCCM | 2 |
| 2009 | Macs: A Minimal Adaptive routing circuit-switched architecture for scalable and parametric NoCsabstractNetworks-on-chips (NoCs) are an emerging communication topology paradigm in single chip VLSI design, enhancing parallelism and system scalability. Processing units (PUs) connect to the communication topology via routers, which are responsible for runtime establishment and management of inter-PU communication channels. Router design directly affects overall system performance and exploited parallelism. In this paper, we present a highly parametric NoC architecture, MACS, providing increased system speed, designer flexibility, and scalability as compared to previous methods. In addition, MACS enhances inter-PU communication using a circuit-switching technique with dedicated, high frequency communication channels. Compared to previous work, MACS offers a 5x increase in operating frequency and a 2x reduction in area overhead. Ann Gordon-Ross |
FPL | 2 |
| 2009 | Fast Configurable-Cache Tuning With a Unified Second-Level CacheabstractTuning a configurable cache subsystem to an application can greatly reduce memory hierarchy energy consumption. Previous tuning methods use a level one configurable cache only, or a second level with separate instruction and data configurable caches. We instead use a commercially-common unified second level cache, a seemingly minor difference that actually expands the configuration space from 500 to about 20 000. We develop additive way tuning for tuning a cache subsystem with this large space, yielding 61% energy savings and 9% performance improvements over a nonconfigurable cache, greatly outperforming an extension of a previous method. Ann Gordon-Ross, Frank Vahid, Nikil Dutt |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2008 | Phase-based cache reconfiguration for a highly-configurable two-level cache hierarchyabstractPhase-based tuning methodologies specialize system parameters for each application phase of execution. Parameters are varied during execution, as opposed to remaining fixed as in an application-based tuning methodology. Prior work and logic suggests phase-based tuning may provide significant savings over application-based tuning. We investigate this hypothesis using a detailed cache model and tune a highly-configurable cache on a per-phase basis compared to tuning once per application, and found phase-based tuning to yield improvements of up to 37% in performance and 20% in energy over application-based tuning. Furthermore, we extend previous phase-based tuning of a configurable cache by significantly increasing configurability and show 14% energy improvement compared to previous methods. In addition, we quantify the overhead imposed due to cache reconfiguration. Ann Gordon-Ross, Jeremy Lau, Brad Calder |
ACM Great Lakes Symposium on VLSI | 1 |
| 2008 | A table-based method for single-pass cache optimizationabstractDue to the large contribution of the memory subsystem to total system power, the memory subsystem is highly amenable to customization for reduced power/energy and/or improved performance. Cache parameters such as total size, line size, and associativity can be specialized to the needs of an application for system optimization. In order to determine the best values for cache parameters, most methodologies utilize repetitious application execution to individually analyze each configuration explored. In this paper we propose a simplified yet efficient technique to accurately estimate the miss rate of many different cache configurations in just one single-pass of execution. The approach utilizes simple data structures in the form of a multi-layered table and elementary bitwise operations to capture the locality characteristics of an application's addressing behavior. The proposed technique intends to ease miss rate estimation and reduce cache exploration time. Pablo Viana, Ann Gordon-Ross, Edna Barros, Frank Vahid |
ACM Great Lakes Symposium on VLSI | 2 |
| 2008 | A resource efficient content inspection system for next generation Smart NICsabstractThe aggregate power consumption of the Internet is increasing at an alarming rate, due in part to the rapid increase in the number of connected edge devices such as desktop PCs. Despite being left idle 75% of the time, 90% of PCs have their power management features disabled. Consequently, much recent research has focused on reducing power consumption of Internet edge devices. One such method for reducing PC power consumption is by augmenting the network interface card (NIC) with enhanced processing capabilities. These capabilities pave the way for green computing by allowing the PC to transition to a low-power sleep state while the NIC responds to network traffic on behalf of the PC - a technique known as power proxying. However, such a Smart-NIC (SNIC) requires specialized low-power, resource-constrained processing, and architectural features in order to realize such capabilities. In this paper, we present a NIC-based packet content inspection system for power proxying and network intrusion detection. We use a novel partitioned TCAM technique that results in 87% energy savings and a 62% lower energy-delay product than existing non-partitioned router-based techniques, thus making our technique highly suitable for SNIC-based deployment. Karthik Sabhanatarajan, Ann Gordon-Ross |
ICCD | 2 |
| 2008 | Real-time performance analysis of Adaptive Link RateabstractHigh speed links are widely deployed in modern day computer networks to meet the ever growing needs for increasing data bandwidth. However, with the increase in the link rate, the power consumption of the network interfaces increases exponentially, compounding growing concerns about network power consumption. Fortunately, network traffic characteristics show that rapid link rates are not always required. During times of reduced network traffic, the Adaptive Link Rate (ALR) mechanism allows link rates to be reduced with little impact on network performance. Current research has focused on policies to control when and how to change link rates, and have shown promising energy savings. However, these works have been largely simulative, and have not addressed many of the challenges involved in implementation. In this paper, we develop a hardware prototype ALR system and address real-time challenges involved in realizing such an implementation. We also identify new considerations for control policy development given current technology capabilities as well as future projections. Baoke Zhang, Karthik Sabhanatarajan, Ann Gordon-Ross, Alan D. George |
LCN | 3 |
| 2007 | A Self-Tuning Configurable CacheabstractThe memory hierarchy of a system can consume up to 50% of microprocessor system power. Previous work has shown that tuning a configurable cache to a particular application can reduce memory subsystem energy by 62% on average. We introduce a self-tuning cache that performs transparent runtime cache tuning, thus relieving the application designer and/or compiler from predetermining an application's cache configuration. The self-tuning cache applies tuning at a determined tuning interval. A good interval balances tuning process energy overhead against the energy overhead of running in a sub-optimal cache configuration, which we show wastes much energy. We present a self-tuning cache that dynamically varies the tuning interval, resulting in average energy reduction of as much as 29%, falling within 13% of an oracle-based optimal method. Ann Gordon-Ross, Frank Vahid |
DAC | 1 |
| 2007 | A one-shot configurable-cache tuner for improved energy and performanceabstractWe introduce a new non-intrusive on-chip cache-tuning hardware module capable of accurately predicting the best configuration of a configurable cache for an executing application. Previous dynamic cache tuning approaches change the cache configuration several times as part of the tuning search process, executing the application using inferior configurations and temporarily causing energy and performance overhead. The introduced tuner uses a different approach, which non-intrusively collects data on addresses issued by the microprocessor, analyzes that data to predict the best cache configuration, and then updates the cache to the new best configuration in "one-shot", without ever having to examine inferior configurations. The result is less energy and less performance overhead, meaning that cache tuning can be applied more frequently. We show through experiments that the one-shot cache tuner can reduce memory-access related energy for instructions by 35% and comes within 4% of a previous intrusive approach, and results in 4.6 times less energy overhead and a 7.7 times speedup in tuning time compared to a previous intrusive approach, at the main expense of 12% larger size Ann Gordon-Ross, Pablo Viana, Frank Vahid, Walid A. Najjar, Edna Barros |
DATE | 1 |
| 2006 | Configurable cache subsetting for fast cache tuningabstractNumerous variations of configurable caches, having variable parameters like total size, line size, and associativity, have been proposed in commercial microprocessors in recent years. Tuning a configurable cache to a target application has been shown to reduce memory-access power by over 50%. However, searching the configuration space for the best configuration can require much time or power, even when using recent cache tuning heuristics. We sought to determine, for a particular domain of applications, the smallest subset of cache configurations that would still enable effective tuning. For a suite of 34 benchmarks and a cache with 18 possible configurations, we determine through an exhaustive search of all possible subsets, that only 3 or 4 candidate configurations are necessary to support tuning. We introduce a new heuristic, adapted from an efficient and effective heuristic developed for data mining, to quickly determine the best configurations for any sized subset, with near optimal results. We then consider a configurable cache with 17,640 possible configurations and improve our heuristic to include a pre-pruning step, yielding near optimal tuning results. We conclude that only 3 or 4 possible cache configurations are needed to offer a near optimal configuration for every benchmark in our suite - resulting in a 91% reduction in design space exploration time over a state-of-the-art cache tuning heuristic. Pablo Viana, Ann Gordon-Ross, Eamonn J. Keogh, Edna Barros, Frank Vahid |
DAC | 2 |
| 2005 | A first look at the interplay of code reordering and configurable cachesabstractThe instruction cache is a popular target for optimizations of microprocessor-based systems because of the cache's high impact on system performance and power, and because of the cache's predictable temporal and spatial locality. Optimization techniques can be designed based on this predictability. We explore for the first time the interplay of two popular instruction cache optimization techniques: the long-known technique of code reordering and the relatively-new technique of cache configuration. We address the question of whether those two optimizations complement each other or if one optimization dominates the other. Through experiments using embedded system benchmarks, we show that cache configuration dominates a particular category of code reordering techniques with respect to optimizing performance and energy, obviating the need for reordering. We also examine the modern scenario of synthesized custom caches, and show that combining cache configuration with code reordering results in cache size reductions of 13% on average, and up to 89% in some benchmarks, beyond just cache configuration alone. Ann Gordon-Ross, Frank Vahid, Nikil Dutt |
ACM Great Lakes Symposium on VLSI | 1 |
| 2005 | Fast configurable-cache tuning with a unified second-level cacheabstractTuning a configurable cache subsystem to an application can greatly reduce memory hierarchy energy consumption. Previous tuning methods use a level one configurable cache only, or a second level with separate instruction and data configurable caches. We instead use a commercially-common unified second level, a seemingly minor difference that actually expands the configuration space from 500 to about 20,000. We develop additive way tuning for tuning a cache subsystem with this large space, yielding 62% energy savings and 35% performance improvements over a non-configurable cache, greatly outperforming an extension of a previous method Ann Gordon-Ross, Frank Vahid, Nikil Dutt |
ISLPED | 1 |
| 2005 | Frequent Loop Detection Using Efficient Nonintrusive On-Chip HardwareabstractDynamic software optimization methods are becoming increasingly popular for improving software performance and power. The first step in dynamic optimization consists of detecting frequently executed code, or "critical regions." Most previous critical region detectors have been targeted to desktop processors. We introduce a critical region detector targeted to embedded processors, with the unique features of being very size and power efficient and being completely nonintrusive to the software's execution-features needed in timing-sensitive embedded systems. Our detector not only finds the critical regions, but also determines their relative frequencies, a potentially important feature for selecting among alternative dynamic optimization methods. Our detector uses a tiny cache-like structure coupled with a small amount of logic. We provide results of extensive explorations across 19 embedded system benchmarks. We show that highly accurate results can be achieved with only a 0.02 percent power overhead, acceptable size overhead; and zero runtime overhead. Our detector is currently being used as part of a dynamic hardware/software partitioning approach, but is applicable to a wide variety of situations. Ann Gordon-Ross, Frank Vahid |
IEEE Trans. Computers | 1 |
| 2004 | Automatic Tuning of Two-Level Caches to Embedded ApplicationsabstractThe power consumed by the memory hierarchy of a microprocessor can contribute to as much as 50% of the total microprocessor system power, and is thus a good candidate for optimizations. We present an automated method for tuning two-level caches to embedded applications for reduced energy consumption. The method is applicable to both a simulation-based exploration environment and a hardware-based system prototyping environment. We introduce the two-level cache tuner, or TCaT - a heuristic for searching the huge solution space of possible configurations. The heuristic interlaces the exploration of the two cache levels and searches the various cache parameters in a specific order based on their impact on energy. We show the integrity of our heuristic across multiple memory configurations and even in the presence of hardware/software partitioning - a common optimization capable of achieving significant speedups and/or reduced energy consumption. We apply our exploration heuristic to a large set of embedded applications. Our experiments demonstrate the efficacy of our heuristic: on average the heuristic examines only 7% of the possible cache configurations, but results in cache sub-system energy savings of 53%, only 1% more than the optimal cache configuration. In addition, the configured cache achieves an average speedup of 30% over the base cache configuration due to tuning of cache line size to the application's needs. Ann Gordon-Ross, Frank Vahid, Nikil Dutt |
DATE | 1 |
| 2003 | Frequent loop detection using efficient non-intrusive on-chip hardwareabstractDynamic software optimization methods are becoming increasingly popular for improving software performance and power. The first step in dynamic optimization consists of detecting frequently executed code, or "critical regions." Previous critical region detectors have been targeted to desktop processors. We introduce a critical region detector targeted to embedded processors, with the unique features of being very size and power efficient, and being completely non-intrusive to the software's execution - features needed in timing-sensitive embedded systems. Our detector not only finds the critical regions, but also determines their relative frequencies, a potentially important feature for selecting among alternative dynamic optimization methods. Our detector uses a tiny cache coupled with a small amount of logic. We provide results of extensive explorations across seventeen embedded system benchmarks. We show that highly accurate results can be achieved with only a 0.02% power overhead and acceptable size overhead. Our detector is currently being used as part of a dynamic hardware/software partitioning approach, but is applicable to a wide-variety of situations. Ann Gordon-Ross, Frank Vahid |
CASES | 1 |
| 2003 | Tiny instruction caches for low power embedded systemsabstractInstruction caches have traditionally been used to improve software performance. Recently, several tiny instruction cache designs, including filter caches and dynamic loop caches, have been proposed to instead reduce software power. We propose several new tiny instruction cache designs, including preloaded loop caches, and one-level and two-level hybrid dynamic/preloaded loop caches. We evaluate the existing and proposed designs on embedded system software benchmarks from both the Powerstone and MediaBench suites, on two different processor architectures, for a variety of different technologies. We show on average that filter caching achieves the best instruction fetch energy reductions of 60--80%, but at the cost of about 20% performance degradation, which could also affect overall energy savings. We show that dynamic loop caching gives good instruction fetch energy savings of about 30%, but that if a designer is able to profile a program, preloaded loop caching can more than double the savings. We describe automated methods for quickly determining the best loop cache configuration, methods useful in a core-based design flow. Ann Gordon-Ross, Susan Cotterell, Frank Vahid |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2002 | Dynamic Loop Caching Meets Preloaded Loop Caching - A Hybrid ApproachabstractDynamically-loaded tagless loop caching reduces instruction fetch power for embedded software with small loops, but only supports simple loops without taken branches. Preloaded tagless loop caching supports complex loops with branches and thus can reduce power further, but has a limit on the total number of instructions cached. We show that each does well on particular benchmarks, but neither is best across all of those benchmarks. We present a new hybrid loop cache that only preloads the complex loops, while dynamically loading other loops, thus achieving the strengths of each approach. We demonstrate better power savings than either previous approach alone. Ann Gordon-Ross, Frank Vahid |
ICCD | 1 |
| 2001 | A self-optimizing embedded microprocessor using a loop table for low powerabstractWe describe an approach for a microprocessor to tune itself to its fixed application to reduce power in an embedded system. We define a basic architecture and methodology supporting a microprocessor self-optimizing mode. We also introduce a loop table as a tunable component, although self-optimization can be done for other tunable components too. We highlight experimental results illustrating good power reductions with no performance penalty. Keywords System-on-a-chip, self-optimizing architecture, embedded systems, parameterized architectures, cores, low-power, tuning, platforms. Frank Vahid, Ann Gordon-Ross |
ISLPED | 2 |