Bruce R. Childers

dblp:00/91 · also Bruce Robert Childers · DBLP profile ↗
← Back
88ranked-venue papers
4as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 60 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 33 · 2 first-authorSecurity and privacy · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
24 papers
Memory systems · 60% GPUs and heterogeneous computing · 8% Hardware reliability and fault tolerance · 8%
Software engineering, system software, and programming languages
6 papers
Runtime systems and virtual machines · 30% Empirical software engineering · 23% Program analysis · 23%

Topics — the 30 heaviest of 79, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
non-volatile memory
1.062016
Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016
Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory · ACM Trans. Archit. Code Optim. 2013
Bit mapping for balanced PCM cell programming · ISCA 2013
Memory systems › non-volatile memory
phase change memory
0.852013
Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory · ACM Trans. Archit. Code Optim. 2013
Bit mapping for balanced PCM cell programming · ISCA 2013
Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems · ACM Trans. Archit. Code Optim. 2012
Memory systems
cache management
0.642016
Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016
Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems · ACM Trans. Archit. Code Optim. 2012
CloudCache: Expanding and shrinking private caches · HPCA 2011
GPUs and heterogeneous computing
GPU sharing
0.522017
Quality of Service Support for Fine-Grained Sharing on GPUs · ISCA 2017
Simultaneous Multikernel GPU: Multi-tasking throughput processors via fine-grained sharing · HPCA 2016
Memory systems
memory management
0.412020
HPE: Hierarchical Page Eviction Policy for Unified Memory in GPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Memory systems › memory management › virtual memory
page replacement
0.412020
HPE: Hierarchical Page Eviction Policy for Unified Memory in GPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Memory systems › cache management › cache partitioning
last-level cache partitioning
0.422016
Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016
Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems · ACM Trans. Archit. Code Optim. 2012
Memory systems
DRAM
0.432016
Restore truncation for performance improvement in future DRAM systems · HPCA 2016
Near-Memory Caching for Improved Energy Consumption · IEEE Trans. Computers 2007
COMeT+: Continuous Online Memory Testing with Multi-Threading Extension · IEEE Trans. Computers 2014
Empirical software engineering › reproducibility
artifact evaluation
0.312018
Artifact Evaluation: FAD or Real News? · ICDE 2018
Energy-efficient computing
power management
0.332016
FPB: Fine-grained Power Budgeting to Improve Write Throughput of Multi-level Cell Phase Change Memory · MICRO 2012
Restore truncation for performance improvement in future DRAM systems · HPCA 2016
Near-Memory Caching for Improved Energy Consumption · IEEE Trans. Computers 2007
Cloud and datacenter computing
quality of service
0.312017
Quality of Service Support for Fine-Grained Sharing on GPUs · ISCA 2017
Distributed systems
fault tolerance
0.322013
Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory · ACM Trans. Archit. Code Optim. 2013
PERFECTORY: A Fault-Tolerant Directory Memory Architecture · IEEE Trans. Computers 2010
Processor architecture and microarchitecture
chip multiprocessor
0.332011
CloudCache: Expanding and shrinking private caches · HPCA 2011
StimulusCache: Boosting performance of chip multiprocessors with excess cache · HPCA 2010
PERFECTORY: A Fault-Tolerant Directory Memory Architecture · IEEE Trans. Computers 2010
GPUs and heterogeneous computing › GPU performance analysis
GPU utilization
0.212016
Simultaneous Multikernel GPU: Multi-tasking throughput processors via fine-grained sharing · HPCA 2016
Memory systems
memory bandwidth management
0.212016
Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016
Hardware reliability and fault tolerance
process variation
0.212016
Restore truncation for performance improvement in future DRAM systems · HPCA 2016
Storage systems › magnetic recording
timing recovery
0.212016
Restore truncation for performance improvement in future DRAM systems · HPCA 2016
Memory systems › memory hierarchy › cache hierarchy
l2 cache
0.222011
CloudCache: Expanding and shrinking private caches · HPCA 2011
StimulusCache: Boosting performance of chip multiprocessors with excess cache · HPCA 2010
Runtime systems and virtual machines › binary translation
dynamic binary translation
0.222011
Evaluating indirect branch handling mechanisms in software dynamic translation systems · ACM Trans. Archit. Code Optim. 2011
Heterogeneous code cache: using scratchpad and main memory in dynamic binary translators · DAC 2009
Memory systems › memory management › virtual memory
address translation
0.212015
Supporting superpages in non-contiguous physical memory · HPCA 2015
Memory systems › memory management › virtual memory
huge pages
0.212015
Supporting superpages in non-contiguous physical memory · HPCA 2015
Memory systems › memory management
physical memory management
0.212015
Supporting superpages in non-contiguous physical memory · HPCA 2015
Memory systems › memory management
virtual memory
0.212015
Supporting superpages in non-contiguous physical memory · HPCA 2015
Hardware reliability and fault tolerance
memory reliability
0.212014
COMeT+: Continuous Online Memory Testing with Multi-Threading Extension · IEEE Trans. Computers 2014
Storage systems › data compression
delta compression
0.212013
Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory · ACM Trans. Archit. Code Optim. 2013
Memory systems › cache
DRAM cache
0.212013
Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory · ACM Trans. Archit. Code Optim. 2013
Memory systems › hybrid memory
hybrid main memory
0.212013
Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory · ACM Trans. Archit. Code Optim. 2013
Storage systems › flash and SSD › flash memory management
wear leveling
0.212013
Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory · ACM Trans. Archit. Code Optim. 2013
Memory systems
cache coherence
0.122011
PERFECTORY: A Fault-Tolerant Directory Memory Architecture · IEEE Trans. Computers 2010
CloudCache: Expanding and shrinking private caches · HPCA 2011
Hardware reliability and fault tolerance
error correction
0.112012
Improving write operations in MLC phase change memory · HPCA 2012

Methods — techniques the papers use, named apart from their topics

simulation · 0.5resource management · 0.3per-cycle progress control · 0.3theoretical model · 0.2restore truncation · 0.2fairness-aware resource allocation · 0.2exhaustive search · 0.2dynamic sharing · 0.2approximate scheme · 0.2gap-tolerant sequential mapping · 0.2profiling · 0.1benchmarking · 0.1static checking · 0.1data flow analysis · 0.1code cache management policies · 0.1value numbering · 0.1partial redundancy elimination · 0.1loop invariant code motion · 0.1
YearPublicationVenuePosition
2023 IEEE TC Special Issue on Real-Time Systems
abstract
The fifteen papers in this special section focus on real-time systems. They present state-of-the-art work in theory, design, analysis, implementation, and evaluation of real-time systems. All the papers address some form of real-time requirements such as deadlines, response times or delays/latency and consider not only hard real-time systems but also time-sensitive systems in general. Following an open call for papers, authors from all over the globe sent 53 submissions on a broad range of topics. The review committee of top experts worldwide conducted rigorous professional reviews. Each paper at least 3 reviews in the first round. Approximately 50 reviews were performed in the second round to evaluate the revised submissions.
Enrico Bini, Thidapat Chantem, Bruce R. Childers, Daniel Mossé
IEEE Trans. Computers3
2020 Coordinated Page Prefetch and Eviction for Memory Oversubscription Management in GPUs
abstract
The adoption of unified memory and demand paging has simplified programming and eased memory management in discrete GPUs. However, long-latency page faults cause significant performance overhead. While several software-based mechanisms have been proposed to address this issue, they suffer from inefficiency when page prefetching and pre-eviction are combined. For example, a state-of-the-art page replacement policy, hierarchical page eviction (HPE), is inefficient when prefetching is enabled. Furthermore, the prefetcher semantics-aware pre-evicting policy, which pre-evicts continuous pages in bulk the way they were brought in by the prefetcher, may cause thrashing for some irregular applications.In this paper, coordinated page prefetch and eviction (CPPE) is proposed to manage memory oversubscription in GPUs with unified memory. CPPE incorporates a modified page eviction policy, MHPE, and an access pattern-aware prefetcher in a fine-grained manner: MHPE is aware of prefetch semantics and the prefetcher prefetches pages according to access patterns in eviction candidates selected by MHPE. Simulation results show that, when the GPU memory is 75% and 50% oversubscribed, CPPE achieves an average speedup of 1.56x and 1.64x (up to 10.97x) over the state-of-the-art baseline, which combines a sequential-local prefetcher and LRU pre-eviction policy. CPPE also outperforms other approaches, including Random/reserved LRU with the sequential-local prefetcher, and simply disabling prefetching under memory oversubscription.
Qi Yu 0003, Bruce R. Childers, Libo Huang 0002, Cheng Qian 0006, Hui Guo 0004, Zhiying Wang 0003
IPDPS2
2020 HPE: Hierarchical Page Eviction Policy for Unified Memory in GPUs
abstract
Recent support for unified memory and demand paging has improved graphics processing unit (GPU) programmability and enabled memory oversubscription. However, this support introduces high overhead when page faults occur. Therefore, when the GPU memory fills to capacity, an important issue is how to select eviction candidates. The widely used policy, LRU, and the advanced replacement policies, RRIP and CLOCK-Pro, suffer from inefficiency when dealing with thrashing access patterns. They also incur significant overhead due to managing metadata at page level. In this article, we propose hierarchical page eviction (HPE), a new replacement policy for GPUs with unified memory. Aided by page walk hit information, HPE manages a page set chain dynamically. It uses statistics to classify applications into three categories and selects an appropriate eviction strategy for each category. It also applies dynamic adjustment to switch the eviction strategy when necessary. The simulation results show that, on average, HPE achieves 1.34× and 1.16× speedup (up to 2.81×) over LRU when the oversubscription rate is 75% and 50%, respectively. HPE also outperforms RRIP and CLOCK-Pro.
Qi Yu 0003, Bruce R. Childers, Libo Huang 0002, Cheng Qian 0006, Zhiying Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 A quantitative evaluation of unified memory in GPUs
Qi Yu 0003, Bruce R. Childers, Libo Huang 0002, Cheng Qian 0006, Zhiying Wang 0003
J. Supercomput.2
2019 Hierarchical Page Eviction Policy for Unified Memory in GPUs
abstract
The introduction of unified memory in discrete GPUs not only improves programmability but also enables oversubscription. However, it introduces high overhead when page faults occur. Therefore, when GPU memory is full, how to select eviction candidates becomes an important issue. The widely used policy LRU performs poorly for workloads with thrashing access patterns, and the advanced cache replacement policy RRIP incurs thrashing when directly applied to GPU memory. In this paper, we propose hierarchical page eviction policy for GPU memory, which relies on a software-managed page set chain to select eviction candidates. Results show that for 15 selected applications, our policy achieves an average speedup of 1.44 and 1.2 over LRU when the oversubscription rate is 75% and 50 %, respectively.
Qi Yu 0003, Bruce R. Childers, Libo Huang 0002, Cheng Qian 0006, Zhiying Wang 0003
ISPASS2
2018 CMH: compression management for improving capacity in the hybrid memory cube
abstract
The Hybrid Memory Cube (HMC) is a novel 3D memory architecture that efficiently improves bandwidth and saves energy. However, due to limitations in scalability and power density of a DRAM bit cell, the physical data capacity of an individual HMC is relatively modest and unlikely to grow significantly and it is likely to be a challenge in adopting the HMC for big data in high-performance computing. In this paper, we propose a new strategy to increase the effective data capacity of the HMC, called Compression Management for HMC (CMH). CMH is incorporated in the logic layer of the HMC. By selectively compressing data during transmission and storing the selectively compressed data in the 3D memory stack, CMH increases data capacity while also improving effective bandwidth. For several memory-intensive benchmarks, our results show that CMH reduces pressure on memory capacity by 64.4%, and improves bandwidth by 42.4%. Similarly good results are observed for multi-programmed workloads, reducing capacity 66.2% and improving bandwidth 47.8%. Although compression has latency overhead, by introducing a small cache in the HMC logic layer to store metadata for compression, CMH mitigates any increase in transaction latency. The overhead in instructions per cycle is a minimal 1.2% and 1.5%, respectively, for single-core and multi-core workloads. The IPC is stable and is not harmed by the inclusion of compression.
Cheng Qian 0006, Libo Huang 0002, Qi Yu 0003, Zhiying Wang 0003, Bruce R. Childers
CF5
2018 Occam: Software Environment for Creating Reproducible Research
abstract
We have implemented an opensource prototype system called Occam1 to define, conduct, and share artifacts as executable content. Occam is a platform to create, run and share experiments. It allows users to contribute their own artifacts, which can be composed with other artifacts and used in workflows to define and conduct experiments. Workflows are directed acyclic graphs that describe the data flow within experiments. The execution of the experiment is automated by Occam, according to the workflow that describes it. Occam preserves provenance information and allows the inspection of the source code, configuration parameters, and datasets used in experiments; more importantly, it allows access to all this information from simply clicking on (the PDF-embedded) plots/results. In fact, with appropriate support in a digital library, we have implemented a mechanism in which clicking on a plot in a published article takes the user to the experiment setup in Occam. This strict approach to executable content preservation imposes little overhead on developers, but it greatly improves the preservation, reuse, and extensibility of software experiments. Consequently, Occam improves science by creating a tool and process that embodies the scientific method.
Luis Oliveira 0002, David Wilkinson, Daniel Mossé, Bruce R. Childers
eScience4
2018 Artifact Evaluation: FAD or Real News?
abstract
Data Management (DM), like many areas of computer science (CS), relies on empirical evaluation that uses software, data sets and benchmarks to evaluate new ideas and compare with past innovation. Despite the importance of these artifacts and associated information about experimental evaluations, few researchers make these available in a findable, accessible, interoperable and reusable (FAIR) manner, in this way hindering the scientific process by limiting open collaboration, credibility of published outcomes, and research progress. Fortunately, this problem is recognized and many CS communities, including the DM one, are advocating and providing incentives for software and analysis papers to follow FAIR principles and be treated equally to traditional publications. Some ACM/IEEE conferences adopted Artifact Evaluation (AE) to reward authors for doing a great job in conducting experiments with FAIR software and data. After half a decade since AE's inception, the question is whether the emerging emphasis on artifacts, is having a real impact in CS research.
Bruce R. Childers, Panos K. Chrysanthis
ICDE1
2018 HMCSP: Reducing Transaction Latency of CSR-based SPMV in Hybrid Memory Cube
abstract
Sparse Matrix Multiplication Vector (SPMV) plays a significant role in sparse linear algebra. Based on the high parallelization of matrix multiplication, SPMV has been accelerated with GPUs, Intel MIC, and FPGAs. The Micron Hybrid Memory Cube (HMC) is a highly parallel device that has atomic operations which support processing in memory (PIM). In this paper, we propose HMCSP, which extends the HMC's existing PIM capability to reduce the memory transaction latency of SPMV. By taking advantage of atomic operations and data prefetch, HMCSP reduces memory transaction latency of SPMV by 49.7% compared to a conventional HMC.
Cheng Qian 0006, Bruce R. Childers, Libo Huang 0002, Qi Yu 0003, Zhiying Wang 0003
ISPASS2
2017 DrMP: Mixed Precision-Aware DRAM for High Performance Approximate and Precise Computing
abstract
Recent studies showed that DRAM restore time degrades as technology scales, which imposes large performance and energy overheads. This problem, prolonged restore time (PRT), has been identified by the DRAM industry as one of three major scaling challenges. This paper proposes DrMP, a novel fine-grained precision-aware DRAM restore scheduling approach, to mitigate PRT. The approach exploits process variations (PVs) within and across DRAM rows to save data with mixed precision. The paper describes three variants of the approach: DrMP-A, DrMP-P, and DrMP-U. DrMP-A supports approximate computing by mapping important data bits to fast row segments to reduce restore time for improved performance at a low application error rate. DrMP-P pairs memory rows together to reduce the average restore time for precise computing. DrMP-U combines DrMP-A and DrMP-P to better trade performance, energy consumption, and computation precision. Our experimental results show that, on average, DrMP achieves 20% performance improvement and 15% energy reduction over a precision-oblivious baseline. Further, DrMP achieves an error rate less than 1% at the application level for a suite of benchmarks, including applications that exhibit unacceptable error rates under simple approximation that does not differentiate the importance of different bits.
Xianwei Zhang 0001, Youtao Zhang, Bruce R. Childers, Jun Yang 0002
PACT3
2017 Artifact Evaluation: Is It a Real Incentive?
abstract
It is well accepted that we learn hard lessons when implementing and re-evaluating systems, yet it is also acknowledged that science faces a crisis in reproducibility. Experimental computer science is far from immune, although it should be easier for CS than other sciences, given the emphasis on experimental artifacts, such as source code, data sets, workflows, parameters, etc. The data management community pioneered methods at ACM SIGMOD 2007 and 2008 to encourage and incentivize authors to improve their software development and experimental practices. Now, after 10 years, the broader CS community has started to adopt Artifact Evaluation (AE) to review artifacts along with papers. In this paper, we examine how AE has incentivized authors, and whether the process is having a measurable impact. Our answer can help guide CS, and more broadly, other computationally-oriented sciences, in encouraging peer-review of software artifacts and developing additional community practices for incentives.
Bruce R. Childers, Panos K. Chrysanthis
eScience1
2017 Quality of Service Support for Fine-Grained Sharing on GPUs
abstract
GPUs have been widely adopted in data centers to provide acceleration services to many applications. Sharing a GPU is increasingly important for better processing throughput and energy efficiency. However, quality of service (QoS) among concurrent applications is minimally supported. Previous efforts are too coarse-grained and not scalable with increasing QoS requirements. We propose QoS mechanisms for a fine-grained form of GPU sharing. Our QoS support can provide control over the progress of kernels on a per cycle basis and the amount of thread-level parallelism of each kernel. Due to accurate resource management, our QoS support has significantly better scalability compared with previous best efforts. Evaluations show that, when the GPU is shared by three kernels, two of which have QoS goals, the proposed techniques achieve QoS goals 43.8% more often than previous techniques and have 20.5% higher throughput.
Zhenning Wang, Jun Yang 0002, Rami G. Melhem, Bruce R. Childers, Youtao Zhang, Minyi Guo
ISCA4
2017 On the Restore Time Variations of Future DRAM Memory
abstract
As the de facto main memory standard, DRAM (Dynamic Random Access Memory) has achieved dramatic density improvement in the past four decades, along with the advancements in process technology. Recent studies reveal that one of the major challenges in scaling DRAM into the deep sub-micron regime is its significant variations on cell restore time, which affect timing constraints such as write recovery time. Adopting traditional approaches results in either low yield rate or large performance degradation. In this article, we propose schemes to expose the variations to the architectural level. By constructing memory chunks with different access speeds and, in particular, exploiting the performance benefits of fast chunks, a variation-aware memory controller can effectively mitigate the performance loss due to relaxed timing constraints. We then proposed restore-time-aware rank construction and page allocation schemes to make better use of fast chunks. Our experimental results show that, compared to traditional designs such as row sparing and Error Correcting Codes, the proposed schemes help to improve system performance by about 16% and 20%, respectively, for 20nm and 14nm technology nodes on a four-core multiprocessor system.
Xianwei Zhang 0001, Youtao Zhang, Bruce R. Childers, Jun Yang 0002
ACM Trans. Design Autom. Electr. Syst.3
2016 Simultaneous Multikernel GPU: Multi-tasking throughput processors via fine-grained sharing
abstract
Studies show that non-graphics programs can be less optimized for the GPU hardware, leading to significant resource under-utilization. Sharing the GPU among multiple programs can effectively improve utilization, which is particularly attractive to systems where many applications require access to the GPU (e.g., cloud computing). However, current GPUs lack proper architecture features to support sharing. Initial attempts are preliminary: They either provide only static sharing, which requires recompilation or code transformation, or they do not effectively improve GPU resource utilization. We propose Simultaneous Multikernel (SMK), a fine-grain dynamic sharing mechanism, that fully utilizes resources within a streaming multiprocessor by exploiting heterogeneity of different kernels. We propose several resource allocation strategies to improve system throughput while maintaining fairness. Our evaluation shows that for shared workloads with complementary resource occupancy, SMK improves GPU throughput by 52% over non-shared execution and 17% over a state-of-the-art design.
Zhenning Wang, Jun Yang 0002, Rami G. Melhem, Bruce R. Childers, Youtao Zhang, Minyi Guo
HPCA4
2016 Restore truncation for performance improvement in future DRAM systems
abstract
Scaling DRAM below 20nm has become a major challenge due to intrinsic limitations in the structure of a bit cell. Future DRAM chips are likely to suffer from significant variations and degraded timings, such as taking much more time to restore cell data after read and write access. In this paper, we propose restore truncation (RT), a low-cost restore strategy to improve performance of DRAM modules that adopt relaxed restore timing. After an access, RT restores a bit cell's voltage only to the level required to persist data to the next scheduled refresh rather than to the default full voltage. Because restore time is shortened, the performance of the cell is improved under process variations. We devise two schemes to balance performance, energy consumption, and hardware overhead. We simulate our proposed RT schemes and compare them with the state of the art. Experimental results show that, on average, RT improves performance by 19.5% and reduces energy consumption by 17%.
Xianwei Zhang 0001, Youtao Zhang, Bruce R. Childers, Jun Yang 0002
HPCA3
2016 Concurrent Migration of Multiple Pages in software-managed hybrid main memory
abstract
This paper describes Concurrent Migration of Multiple Pages (CMMP), a new hardware-software mechanism for managing hybrid main memory (DRAM+PCM). CMMP migrates multiple pages concurrently without significantly affecting the memory bandwidth available to applications. CMMP provides a simple interface for the OS to observe memory access patterns. CMMP reduces PCM-to-DRAM transfer bandwidth by copying blocks on-demand. It also reduces DRAM-to-PCM bandwidth by suppressing the transfer of untouched blocks back to PCM. Compared to a state-of-the-art page migration approach for hybrid memory, CMMP improves performance by 14% and reduces energy consumption by 29% on average.
Santiago Bock, Bruce R. Childers, Rami G. Melhem, Daniel Mossé
ICCD2
2016 Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems
abstract
In a multicore system, many applications share the last-level cache (LLC) and memory bandwidth. These resources need to be carefully managed in a coordinated way to maximize performance. DRAM is still the technology of choice in most systems. However, as traditional DRAM technology faces energy, reliability, and scalability challenges, nonvolatile memory (NVM) technologies are gaining traction. While DRAM is read/write symmetric (a read operation has comparable latency and energy consumption as a write operation), many NVM technologies (such as Phase-Change Memory, PCM) experience read/write asymmetry: write operations are typically much slower and more power hungry than read operations. Whether the memory’s characteristics are symmetric or asymmetric influences the way shared resources are managed. We propose two symmetry-agnostic schemes to manage a shared LLC through way partitioning and memory through bandwidth allocation. The proposals work well for both symmetric and asymmetric memory. First, an exhaustive search is proposed to find the best combination of a cache way partition and bandwidth allocation. Second, an approximate scheme, derived from a theoretical model, is proposed without the overhead of exhaustive search. Simulation results show that the approximate scheme improves weighted speedup by at least 14% on average (regardless of the memory symmetry) over a state-of-the-art way partitioning and memory bandwidth allocation. Simulation results also show that the approximate scheme achieves comparable weighted speedup as a state-of-the-art multiple resource management scheme, XChange, for symmetric memory, and outperforms it by an average of 10% for asymmetric memory.
Miao Zhou, Yu Du 0002, Bruce R. Childers, Daniel Mossé, Rami G. Melhem
ACM Trans. Archit. Code Optim.3
2015 Exploiting DRAM restore time variations in deep sub-micron scaling
Xianwei Zhang 0001, Youtao Zhang, Bruce R. Childers, Jun Yang 0002
DATE3
2015 Supporting superpages in non-contiguous physical memory
abstract
For memory-intensiv e workloads with large memory footprints, superpages are effective to avoid address translation overhead, which can be a critical performance bottleneck. A superpage is a large virtual memory page that is mapped to an equivalently-sized amount of contiguous physical memory pages. Superpage mapping assumes physical memory does not contain retired pages, which is an important technique to improve memory resilience: the OS avoids allocating physical pages that have detected errors. Retired pages create unusable "holes" in the physical memory. We show that even a small percentage of retired pages makes it very difficult to find enough contiguous memory to form superpages. To address this problem, we propose GTSM, or gap-tolerant sequential mapping, that allows superpages to be formed even in the presence of retired physical pages. A new page table format is also proposed to support GTSM. This format has similar storage efficiency as traditional superpaging to hold address translations in the last-level cache. To further compress the page table and improve cache hit rates for address translation in large memory footprint workloads, we also propose an extended format that reduces the page table size by 50%. In comparison to an ideal memory without any retired physical pages, we show that our technique, with retired pages, achieves nearly 96.8% of the performance of traditional 2MB superpaging.
Yu Du 0002, Miao Zhou, Bruce R. Childers, Daniel Mossé, Rami G. Melhem
HPCA3
2015 Characterizing the Overhead of Software-Managed Hybrid Main Memory
abstract
The size of main memory in modern computers is approaching energy and scalability limits. Combining DRAM and non-volatile memory (NVM) has been proposed to increase capacity and reliability, and to decrease energy consumption. Software-managed hybrid memory is a promising way to incorporate NVM in main memory due to its architectural simplicity. However, there are significant performance issues caused by interference due to data migration between DRAM and NVM and a lack of effective migration policies. To aid in the development of migration policies and hardware mechanisms for incorporating NVM in main memory, we propose new analysis and simulation techniques to understand the behavior of software-managed hybrid memory. These techniques allow us to characterize the overhead experienced by requests in the memory hierarchy and identify the factors that limit performance in software-managed hybrid memory. Using our techniques, we show that queuing delays at the NVM banks and NVM bus are the main limiting factors, and that there is significant potential to improve performance with better migration policies.
Santiago Bock, Bruce R. Childers, Rami G. Melhem, Daniel Mossé
MASCOTS2
2014 Program affinity performance models for performance and utilization
abstract
Multithreaded applications have a wide variety of behavior, causing complex interactions with today's chip multiprocessor machines. Application threads may have large private working sets, and may compete for cache space and memory bandwidth. These threads benefit from large private caches. Other threads may share data or communicate, and thus, execute more quickly if using shared caches. Many applications fall somewhere in between, requiring careful thread-to-core assignments to maximize performance. Yet because of the large number of thread-to-core assignments on today's chip multiprocessors, it is time and energy prohibitive to exhaustively try and determine the best assignment. In this paper, we present and demonstrate application performance models that predict application performance given a proposed thread-to-core assignment. We show how these models can be quickly built and used to select thread-to-core assignments for multiple programs and to improve system utilization.
Ryan W. Moore, Bruce R. Childers
DATE2
2014 COMeT+: Continuous Online Memory Testing with Multi-Threading Extension
abstract
Today’s computers have gigabytes of main memory due to improved DRAM density. As density increases, smaller bit cells become more susceptible to errors. With an increase in error susceptibility, the need for memory resiliency also increases. Self-testing of memory health can proactively check for errors to improve resiliency. This paper describes a software-only self-test to continuously test memory. We present the challenges and design for an approach, called Continuous Online Memory Testing with Multi-threading Extension (COMeT+), that targets chip multiprocessors. COMeT+ tests memory health simultaneously with execution of single and multi-threaded applications in anticipation of allocation requests. The approach guarantees that memory is tested within a fixed time interval to limit exposure to lurking errors. We developed and evaluated an implementation of COMeT+. On the SPEC CPU2006 and the PARSEC benchmarks, COMeT+ has a low 4% average performance overhead. On the PARSEC benchmarks, the effect of TLB shootdowns on application performance due to additional page migrations caused by COMeT+ was insignificant. When emulated errors were injected into physical memory, applications executed 1.13x to 4.41x longer with COMeT+ than without it.
Musfiq Rahman, Bruce R. Childers, Sangyeun Cho
IEEE Trans. Computers2
2014 Building and using application utility models to dynamically choose thread counts
Ryan W. Moore, Bruce R. Childers
J. Supercomput.2
2013 Writeback-aware bandwidth partitioning for multi-core systems with PCM
abstract
Phase-Change Memory (PCM) has emerged as a promising low-power candidate to replace DRAM in main memory. Hybrid memory architecture comprised of a large PCM and a small DRAM is a popular solution to mitigate undesirable characteristics of PCM writes. Because PCM writes are much slower than reads, writebacks from the last-level cache consume a large portion of memory bandwidth, and thus, impact performance. Effectively utilizing shared resources, such as the last-level cache and the memory bandwidth, is crucial to achieving high performance for multi-core systems. Although existing memory bandwidth allocation schemes improve system performance, no current approach uses writeback information to partition bandwidth for hybrid memory. We use a writeback-aware analytic model to derive the allocation strategy for bandwidth partitioning of phase-change memory. From the derivation of the model, Writeback-aware Bandwidth Partitioning (WBP) is proposed as a new runtime mechanism to partition PCM service cycles among applications. WBP uses a partitioning weight to indicate the importance of writebacks (in addition to LLC misses) to bandwidth allocation. A companion Dynamic Weight Adjustment (DWA) scheme dynamically selects the partitioning weight to maximize system performance. Simulation results show that WBP and DWA improve performance by 24.9% (weighted speedup) over bandwidth partitioning schemes that do not take writebacks into consideration in a 8-core system.
Miao Zhou, Yu Du 0002, Bruce R. Childers, Rami G. Melhem, Daniel Mossé
PACT3
2013 Automatic Generation of Program Affinity Policies Using Machine Learning
Ryan W. Moore, Bruce R. Childers
CC2
2013 Bit mapping for balanced PCM cell programming
abstract
Write bandwidth is an inherent performance bottleneck for Phase Change Memory (PCM) for two reasons. First, PCM cells have long programming time, and second, only a limited number of PCM cells can be programmed concurrently due to programming current and write circuit constraints,
Yu Du 0002, Miao Zhou, Bruce R. Childers, Daniel Mossé, Rami G. Melhem
ISCA3
2013 Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory
abstract
Limited PCM write bandwidth is a critical obstacle to achieve good performance from hybrid DRAM/PCM memory systems. The write bandwidth is severely restricted in PCM devices, which harms application performance. Indeed, as we show, it is more important to reduce PCM write traffic than to reduce PCM read latency for application performance. To reduce the number of PCM writes, we propose a DRAM cache organization that employs compression. A new delta compression technique for modified data is used to achieve a large compression ratio. Our approach can selectively and predictively apply compression to improve its efficiency and performance. Our approach is designed to facilitate adoption in existing main memory compression frameworks. We describe an instance of how to incorporate delta compression in IBM's MXT memory compression architecture when used for DRAM cache in a hybrid main memory. For fourteen representative memory-intensive workloads, on average, our delta compression technique reduces the number of PCM writes by 54.3%, and improves IPC performance by 24.4%.
Yu Du 0002, Miao Zhou, Bruce R. Childers, Rami G. Melhem, Daniel Mossé
ACM Trans. Archit. Code Optim.3
2013 Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory
abstract
Phase Change Memory (PCM) has recently emerged as a promising memory technology. However, PCM’s limited write endurance restricts its immediate use as a replacement for DRAM. To extend the lifetime of PCM chips, wear-leveling and salvaging techniques have been proposed. Wear-leveling balances write operations across different PCM regions while salvaging extends the duty cycle and provides graceful degradation for a nonnegligible number of failures. Current wear-leveling and salvaging schemes have not been designed and integrated to work cooperatively to achieve the best PCM device lifetime. In particular, a noncontiguous PCM space generated from salvaging complicates wear-leveling and incurs large overhead. In this article, we propose LLS, a Line-Level mapping and Salvaging design. By allocating a dynamic portion of total space in a PCM device as backup space, and mapping failed lines to backup PCM, LLS constructs a contiguous PCM space and masks lower-level failures from the OS and applications. LLS integrates wear-leveling and salvaging and copes well with modern OSes. Our experimental results show that LLS achieves 31% longer lifetime than the state-of-the-art. It has negligible hardware cost and performance overhead.
Lei Jiang 0001, Yu Du 0002, Bo Zhao 0007, Youtao Zhang, Bruce R. Childers, Jun Yang 0002
ACM Trans. Archit. Code Optim.5
2012 Improving write operations in MLC phase change memory
abstract
Phase change memory (PCM) recently has emerged as a promising technology to meet the fast growing demand for large capacity memory in modern computer systems. In particular, multi-level cell (MLC) PCM that stores multiple bits in a single cell, offers high density with low per-byte fabrication cost. However, despite many advantages, such as good scalability and low leakage, PCM suffers from exceptionally slow write operations, which makes it challenging to be integrated in the memory hierarchy. In this paper, we propose architectural innovations to improve the access time of MLC PCM. Due to cell process variation, composition fluctuation and the relatively small differences among resistance levels, MLC PCM typically employs an iterative write scheme to achieve precise control, which suffers from large write access latency. To address this issue, we propose write truncation (WT) to reduce the number of write iterations with the assistance of an extra error correction code (ECC). We also propose form switch (FS) to reduce the storage overhead of the ECC. By storing highly compressible lines in SLC form, FS improves read latency as well. Our experimental results show that WT and FS improve the effective write/read latency by 57%/28% respectively, and achieve 26% performance improvement over the state of the art.
Lei Jiang 0001, Bo Zhao 0007, Youtao Zhang, Jun Yang 0002, Bruce R. Childers
HPCA5
2012 Using utility prediction models to dynamically choose program thread counts
abstract
Multithreaded applications can simultaneously execute on a chip multiprocessor computer, starting and stopping without warning or pattern. The behavior of each program can be different, interacting in unexpected ways, including causing competition for CPU cycles, which harms performance.
Ryan W. Moore, Bruce R. Childers
ISPASS2
2012 FPB: Fine-grained Power Budgeting to Improve Write Throughput of Multi-level Cell Phase Change Memory
abstract
As a promising nonvolatile memory technology, Phase Change Memory (PCM) has many advantages over traditional DRAM. Multi-level Cell PCM (MLC) has the benefit of increased memory capacity with low fabrication cost. Due to high per-cell write power and long write latency, MLC PCM requires careful power management to ensure write reliability. Unfortunately, existing power management schemes applied to MLC PCM result in low write throughput and large performance degradation. In this paper, we propose Fine-grained write Power Budgeting (FPB) for MLC PCM. We first identify two major problems for MLC write operations: (i) managing write power without consideration of the iterative write process used by MLC is overly pessimistic, (ii) a heavily written (hot) chip may block the memory from accepting further writes due to chip power restrictions, although most chips may be available. To address these problems, we propose two FPB schemes. First, FPB-IPM observes a global power budget and regulates power across write iterations according to the step-down power demand of each iteration. Second, FPB-GCP integrates a global charge pump on a DIMM to boost power for hot PCM chips while staying within the global power budget. Our experimental results show that these techniques achieve significant improvement on write throughput and system performance. Our schemes also interact positively with PCM effective read latency reduction techniques, such as write cancellation, write pausing and write truncation.
Lei Jiang 0001, Youtao Zhang, Bruce R. Childers, Jun Yang 0002
MICRO3
2012 REEact: a customizable virtual execution manager for multicore platforms
abstract
With the shift to many-core chip multiprocessors (CMPs), a critical issue is how to effectively coordinate and manage the execution of applications and hardware resources to overcome performance, power consumption, and reliability challenges stemming from hardware and application variations inherent in this new computing environment. Effective resource and application management on CMPs requires consideration of user/application/hardware-specific requirements and dynamic adaption of management decisions based on the actual run-time environment. However, designing an algorithm to manage resources and applications that can dynamically adapt based on the run-time environment is difficult because most resource and application management and monitoring facilities are only available at the operating system level. This paper presents REEact, an infrastructure that provides the capability to specify user-level management policies with dynamic adaptation. REEact is a virtual execution environment that provides a framework and core services to quickly enable the design of custom management policies for dynamically managing resources and applications. To demonstrate the capabilities and usefulness of REEact, this paper describes three case studies--each illustrating the use of REEact to apply a specific dynamic management policy on a real CMP. Through these case studies, we demonstrate that REEact can effectively and efficiently implement policies to dynamically manage resources and adapt application execution.
Wei Wang 0054, Tanima Dey, Ryan W. Moore, Mahmut Aktasoglu, Bruce R. Childers, Jack W. Davidson, Mary Jane Irwin, Mahmut T. Kandemir, Mary Lou Soffa
VEE5
2012 Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems
abstract
Phase-Change Memory (PCM) has emerged as a promising low-power main memory candidate to replace DRAM. The main problems of PCM are that writes are much slower and more power hungry than reads, write bandwidth is much lower than read bandwidth, and limited write endurance. Adding an extra layer of cache, which is logically the last-level cache (LLC), can mitigate the drawbacks of PCM. However, writebacks from the LLC might (a) overwhelm the limited PCM write bandwidth and stall the application, (b) shorten lifetime, and (c) increase energy consumption. Cache partitioning and replacement schemes are important to achieve high throughput for multi-core systems. However, we noted that no existing partitioning and replacement policy takes into account the writeback information. This paper proposes two writeback-aware schemes to manage the LLC for PCM main memory systems. Writeback-aware Cache Partitioning (WCP) is a runtime mechanism that partitions a shared LLC among multiple applications. Unlike past partitioning schemes, our scheme considers the reduction in cache misses as well as writebacks. Write Queue Balancing (WQB) replacement policy manages the cache partition of each application intelligently so that the writebacks are distributed evenly among PCM write queues. In this way, applications rarely stall due to unbalanced PCM write traffic among write queues. Our evaluation shows that WCP and WQB result in, on average, 21% improvement in throughput, 49% reduction in PCM writes, and 14% reduction in energy over a state-of-the-art cache partitioning scheme.
Miao Zhou, Yu Du 0002, Bruce R. Childers, Rami G. Melhem, Daniel Mossé
ACM Trans. Archit. Code Optim.3
2012 Enabling dynamic binary translation in embedded systems with scratchpad memory
abstract
Important challenges for embedded systems can be addressed by dynamic binary translation. A dynamic binary translator stores translated instructions in a software-managed code cache, which is usually large to minimize overhead. This article shows how to use a small scratchpad memory for the code cache. A small code cache may require frequent code evictions and retranslation, which degrade performance. We propose techniques to reduce the number of instructions inserted by the translator and a way to form fragments that minimizes translated code size. With our techniques, a much smaller code cache can hold a program's translated code working set.
José Baiocchi, Bruce R. Childers, Jack W. Davidson, Jason Hiser
ACM Trans. Embed. Comput. Syst.2
2011 Demand code paging for NAND flash in MMU-less embedded systems
abstract
NAND flash is preferred for code and data storage in embedded devices due to its high density and low cost. However, NAND flash requires code to be copied to main memory for execution. In inexpensive devices without hardware memory management, full shadowing of an application binary is commonly used to load the program. This approach can lead to a high initial application start-up latency and poor amortization of copy overhead. To overcome these problems, we describe a software-only demand-paging approach that incrementally copies code to memory with a dynamic binary translator (DBT). This approach does not require hardware or operating system support. With careful management, a savings can be achieved in total code footprint, which can offset the size of data structures used by DBT. For applications that cannot amortize full shadowing cost, our approach can reduce start-up latency by 50% or more, and improve performance by 11% on average.
José Baiocchi, Bruce R. Childers
DATE2
2011 Impact of process variation on endurance algorithms for wear-prone memories
abstract
Non-volatile memories, such as Flash and Phase-Change Memory, are replacing other memory and storage technologies. Although these new technologies have desirable energy and scalability properties, they are prone to wear-out due to excessive write operations. Because wear-out is an important phenomenon, a number of endurance management schemes have been proposed. There is a trade-off between what techniques to use, depending on the range of bit cell lifetime within a device. This range in cell durability arises from effects due to process variation. In this paper, we describe modeling techniques to analyze trade-offs for endurance management based on the anticipated distribution of cell lifetime. This analysis considers two general endurance strategies (physical capacity degradation and physical sparing) under four distributions of cell lifetime (constant, linear, normal, and bimodal). The modeling techniques can be used to determine how much redundancy is needed when a sparing endurance strategy is adopted. With the correct choice of technique, the device lifetime can be doubled.
Alexandre Peixoto Ferreira, Santiago Bock, Bruce R. Childers, Rami G. Melhem, Daniel Mossé
DATE3
2011 LLS: Cooperative integration of wear-leveling and salvaging for PCM main memory
abstract
Phase change memory (PCM) has emerged as a promising technology for main memory due to many advantages, such as better scalability, non-volatility and fast read access. However, PCM's limited write endurance restricts its immediate use as a replacement for DRAM. Recent studies have revealed that a PCM chip which integrates millions to billions of bit cells has non-negligible variations in write endurance. Wear leveling techniques have been proposed to balance write operations to different PCM regions. To further prolong the lifetime of a PCM device after the failure of weak cell, techniques have been proposed to remap failed lines to spares and to salvage a PCM device that has a large number of failed lines or pages with graceful degradation. However, current wear-leveling and salvaging schemes have not been designed and integrated to work cooperatively to achieve the best PCM device lifetime. In particular, a non-contiguous PCM space generated from salvaging complicates wear leveling and incurs large overhead. In this paper, we propose LLS, a Line-Level mapping and Salvaging design. By allocating a dynamic portion of total space in a PCM device as backup space, and mapping failed lines to backup PCM, LLS constructs a contiguous PCM space and masks lower-level failures from the OS and applications. LLS seamlessly integrates wear leveling and salvaging and copes well with modern OSs, including ones that support multiple page sizes. Our experimental results show that LLS achieves 24% longer lifetime than a state-of-the-art technique. It has negligible hardware cost and performance overhead.
Lei Jiang 0001, Yu Du 0002, Youtao Zhang, Bruce R. Childers, Jun Yang 0002
DSN4
2011 CloudCache: Expanding and shrinking private caches
abstract
The number of cores in a single chip multiprocessor is expected to grow in coming years. Likewise, aggregate on-chip cache capacity is increasing fast and its effective utilization is becoming ever more important. Furthermore, available cores are expected to be underutilized due to the power wall and highly heterogeneous future workloads. This trend makes existing L2 cache management techniques less effective for two problems: increased capacity interference between working cores and longer L2 access latency. We propose a novel scalable cache management framework called CloudCache that creates dynamically expanding and shrinking L2 caches for working threads with fine-grained hardware monitoring and control. The key architectural components of CloudCache are L2 cache chaining, inter- and intra-bank cache partitioning, and a performance-optimized coherence protocol. Our extensive experimental evaluation demonstrates that CloudCache significantly improves performance of a wide range of workloads when all or a subset of cores are occupied.
Hyunjin Lee 0006, Sangyeun Cho, Bruce R. Childers
HPCA3
2011 Analyzing the impact of useless write-backs on the endurance and energy consumption of PCM main memory
abstract
Phase Change Memory (PCM) is an emerging technology that has been recently considered as a cost-effective and energy-efficient alternative to traditional DRAM main memory. Due to the high energy consumption of writes and limited number of write cycles, reducing the number of writes to PCM can result in considerable energy savings and endurance improvement. In this paper, we introduce the concept of useless write-backs, which occur when a dirty cache line that belongs to a dead memory region is evicted from the cache (a dead region is a memory location that is not used again by a program). Since the evicted data is not used again, the write-back can be safely avoided to improve endurance and energy consumption. This paper presents a limit study on the improvement that passing information to the memory system about useless writebacks has on the endurance and energy consumption of systems based on PCM main memory. We developed algorithms to measure the number of useless write-backs to PCM for three different types of memory regions and we present an energy model to determine the maximum energy savings that could potentially be achieved through such a scheme. Our results show that avoiding useless write-backs can save up to 19.8% of energy and improve endurance by up to 26.2%.
Santiago Bock, Bruce R. Childers, Rami G. Melhem, Daniel Mossé, Youtao Zhang
ISPASS2
2011 COMeT: Continuous Online Memory Test
abstract
Today's computers have gigabytes of main memory due to improved DRAM density. As density increases, smaller bit cells become more susceptible to errors. With an increase in error susceptibility, the need for memory resiliency also increases. Self-testing of memory health can proactively check for errors to improve resiliency. This paper describes a software-only self test to continuously test memory. We present the challenges and design for an approach, called Continuous Online Memory Testing (COMeT), that targets chip multiprocessors. COMeT tests memory health simultaneously with application execution in anticipation of allocation requests. The approach guarantees that memory is tested within a fixed time interval to limit exposure to lurking errors. We developed and evaluated an implementation of COMeT. On the SPEC CPU2006 benchmarks, COMeT has a low 4% average performance overhead. When emulated errors were injected into physical memory, applications executed 1.13× to 4.41× longer with COMeT than without it.
Musfiq Rahman, Bruce R. Childers, Sangyeun Cho
PRDC2
2011 Real-Time Scheduling for Phase Change Main Memory Systems
abstract
Multi-core processors are effective for reducing energy consumption in computer systems, since modern multi- core chips allow for power management of individual cores. However, multiple cores impose higher demand on the memory subsystem, which is extremely power hungry. In addition to the small steps towards managing power in DRAMs, Phase-Change Memory (PCM) has emerged as a low-power alternative that is especially helpful for energy-aware embedded real-time systems. However, there are three drawbacks to PCM: its high latency, high energy consumption when writing, and low endurance. In real-time systems, the impact of PCM's high access latency is of special interest, as it has a negative effect on the number of deadlines that are met by the system. In this paper, we examine the memory subsystem and add a real-time scheduler for prioritizing requests at the bottleneck resource, the PCM controller. Adding support for external priorities, we use rate monotonic (RM) and earliest deadline first (EDF) prioritization at the PCM and show that it does reduce the number of deadline misses, but not sufficiently. We examine two additional schemes for prioritizing PCM requests (critical read boosting and read over write). We show that the scheduler of the PCM controller has a significant influence on the percentage of missed deadlines: critical read boosting and read over write can reduce the percentage of missed deadlines by 80% in the best case with negligible energy overhead.
Miao Zhou, Santiago Bock, Alexandre Peixoto Ferreira, Bruce R. Childers, Rami G. Melhem, Daniel Mossé
TrustCom4
2011 Evaluating indirect branch handling mechanisms in software dynamic translation systems
abstract
Software Dynamic Translation (SDT) is used for instrumentation, optimization, security, and many other uses. A major source of SDT overhead is the execution of code to translate an indirect branch's target address into the translated destination block's address. This article discusses sources of Indirect Branch (IB) overhead in SDT systems and evaluates techniques for overhead reduction. Measurements using SPEC CPU2000 show that the appropriate choice and configuration of IB translation mechanisms can significantly reduce the overhead. Further, cross-architecture evaluation of these mechanisms reveals that the most efficient implementation and configuration can be highly dependent on the architecture implementation.
Jason Hiser, Daniel W. Williams, Jack W. Davidson, Jason Mars, Bruce R. Childers
ACM Trans. Archit. Code Optim.6
2011 DEFCAM: A design and evaluation framework for defect-tolerant cache memories
abstract
Advances in deep submicron technology call for a careful review of existing cache designs and design practices in terms of yield, area, and performance. This article presents a Design and Evaluation Framework for defect-tolerant Cache Memories (DEFCAM), which enables processor architects to consider yield, area, and performance together in a unified framework. Since there is a complex, changing trade-off among these metrics depending on the technology, the cache organization, and the yield enhancement scheme employed, such a design flow is invaluable to processor architects when they assess a design and explore the design space quickly at an early stage. We develop a complete framework supporting the proposed DEFCAM design flow, from injecting defects into a wafer to evaluating program performance of individual processors on the wafer. Using DEFCAM, interesting interactions between architectural, organizational, and layout/defect related parameters can be easily evaluated. Moreover, we propose practical set remapping schemes to contain hard faults in cache memory. In a set remapping scheme, accesses that would go to an unusable faulty set are directed to a sound set. Case studies are presented to demonstrate the effectiveness of the proposed design flow and developed tools. Experimental results show that a set remapping is the most efficient fault covering method among prevailing strategies.
Hyunjin Lee 0006, Sangyeun Cho, Bruce R. Childers
ACM Trans. Archit. Code Optim.3
2010 Increasing PCM main memory lifetime
abstract
The introduction of Phase-Change Memory (PCM) as a main memory technology has great potential to achieve a large energy reduction. PCM has desirable energy and scalability properties, but its use for main memory also poses challenges such as limited write endurance with at most 107writes per bit cell before failure. This paper describes techniques to enhance the lifetime of PCM when used for main memory. Our techniques are (a) writeback minimization with new cache replacement policies, (b) avoidance of unnecessary writes, which write only the bit cells that are actually changed, and (c) endurance management with a novel PCM-aware swap algorithm for wear-leveling. A failure detection algorithm is also incorporated to improve the reliability of PCM. With these approaches, the lifetime of a PCM main memory is increased from just a few days to over 8 years.
Alexandre Peixoto Ferreira, Miao Zhou, Santiago Bock, Bruce R. Childers, Rami G. Melhem, Daniel Mossé
DATE4
2010 StimulusCache: Boosting performance of chip multiprocessors with excess cache
abstract
Technology advances continuously shrink on-chip devices. Consequently, the number of cores in a single chip multiprocessor (CMP) is expected to grow in coming years. Unfortunately, with smaller device size and greater integration, chip yield degrades significantly. Guaranteeing that all chip components function correctly leads to an unrealistically low yield. Chip vendors have adopted a design strategy to market partially functioning processor chips to combat this problem. The two major components in a multicore chip are compute cores and on-chip memory such as L2 cache. From the viewpoint of the chip yield, the compute cores have a much lower yield than the on-chip memory due to their logic complexity and well-established memory yield enhancing techniques. Therefore, future CMPs are expected to have more available on-chip memories than working cores. This paper introduces a novel on-chip memory utilization scheme called StimulusCache, which decouples the L2 caches of faulty compute cores and employs them to assist applications on other working cores. Our extensive experimental evaluation demonstrates that StimulusCache significantly improves the performance of both single-threaded and multithreaded workloads.
Hyunjin Lee 0006, Sangyeun Cho, Bruce R. Childers
HPCA3
2010 Using PCM in Next-generation Embedded Space Applications
abstract
Dynamic RAM (DRAM) has been the best technology for main memory for over thirty years. In embedded space applications, radiation hardened DRAM is needed because gamma rays cause transient errors; such rad-hard memories are extremely expensive and power hungry, leading to lower life (or increased battery weight) for satellite and other devices operating in space. Despite these problems, DRAM has been the technology of choice because it has better performance and it scales well. New, more energy efficient, non-volatile, scalable, radiation resistant memory technologies are now available, namely phase-change memory (PCM), making the DRAM choice much less compelling. However, current approaches require changes to PCM device internal circuitry, the operating system and/or the CPU cache-memory organization/interface. This paper presents a new, practical, detailed architecture, called PMMA, to effectively use PCM for main memory in next-generation embedded space systems. We designed PMMA avoiding changes to commodity PCM devices, the operating system, and the existing CPU cache-memory interface, enabling plug-in replacement of a conventional DRAM main memory by one constructed with PMMA. Our architecture incorporates novel mechanisms to address PCM’s limitations including expensive write operations, asymmetric read/write latency, and limited endurance. In our evaluation we show that PMMA achieves a 60% improvement in energy-delay over a conventional DRAM main memory.
Alexandre Peixoto Ferreira, Bruce R. Childers, Rami G. Melhem, Daniel Mossé, Mazin Yousif
IEEE Real-Time and Embedded Technology and Applications Symposium2
2010 StealthWorks: Emulating Memory Errors
Musfiq Rahman, Bruce R. Childers, Sangyeun Cho
RV2
2010 PERFECTORY: A Fault-Tolerant Directory Memory Architecture
abstract
The number of CPUs in chip multiprocessors is growing at the Moore's Law rate, due to continued technology advances. However, new technologies pose serious reliability challenges, such as more frequent occurrences of degraded or even nonoperational devices, and they threaten the cost-effectiveness and dependability of future computing systems. This work studies how to protect the on-chip coherence directory from fault occurrences. In a chip multiprocessor, cache coherence mechanisms such as directory memory are critical for offering consistent data view to all CPUs. We propose a novel online fault detection and correction scheme to enhance yield and resilience to runtime errors at a small performance cost. The proposed scheme uses smart encoding and coherence protocol adaptation strategies to salvage faulty directory entries. We also develop an online error recovery scheme that protects the directory memory from soft errors. We call our fault-tolerant directory memory architecture PERFECTORY. Evaluation results show that PERFECTORY achieves very high fault resilience: Over 99 percent chip yield at 0.05 percent hard error ratio and 1,934 years MTTF at 1,000 FIT using a 100-processor cluster configuration. PERFECTORY limits performance degradation to less than 1 percent at 0.05 percent hard error ratio and requires significantly smaller area overheads than existing redundancy approaches.
Hyunjin Lee 0006, Sangyeun Cho, Bruce R. Childers
IEEE Trans. Computers3
2010 Detecting bugs in register allocation
abstract
Although register allocation is critical for performance, the implementation of register allocation algorithms is difficult, due to the complexity of the algorithms and target machine architectures. It is particularly difficult to detect register allocation errors if the output code runs to completion, as bugs in the register allocator can cause the compiler to produce incorrect output code. The output code may even execute properly on some test data, but errors can remain. In this article, we propose novel data flow analyses to statically check that the value flow of the output code from the register allocator is the same as the value flow of its input code. The approach is accurate, fast, and can identify and report error locations and types. It is independent of the register allocator and uses only the input and output code of the register allocator. It can be used with different register allocators, including those that perform coalescing and rematerialization. The article describes our approach, called SARAC, and a tool that statically checks a register allocation and reports the errors and their types that it finds. The tool has an average compile-time overhead of only 8% and a modest average memory overhead of 85KB. Our techniques can be used by compiler developers during regression testing and as a command-line-enabled debugging pass for mysterious compiler behavior.
Yuqiang Huang, Bruce R. Childers, Mary Lou Soffa
ACM Trans. Program. Lang. Syst.2
2009 A Framework for Exploring Optimization Properties
Min Zhao 0009, Bruce R. Childers, Mary Lou Soffa
CC2
2009 Transparent Debugging of Dynamically Optimized Code
abstract
Debugging programs at the source level is essential in the software development cycle. With the growing importance of dynamic optimization, there is a clear need for debugging support in the presence of runtime code transformation. This paper presents a framework, called DeDoc, and lightweight techniques that allow debugging at the source level for programs that have been transformed by a trace-based binary dynamic optimizer. Our techniques provide full transparency and hide from the user the effect of dynamic optimizations on code statements and data values. We describe and evaluate an implementation of DeDoc and its techniques that interface a dynamic optimizer with a native debugger. Our experimental results indicate that DeDoc is able to report over 96% of values, that are otherwise not reportable due to code transformations, and incurs less than 1% performance overhead.
Naveen Kumar 0002, Bruce R. Childers, Mary Lou Soffa
CGO2
2009 Heterogeneous code cache: using scratchpad and main memory in dynamic binary translators
abstract
Dynamic binary translation (DBT) can be used to address important issues in embedded systems. DBT systems store translated code in a software-managed code cache. Unlike general-purpose systems, embedded systems often have specialized memory resources, such as a fast scratchpad memory, that can be used to mitigate DBT performance overhead. This paper presents the Heterogeneous Code Cache (HCC), a code cache split among scratchpad and main memory. We explore several HCC management policies and show that, on average, an HCC outperforms a code cache allocated only to scratchpad or only to main memory.
José Baiocchi, Bruce R. Childers
DAC2
2009 MCP: An Energy-Efficient Code Distribution Protocol for Multi-Application WSNs
Youtao Zhang, Bruce R. Childers
DCOSS3
2009 Addressing the challenges of DBT for the ARM architecture
abstract
Dynamic binary translation (DBT) can provide security, virtualization, resource management and other desirable services to embedded systems. Although DBT has many benefits, its run-time performance overhead can be relatively high. The run-time overhead is important in embedded systems due to their slow processor clock speeds,simple microarchitectures, and small caches.This paper addresses how to implement efficient DBT for ARM-based embedded systems, taking into account instruction set and cache/TLB nuances. We develop several techniques that reduce DBT overhead for the ARM. Our techniques focus on cache and TLB behavior. We tested the techniques on an ARM-based embedded device and found that DBT overhead was reduced by 54 % in comparison to a general-purpose DBT configuration that is known to perform well, thus further enabling DBT for a wide range of purposes. Categories and Subject Descriptors C.3 [Computer Systems Organization]: Special-purpose and application- based systems–Realtime and embedded systems; D.3.4 [Programming Languages]: Processors–Code generation, Compilers, Incremental compilers,
Ryan W. Moore, José Baiocchi, Bruce R. Childers, Jack W. Davidson, Jason Hiser
LCTES3
2008 Reducing pressure in bounded DBT code caches
abstract
Dynamic binary translators (DBT) have recently attracted much attention for embedded systems. The effective implementation of DBT in these systems is challenging due to tight constraints on memory and performance. A DBT uses a software-managed code cache to hold blocks of translated code. To minimize overhead, the code cache is usually large so blocks are translated once and never discarded. However, an embedded system may lack the resources for a large code cache. This constraint leads to significant slowdowns due to the retranslation of blocks prematurely discarded from a small code cache. This paper addresses the problem and shows how to impose a tight size bound on the code cache without performance loss. We show that about 70 % of the code cache is consumed by instructions that the DBT introduces for its own purposes. Based on this observation, we propose novel techniques that reduce the amount of space required by DBT-injected code, leaving more room for actual application code and improving the miss ratio. We experimentally demonstrate that a bounded code cache can have performance on-par with an unbounded one. Categories and Subject Descriptors C.3 [Computer Systems Organization]: Special-purpose and a-pplication-based systems—Real-time and embedded systems; D.3.4 [Programming Languages]: Processors—Code generation, Compilers,
José Baiocchi, Bruce R. Childers, Jack W. Davidson, Jason Hiser
CASES2
2008 Adaptive Buffer Management for Efficient Code Dissemination in Multi-Application Wireless Sensor Networks
abstract
Future wireless sensor networks (WSNs) are projected to run multiple applications in the same network infrastructure. While such multi-application WSNs (MA-WSNs) are economically more efficient and adapt better to the changing environments than traditional single-application WSNs, they usually require frequent code redistribution on wireless sensors, making it critical to design energy efficient post-deployment code dissemination protocols in MA-WSNs. Different applications in MA-WSNs often share some common code segments. Therefore when there is a need to disseminate a new application from the sink node, it is possible to disseminate its shared code segments from peer sensors instead of disseminating everything from the sink node. While dissemination protocols have been proposed to handle code of each single type, it is challenging to achieve energy efficiency when the code contains both types and needs simultaneous dissemination. In this paper we utilize an adaptive buffer management approach to achieve efficient code dissemination in MA-WSNs. Our experimental results show that adaptive buffer management can reduce the completion time and the message overhead up to 10% and 20% respectively.
Yu Du 0002, Youtao Zhang, Bruce R. Childers, Jun Yang 0002
EUC (1)4
2008 Integrated CPU Cache Power Management in Multiple Clock Domain Processors
Nevine AbouGhazaleh, Bruce R. Childers, Daniel Mossé, Rami G. Melhem
HiPEAC2
2008 Running a Java VM inside an operating system kernel
abstract
Operating system extensions have been shown to be beneficial to implement custom kernel functionality. In most implementations, the extensions are made by an administrator with kernel loadable modules. An alternative approach is to provide a run-time system within the operating system itself that can execute user kernel extensions. In this paper, we describe such an approach,where a lightweight Java virtual machine is embedded within the kernel for flexible extension of kernel network I/O. For this purpose, we first implemented a compact Java Virtual Machine with a Just-In-Time compiler on the Intel IA32 instruction set architecture at the user space. Then, the virtual machine was embedded onto the FreeBSDoperating system kernel. We evaluate the system to validate the model, with systematic benchmarking.
Takashi Okumura, Bruce R. Childers, Daniel Mossé
VEE2
2008 Preface
Bruce R. Childers, Mahmut T. Kandemir
Comput. Lang. Syst. Struct.1
2007 Fragment cache management for dynamic binary translators in embedded systems with scratchpad
abstract
Dynamic binary translation (DBT) has been used to achieve numerous goals (e.g., better performance) for general-purpose computers. Recently, DBT has also attracted attention for embedded systems. However, a challenge to DBT in this domain is stringent constraints on memory and performance. The translated code buffer used by DBT may occupy too much memory space. This paper proposes novel schemes to manage this buffer with scratchpad memory. We use footprint reduction to minimize the space needed by the translated code, victim compression to reduce the cost of retranslating previously seen code, and fragment pinning to avoid evicting needed code. We comprehensively evaluate our techniques to demonstrate their effectiveness.
José Baiocchi, Bruce R. Childers, Jack W. Davidson, Jason Hiser, Jonathan Misurda
CASES2
2007 Evaluating Indirect Branch Handling Mechanisms in Software Dynamic Translation Systems
abstract
Software dynamic translation (SDT) systems are used for program instrumentation, dynamic optimization, security, intrusion detection, and many other uses. As noted by many researchers, a major source of SDT overhead is the execution of code which is needed to translate an indirect branch's target address into the address of the translated destination block. This paper discusses the sources of indirect branch (IB) overhead in SDT systems and evaluates several techniques for overhead reduction. Measurements using SPEC CPU2000 show that the appropriate choice and configuration of IB translation mechanisms can significantly reduce the IB handling overhead. In addition, cross-architecture evaluation of IB handling mechanisms reveals that the most efficient implementation and configuration can be highly dependent on the implementation of the underlying architecture
Jason Hiser, Daniel W. Williams, Jack W. Davidson, Jason Mars, Bruce R. Childers
CGO6
2007 Exploring the interplay of yield, area, and performance in processor caches
abstract
The deployment of future deep submicron technology calls for a careful review of existing cache organizations and design practices in terms of yield and performance. This paper presents a cache design flow that enables processor architects to consider yield, area, and performance (YAP) together in a unified framework. Since there is a complex, changing trade-off between these metrics depending on the technology, the cache organization, and the yield enhancement scheme employed, such a design flow becomes invaluable to processor architects when they assess a design and explore the design space quickly at an early stage. We develop a complete set of tools supporting the proposed design flow, from injecting defects into a wafer to evaluating program performance of individual processors in the wafer. A case study is presented to demonstrate the effectiveness of the proposed design flow and developed tools.
Hyunjin Lee 0006, Sangyeun Cho, Bruce R. Childers
ICCD3
2007 Virtual Execution Environments: Support and Tools
abstract
In today's dynamic computing environments, the available resources and even underlying computation engine can change during the execution of a program. Additionally, current trends in software development favor the flexibility and cost-effectiveness of dynamically loaded components and libraries. Because of these trends, there has been increased research interest in virtual execution environments (VEEs) for delivering adaptable software suitable for today's rapidly changing, heterogeneous computing environments. In this project, we have been investigating tools and techniques to support implementation of VEEs using software dynamic translation (SDT). This paper highlights some of our recent results. One significant result is that we have developed novel translation techniques that reduce the memory and runtime overhead of SDT to negligible levels. We have also developed innovative debugging and instrumentation tools for SDT-based software environments. Together, these results make SDT-based systems viable for solving a wide range of pressing problems. The paper concludes with a discussion of how SDT may offer a solution to one such problem-inherent process variation in emerging chip multiprocessors.
Apala Guha, Jason Hiser, Naveen Kumar 0002, Jing Yang 0003, Min Zhao 0009, Shukang Zhou, Bruce R. Childers, Jack W. Davidson, Kim M. Hazelwood, Mary Lou Soffa
IPDPS7
2007 Integrated CPU and l2 cache voltage scaling using machine learning
abstract
Embedded systems serve an emerging and diverse set of applications. As a result, more computational and storage capabilities are added to accommodate ever more demanding applications. Unfortunately, adding more resources typically comes on the expense of higher energy costs. New chip design with Multiple Clock Domains (MCD) opens the opportunity for fine-grain power management within theprocessor chip. When used with dynamic voltage scaling (DVS), we can control the voltage and power of each domain independently. A significant power and energy improvement has been shown when using MCD design in comparison to managing a single voltage domain for the whole chip, as in traditional chips with global DVS.
Nevine AbouGhazaleh, Alexandre Peixoto Ferreira, Cosmin Rusu, Ruibin Xu, Frank Liberato, Bruce R. Childers, Daniel Mossé, Rami G. Melhem
LCTES6
2007 Near-Memory Caching for Improved Energy Consumption
abstract
Main memory has become one of the largest contributors to overall energy consumption and offers many opportunities for power/energy reduction. In this paper, we propose a Power-Aware Cached-DRAM (PA-CDRAM) organization that integrates a moderately sized cache directly into a memory chip. We use this near-memory cache to turn a memory bank off immediately after it is accessed to reduce power consumption.We modify the operation and structure of cached DRAM (CDRAM) with the goal of reducing energy consumption while retaining the performance advantage for which CDRAM was originally proposed. In this paper, we describe our PA-CDRAM organization and show how to incorporate it into Rambus memory. We evaluate the approach using a cycle accurate processor and memory simulator. Our results show that PA-CDRAM achieves up to 84% (28% on average) improvement in the energy-delay product and up to 76% (19% on average) savings in energy when compared to a time-out power management technique.
Nevine AbouGhazaleh, Bruce R. Childers, Daniel Mossé, Rami G. Melhem
IEEE Trans. Computers2
2006 Techniques and tools for dynamic optimization
abstract
Traditional code optimizers have produced significant performance improvements over the past forty years. While promising avenues of research still exist, traditional static and profiling techniques have reached the point of diminishing returns. The main problem is that these approaches have only a limited view of the program and have difficulty taking advantage of the actual run-time behavior of a program. We are addressing this problem through the development of a dynamic optimization system suited for aggressive optimization - using the full power of the most beneficial optimizations. We have designed our optimizer to operate using a software dynamic translation (SDT) execution system. Difficult challenges in this research include reducing SDT overhead and determining what optimizations to apply and where in the code to apply them. Another challenge is having the necessary tools to ensure the reliability of software that is dynamically optimized. In this paper, we describe our efforts in reducing overhead in SDT and efficient techniques for instrumenting the application code. We also describe our approach to determine what and where an optimization should be applied. We discuss other fundamental issues in developing a dynamic optimizer and finally present a basic debugger for SDT systems
Jason Hiser, Naveen Kumar 0002, Min Zhao 0009, Shukang Zhou, Bruce R. Childers, Jack W. Davidson, Mary Lou Soffa
IPDPS5
2006 Catching and Identifying Bugs in Register Allocation
Yuqiang Huang, Bruce R. Childers, Mary Lou Soffa
SAS2
2006 A Speculative Trace Reuse Architecture with Reduced Hardware Requirements
abstract
Trace reuse is an effective way of improving the performance of superscalar processors by skipping the execution of a sequence of instructions with known input and output values. However, the extra hardware complexity is of special concern when implementing such mechanisms. In this paper, we describe ways to reduce these requirements for Reuse through Speculation on Traces (RST). RST combines instruction and trace reuse with value prediction in an integrated mechanism to provide missing trace inputs when execution reaches the beginning of a trace. Speculatively reused traces do not consume resources in the execution pipeline, as they are not executed. In this paper, we study the effects of constraining reuse tables to effectively reduce the number of reuse candidates and comparisons. We compare our approach to instruction reuse, trace reuse and value prediction. We show that RST reuses more instructions and has better performance than traditional trace reuse, with an average speedup over a baseline without reuse of 1.21.
Maurício L. Pilla, Bruce R. Childers, Amarildo T. da Costa, Felipe M. G. França, Philippe Olivier Alexandre Navaux
SBAC-PAD2
2006 Evaluating fragment construction policies for SDT systems
abstract
Software Dynamic Translation (SDT) systems have been used for program instrumentation, dynamic optimization, security policy enforcement, intrusion detection, and many other uses. To be widely applicable, the overhead (runtime, memory usage, and power consumption) should be as low as possible. For instance, if an SDT system is protecting a web server against possible attacks, but causes 30% slowdown, a company may need 30% more machines to handle the web traffic they expect. Consequently, the causes of SDT overhead should be studied rigorously. This work evaluates many alternative policies for the creation of fragments within the Strata SDT framework. In particular, we examine the effects of ending translation at conditional branches; ending translation at unconditional branches; whether to use partial inlining for call instructions; whether to build the target of calls immediately or lazily; whether to align branch targets; and how to place code to transition back to the dynamic translator. We find that effective translation strategies are vital to program performance, improving performance from as much as 28% overhead, to as little as 3% overhead on average for the SPEC CPU2000 benchmark suite. We further demonstrate that these translation strategies are effective across several platforms, including Sun SPARC UltraSparc IIi, AMD Athlon Opteron, and Intel Pentium IV processors.
Jason Hiser, Daniel W. Williams, Adrian Filipi, Jack W. Davidson, Bruce R. Childers
VEE5
2006 An approach toward profit-driven optimization
abstract
Although optimizations have been applied for a number of years to improve the performance of software, problems with respect to the application of optimizations have not been adequately addressed. For example, in certain circumstances, optimizations may degrade performance. However, there is no efficient way to know when a degradation will occur. In this research, we investigate the profitability of optimizations, which is useful for determining the benefit of applying optimizations. We develop a framework that enables us to predict profitability using analytic models. The profitability of an optimization depends on code context, the particular optimization, and machine resources. Thus, our framework has analytic models for each of these components. As part of the framework, there is also a profitability engine that uses models to predict the profit. In this paper, we target scalar optimizations and, in particular, describe the models for partial redundancy elimination (PRE), loop invariant code motion (LICM), and value numbering (VN). We implemented the framework for predicting the profitability of these optimizations. Based on the predictions, we can selectively apply profitable optimizations. We compared the profit-driven approach with an approach that uses a heuristic in deciding when optimizations should be applied. Our experiments demonstrate that the profitability of scalar optimizations can be accurately predicted by using models. That is, without actually applying a scalar optimization, we can determine if an optimization is beneficial and should be applied.
Min Zhao 0009, Bruce R. Childers, Mary Lou Soffa
ACM Trans. Archit. Code Optim.2
2006 Collaborative operating system and compiler power management for real-time applications
abstract
Managing energy consumption has become vitally important to battery-operated portable and embedded systems. Dynamic voltage scaling (DVS) reduces the processor's dynamic power consumption quadratically at the expense of linearly decreasing the performance. When reducing energy with DVS for real-time systems, one must consider the performance penalty to ensure that deadlines can be met. In this paper, we introduce a novel collaborative approach between the compiler and the operating system (OS) to reduce energy consumption. We use the compiler to annotate an application's source code with path-dependent information called power-management hints (PMHs). This fine-grained information captures the temporal behavior of the application, which varies by executing different paths. During program execution, the OS periodically changes the processor's frequency and voltage based on the temporal information provided by the PMHs. These speed adaptation points are called power-management points (PMPs). We evaluate our scheme using three embedded applications: a video decoder, automatic target recognition, and a sub-band tuner. Our scheme shows an energy reduction of up to 57% over no power-management and up to 32% over a static power-management scheme. We compare our scheme to other schemes that solely utilize PMPs for power-management and show experimentally that our scheme achieves more energy savings. We also analyze the advantages and disadvantages of our approach relative to another compiler-directed scheme.
Nevine AbouGhazaleh, Daniel Mossé, Bruce R. Childers, Rami G. Melhem
ACM Trans. Embed. Comput. Syst.3
2005 Jazz: A Tool for Demand-Driven Structural Testing
Jonathan Misurda, James Clause, Juliya L. Reed, Bruce R. Childers, Mary Lou Soffa
CC4
2005 Model-Based Framework: An Approach for Profit-Driven Optimization
abstract
Although optimizations have been applied for a number of years to improve the performance of software, problems that have been long-standing remain, which include knowing what optimizations to apply and how to apply them. To systematically tackle these problems, we need to understand the properties of optimizations. In our current research, we are investigating the profitability property, which is useful for determining the benefit of applying an optimization. Due to the high cost of applying optimizations and then experimentally evaluating their profitability, we use an analytic model framework for predicting the profitability of optimizations. In this paper, we target scalar optimizations, and in particular, describe framework instances for partial redundancy elimination (PRE) and loop invariant code motion (LICM). We implemented the framework for both optimizations and compare profit-driven PRE and LICM with a heuristic-driven approach. Our experiments demonstrate that a model-based approach is effective and efficient in that it can accurately predict the profitability of optimizations with low overhead. By predicting the profitability using models, we can selectively apply optimizations. The model-based approach does not require tuning of parameters used in heuristic approaches and works well across different code contexts and optimizations.
Min Zhao 0009, Bruce R. Childers, Mary Lou Soffa
CGO2
2005 Near-memory Caching for Improved Energy Consumption
abstract
Main memory has become one of the largest contributors to overall energy consumption and offers many opportunities for power/energy reduction. In this paper, we propose a power-aware cached-DRAM (PA-CDRAM) organization that integrates a moderately sized cache directly into a memory module. We use this near-memory cache to turn a memory bank off immediately after it is accessed to reduce power consumption. We modify the structure of cached DRAM (CDRAM) with the goal of reducing energy consumption while retaining the performance advantage for which CDRAM was originally proposed. We evaluate the approach using a cycle accurate processor and memory simulator. Our results show that PACDRAM achieves up to 84% (28% on average) improvement in the energy-delay product and up to 76% (19% on average) savings in energy when compared to a time-out power management technique.
Nevine AbouGhazaleh, Bruce R. Childers, Daniel Mossé, Rami G. Melhem
ICCD2
2005 Demand-driven structural testing with dynamic instrumentation
abstract
Producing reliable and robust software has become one of the most important software development concerns in recent years. Testing is a process by which software quality can be assured through the collection of information. While testing can improve software reliability, current tools typically are inflexible and have high over-heads, making it challenging to test large software projects. In this paper, we describe a new scalable and flexible framework for testing programs with a novel demand-driven approach based on execution paths to implement test coverage. This technique uses dynamic instrumentation on the binary code that can be inserted and removed on-the-fly to keep performance and memory overheads low. We describe and evaluate implementations of the framework for branch, node and defuse testing of Java programs. Experimental results for branch testing show that our approach has, on average, a 1.6 speed up over static instrumentation and also uses less memory.
Jonathan Misurda, James Clause, Juliya L. Reed, Bruce R. Childers, Mary Lou Soffa
ICSE4
2005 Low overhead program monitoring and profiling
abstract
Program instrumentation, inserted either before or during execution, is rapidly becoming a necessary component of many systems. Instrumentation is commonly used to collect information for many diverse analysis applications, such as detecting program invariants, dynamic slicing and alias analysis, software security checking, and computer architecture modeling. Because instrumentation typically has a high run-time overhead, techniques are needed to mitigate the overheads. This paper describes "instrumentation optimizations" that reduce the overhead of profiling for program analysis. Our approach applies transformations to the instrumentation code that reduce the (1) number of instrumentation points executed, (2) cost of instrumentation probes, and (3) cost of instrumentation payload, while maintaining the semantics of the original instrumentation. We present the transformations and apply them for program profiling and computer architecture modeling. We evaluate the optimizations and show that the optimizations improve profiling performance by 1.26-2.63x and architecture modeling performance by 2-3.3x.
Naveen Kumar 0002, Bruce R. Childers, Mary Lou Soffa
PASTE2
2005 Planning for code buffer management in distributed virtual execution environments
abstract
Virtual execution environments have become increasingly useful in system implementation, with dynamic translation techniques being an important component for performance-critical systems. Many devices have exceptionally tight performance and memory constraints (e.g., smart cards and sensors in distributed systems), which require effective resource management. One approach to manage code memory is to download code partitions on-demand from a server and to cache the partitions in the resource-constrained device (client). However, due to the high cost of downloading code and re-translation, it is critical to intelligently manage the code buffer to minimize the overhead of code buffer misses. Yet, intelligent buffer management on the tightly constrained client can be too expensive. In this paper, we propose to move code buffer management to the server, where sophisticated schemes can be employed. We describe two schemes that use profiling information to direct the client in caching code partitions. One scheme is designed for workloads with stable run-time behavior, while the other scheme adapts its decisions for workloads with unstable behaviors. We evaluate and compare our schemes and show they perform well, compared to other approaches, with the adaptive scheme having the best performance overall.
Shukang Zhou, Bruce R. Childers, Mary Lou Soffa
VEE2
2004 Compact Binaries with Code Compression in a Software Dynamic Translator
abstract
Embedded software is becoming more flexible and adaptable, which presents new challenges for management of highly constrained system resources. Software dynamic translation (SDT) has been used to enable software malleability at the instruction level for dynamic code optimizers, security checkers, and binary translators. This paper studies the feasibility of using SDT to manage program code storage in embedded systems. We explore to what extent code compression can be incorporated in a software infrastructure to reduce program storage requirements, while minimally impacting run-time performance and memory resources. We describe two approaches for code compression, called full and partial image compression, and evaluate their compression ratios and performance in a software dynamic translation system. We demonstrate that code decompression is indeed feasible in a SDT.
Stacey Shogan, Bruce R. Childers
DATE2
2004 Profile Guided Management of Code Partitions for Embedded Systems
abstract
Researchers have proposed to divide embedded applications into code partitions and to download partitions on demand from a wireless code server to enable a diverse set of applications for very tightly constrained embedded systems. This paper describes a new approach for managing the request and storage of code partitions and we explore the benefits of our scheme.
Shukang Zhou, Bruce R. Childers, Naveen Kumar 0002
DATE2
2004 Value Predictors for Reuse through Speculation on Traces
abstract
Reusing dynamic sequences of instructions - i.e., traces - improves performance for many benchmarks. However, many traces are not reused because of unavailable inputs in the reuse test. Reuse through speculation on traces (RST) aims to increase the number of reused traces by predicting those inputs when necessary, with minimal additional hardware when compared to nonspeculative trace reuse. In this paper, we compare last n-value and stride-aware prediction for trace inputs. Last n-value prediction uses the last recorded values as predictions, while stride-aware prediction identifies and uses strides to compute new predictions. Stride-aware RST has a higher hardware cost than last n-value RST and has also the shortcoming of not allowing branches inside predicted traces. This paper aims to determine which scheme is the most beneficial for RST. We show that stride values are important for reuse in RST and that last n-value prediction works as well as the more sophisticated stride-aware approach with simpler hardware.
Maurício L. Pilla, Philippe Olivier Alexandre Navaux, Bruce R. Childers, Amarildo T. da Costa, Felipe M. G. França
SBAC-PAD3
2004 Custom Wide Counterflow Pipelines for High-Performance Embedded Applications
abstract
Application-specific instruction set processor (ASIP) design is a promising technique to meet the performance and cost goals of high-performance systems. ASIPs are especially valuable for embedded computing applications (e.g., digital cameras, color printers, cellular phones, etc.) where a small increase in performance and decrease in cost can have a large impact on a product's viability. Sutherland, Sproull, and Molnar originally proposed a processor organization called the counterflow pipeline (CFP) as a general-purpose architecture. We observed that the CFP is appropriate for ASIP design due to its simple and regular structure, local control and communication, and high degree of modularity. We describe a new CFP architecture, called the wide counterflow pipeline (WCFP), that extends the original proposal to be better suited for custom embedded instruction-level parallel processors. This presents a novel and practical application of the CFP to automatic and quick turnaround design of ASIPs. We introduce the WCFP architecture and describe several microarchitecture capabilities needed to get good performance from custom WCFPs. We demonstrate that custom WCFPs have performance that is up to four times better than that of ASIPs based on the CFP. Using an analytic cost model, we show that custom WCFPs do not unduly increase the cost of the original counterflow pipeline architecture, yet they retain the simplicity of the CFP. We also compare custom WCFPs to custom VLIW architectures and demonstrate that the WCFP is performance competitive with traditional VLIWs without requiring complicated global interconnection of functional devices.
Bruce R. Childers, Jack W. Davidson
IEEE Trans. Computers1
2003 Retargetable and Reconfigurable Software Dynamic Translation
abstract
Software dynamic translation (SDT) is a technology that permits the modification of an executing program's instructions. In recent years, SDT has received increased attention, from both industry and academia, as a feasible and effective approach to solving a variety of significant problems. Despite this increased attention, the task of initiating a new project in software dynamic translation remains a difficult one. To address this concern, and in particular, to promote the adoption of SDT technology into an even wider range of applications, we have implemented Strata, a cross-platform infrastructure for building software dynamic translators. This paper describes Strata's architecture, our experience retargeting it to three different processors, and our use of Strata to build two novel SDT systems - one for safe execution of untrusted binaries and one for fast prototyping of architectural simulators.
Kevin Scott, Naveen Kumar 0002, S. Velusamy, Bruce R. Childers, Jack W. Davidson, Mary Lou Soffa
CGO4
2003 Energy management for real-time embedded applications with compiler support
Nevine AbouGhazaleh, Bruce R. Childers, Daniel Mossé, Rami G. Melhem, Matthew Craven
LCTES2
2003 Predicting the impact of optimizations for embedded systems
abstract
When applying optimizations, a number of decisions are made using fixed strategies, such as always applying an optimization if it is applicable, applying optimizations in a fixed order and assuming a fixed configuration for optimizations such as tile size and loop unrolling factor. While it is widely recognized that these fixed strategies may not be the most appropriate for producing high quality code, especially for embedded systems, there are no general and automatic strategies that do otherwise. In this paper, we present a framework that enables these decisions to be made based on predicting the impact of an optimization, taking into account resources and code context. The framework consists of optimization models, code models and resource models, which are integrated for predicting the impact of applying optimizations. Because data cache performance is important to embedded codes, we focus on cache performance and present an instance of the framework for cache performance in this paper. Since most opportunities for cache improvement come from loop optimizations, we describe code, optimization and cache models tailored to predict the impact of applying loop optimizations for data locality. Experimentally we demonstrate the need to selectively apply optimizations and show the performance benefit of our framework in predicting when to apply an optimization. We also show that our framework can be used to choose the most beneficial optimization when a number of optimizations can be applied to a loop nest. And lastly, we show that we can use the framework to combine optimizations on a loop nest.
Min Zhao 0009, Bruce R. Childers, Mary Lou Soffa
LCTES2
2003 The Limits of Speculative Trace Reuse on Deeply Pipelined Processors
abstract
Trace reuse improves the performance of processors by skipping the execution of sequences of redundant instructions. However, many reusable traces do not have all of their inputs ready by the time the reuse test is done. For these cases, we developed a new technique called reuse through speculation on traces (RST), where trace inputs may be predicted. We study the limits of RST for modern processors with deep pipelines, as well as the effects of constraining resources on performance. We show that our approach reuses more traces than the nonspeculative trace reuse technique, with speedups of 43% over a nonspeculative trace reuse and 57% when memory accesses are reused.
Maurício L. Pilla, Amarildo T. da Costa, Felipe M. G. França, Bruce R. Childers, Mary Lou Soffa
SBAC-PAD4
2003 Scheduling with Dynamic Voltage/Speed Adjustment Using Slack Reclamation in Multiprocessor Real-Time Systems
abstract
The high power consumption of modern processors becomes a major concern because it leads to decreased mission duration (for battery-operated systems), increased heat dissipation, and decreased reliability. While many techniques have been proposed to reduce power consumption for uniprocessor systems, there has been considerably less work on multiprocessor systems. In this paper, based on the concept of slack sharing among processors, we propose two novel power-aware scheduling algorithms for task sets with and without precedence constraints executing on multiprocessor systems. These scheduling techniques reclaim the time unused by a task to reduce the execution speed of future tasks and, thus, reduce the total energy consumption of the system. We also study the effect of discrete voltage/speed levels on the energy savings for multiprocessor systems and propose a new scheme of slack reservation to incorporate voltage/speed adjustment overhead in the scheduling algorithms. Simulation and trace-based results indicate that our algorithms achieve substantial energy savings on systems with variable voltage processors. Moreover, processors with a few discrete voltage/speed levels obtain nearly the same energy savings as processors with continuous voltage/speed, and the effect of voltage/speed adjustment overhead on the energy savings is relatively small.
Dakai Zhu 0001, Rami G. Melhem, Bruce R. Childers
IEEE Trans. Parallel Distributed Syst.3
2001 Scheduling with Dynamic Voltage/Speed Adjustment Using Slack Reclamation in Multi-Processor Real-Time Systems
abstract
The power consumption of modern high-performance processors is becoming a major concern because it leads to increased heat dissipation and decreased reliability. While many techniques have been proposed to reduce power consumption for uni-processors, there has been considerably less work on multi-processor systems. In this paper we focus on power-aware scheduling for multi-processor real-time systems. Based on the idea of slack sharing among processors, we propose two novel scheduling algorithms for task sets with and without precedence constraints. These scheduling techniques reclaim the time unused by a task to reduce the execution speed of future tasks, and thus reduce the total energy consumption of the system. Simulation results indicate that our algorithms achieve up to 60% energy savings on multi-processor systems with variable voltage processors.
Dakai Zhu 0001, Rami G. Melhem, Bruce R. Childers
RTSS3
2001 Message from the Guest Editors
abstract
HIS special issue of the IEEE Transactions on Computers is the result of the Parallel Architecture and Compilation Techniques (PACT2000) Conference which was successfully held in Philadelphia, Pennsylvania, during the third week of October 2000. To follow the PACT tradition, this special issue is also intended to act as a forum that brings together research in parallel architectures and compilation. As advances in technology increase processing power and computing speed, the need for new architectures and compilers to harness such a potential also increases. This is exactly the goal of this special issue in providing a forum for computer architects and compiler designers to respond to the ever increasing need for higher performance. The major theme of the special issue was composed of 13 topics ranging from performance characterization and analytical modeling, prediction/speculation mechanisms, novel architectures, and software and hardware compilation techniques. We received 115 papers and five referees reviewed each paper, on average. The program committee met on 24 and 25 June 24 2000 in Pittsburgh. A great effort was given to select high quality of papers for presentation during the conference. After a long debate about each individual paper, the program committee recommended about 15 percent of the total papers for this special issue. Based on such a recommendation, authors of these papers were contacted and encouraged to revise and expand their papers for possible inclusion in this special issue. However, to be fair and to increase the quality of this special issue, a general call for papers was also issued. The 25 papers submitted for this special issue were evaluated and refereed by another set of reviewers. Late January of 2001, the guest editors spent a great deal of time making the final decision about the papers in this special issue. It was a challenging and rewarding experience for us to select the best among the best. This effort resulted in the selection of eight papers (six regular and two brief) in this special issue.
Ali R. Hurson, Bruce R. Childers
IEEE Trans. Computers2