Antonia Zhai

dblp:68/1545 · DBLP profile ↗
← Back
43ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0002-8921-1415ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 38 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 12 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2Computer networks · 1Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
16 papers
Processor architecture and microarchitecture · 28% Parallel and multicore computing · 23% Memory systems · 22%
Software engineering, system software, and programming languages
7 papers
Runtime systems and virtual machines · 58% Compilers and program optimization · 26% Program analysis · 16%

Topics — the 30 heaviest of 42, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Runtime systems and virtual machines › binary translation
dynamic binary translation
1.242019
Unleashing the Power of Learning: An Enhanced Learning-Based Approach for Dynamic Binary Translation · USENIX ATC 2019
Enhancing Cross-ISA DBT Through Automatically Learned Translation Rules · ASPLOS 2018
Enabling Cross-ISA Offloading for COTS Binaries · MobiSys 2017
Parallel and multicore computing › speculative parallelization
thread-level speculation
0.682013
The design and implementation of heterogeneous multicore systems for energy-efficient speculative thread execution · ACM Trans. Archit. Code Optim. 2013
Dynamically dispatching speculative threads to improve sequential execution · ACM Trans. Archit. Code Optim. 2012
Dynamic performance tuning for speculative threads · ISCA 2009
Processor architecture and microarchitecture › multicore design
heterogeneous multicore
0.422015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
The design and implementation of heterogeneous multicore systems for energy-efficient speculative thread execution · ACM Trans. Archit. Code Optim. 2013
Hardware accelerators and domain-specific architectures
spatial architecture
0.422015
Efficient Control and Communication Paradigms for Coarse-Grained Spatial Architectures · ACM Trans. Comput. Syst. 2015
Triggered instructions: a control paradigm for spatially-programmed architectures · ISCA 2013
Program analysis › symbolic execution
binary symbolic execution
0.312018
Enhancing Cross-ISA DBT Through Automatically Learned Translation Rules · ASPLOS 2018
Processor architecture and microarchitecture
multicore design
0.322013
The design and implementation of heterogeneous multicore systems for energy-efficient speculative thread execution · ACM Trans. Archit. Code Optim. 2013
Dynamically dispatching speculative threads to improve sequential execution · ACM Trans. Archit. Code Optim. 2012
Cloud and datacenter computing
computation offloading
0.312017
Enabling Cross-ISA Offloading for COTS Binaries · MobiSys 2017
Memory systems
cache management
0.212015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
Memory systems › cache management
cache partitioning
0.212015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
0.212015
Efficient Control and Communication Paradigms for Coarse-Grained Spatial Architectures · ACM Trans. Comput. Syst. 2015
Memory systems › memory consistency
memory consistency model
0.212015
Leveraging Transactional Execution for Memory Consistency Model Emulation · ACM Trans. Archit. Code Optim. 2015
Memory systems › cache
shared last-level cache
0.212015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
Processor architecture and microarchitecture
transactional execution
0.212015
Leveraging Transactional Execution for Memory Consistency Model Emulation · ACM Trans. Archit. Code Optim. 2015
Parallel and multicore computing
transactional memory
0.212015
Leveraging Transactional Execution for Memory Consistency Model Emulation · ACM Trans. Archit. Code Optim. 2015
Memory systems › memory hierarchy
cache hierarchy
0.212014
Measuring Microarchitectural Details of Multi- and Many-Core Memory Systems through Microbenchmarking · ACM Trans. Archit. Code Optim. 2014
Performance modeling and evaluation › benchmarking
microbenchmarking
0.212014
Measuring Microarchitectural Details of Multi- and Many-Core Memory Systems through Microbenchmarking · ACM Trans. Archit. Code Optim. 2014
Compilers and program optimization
instruction scheduling
0.122008
Compiler and hardware support for reducing the synchronization of speculative threads · ACM Trans. Archit. Code Optim. 2008
Compiler optimization of scalar value communication between speculative threads · ASPLOS 2002
Processor architecture and microarchitecture › instruction set architecture
cross-ISA translation
0.112018
Enhancing Cross-ISA DBT Through Automatically Learned Translation Rules · ASPLOS 2018
Processor architecture and microarchitecture
instruction set architecture
0.112018
Enhancing Cross-ISA DBT Through Automatically Learned Translation Rules · ASPLOS 2018
Energy-efficient computing › energy-aware mobile computing
energy-aware offloading
0.112017
Enabling Cross-ISA Offloading for COTS Binaries · MobiSys 2017
Compilers and program optimization
parallelizing compiler
0.112008
Compiler optimizations for parallelizing general-purpose applications under thread-level speculation · PPoPP 2008
Compilers and program optimization › parallelization
thread-level speculation
0.112008
Compiler optimizations for parallelizing general-purpose applications under thread-level speculation · PPoPP 2008
Parallel and multicore computing
speculative parallelization
0.112008
Compiler optimizations for parallelizing general-purpose applications under thread-level speculation · PPoPP 2008
Memory systems
cache coherence
0.122005
The STAMPede approach to thread-level speculation · ACM Trans. Comput. Syst. 2005
A scalable approach to thread-level speculation · ISCA 2000
Memory systems › cache coherence
writeback invalidation-based coherence
0.122005
The STAMPede approach to thread-level speculation · ACM Trans. Comput. Syst. 2005
A scalable approach to thread-level speculation · ISCA 2000
Energy-efficient computing › power management › memory power management
cache energy reduction
0.112015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
Electronic design automation › hardware verification and test › functional verification › emulation
full-system emulation
0.112015
Leveraging Transactional Execution for Memory Consistency Model Emulation · ACM Trans. Archit. Code Optim. 2015
Parallel and multicore computing
pipeline parallelism
0.112015
Efficient Control and Communication Paradigms for Coarse-Grained Spatial Architectures · ACM Trans. Comput. Syst. 2015
Processor architecture and microarchitecture › multithreading
simultaneous multithreading
0.112014
Measuring Microarchitectural Details of Multi- and Many-Core Memory Systems through Microbenchmarking · ACM Trans. Archit. Code Optim. 2014
Processor architecture and microarchitecture
multithreading
0.112005
The STAMPede approach to thread-level speculation · ACM Trans. Comput. Syst. 2005

Methods — techniques the papers use, named apart from their topics

dynamic binary translation · 0.8translation rule learning · 0.7symbolic execution · 0.7runtime offloading · 0.6machine learning · 0.4code caching · 0.2transactional execution · 0.2simulation · 0.2runtime thread-level parallelism measurement · 0.2memory fences · 0.2latency-insensitive channels · 0.2microbenchmarking · 0.2instruction scheduling · 0.1dataflow algorithms · 0.1thread-level speculation · 0.1compiler optimization · 0.1
YearPublicationVenuePosition
2025 DeCOS: Data-Efficient Reinforcement Learning for Compiler Optimization Selection Ignited by LLM
abstract
Machine learning methods have proven their effectiveness in a wide range of program optimization tasks.These methods selectively map program feature spaces to carefully defined optimization spaces to identify effective optimizations.However, the size and complexity of these spaces often necessitate large amounts of training data to achieve effective mappings.For certain optimization tasks, obtaining accurate training data can be costly, making data efficiency a critical concern.Reinforcement learning (RL) offers a promising solution by dynamically adjusting exploration strategies and selectively requesting training data.In this paper, we propose leveraging reinforcement learning to optimize compilation sequences.This paper presents the Data-efficient Compiler Optimization Selection (DeCOS) system, which utilizes a reinforcement learning engine to perform a guided search of the optimization spaces.To improve the data efficiency in training DeCOS, we utilize synthesized data to configure the RL-architecture; and incorporate simulation results to refine profiling information.To overcome the slow start-up issue in RL-processes, we integrate an LLM into the workflow, leveraging its knowledge to accelerate the initial training phase of the RL-agent.Our experiments show that DeCOS efficiently generates compiler optimization sequences that
Tianming Cui, Pen-Chung Yew, Stephen McCamant, Antonia Zhai
ICS4
2024 Non-Fusion Based Coherent Cache Randomization Using Cross-Domain Accesses
abstract
Randomization has proven to be a effective defense against conflict-based side-channel attacks in a shared cache. It improves security by assigning a unique randomization scheme to each security domain, e.g., though a different hashing function. However, if two domains have shared data, the domains must be fused in order to guarantee correctness (i.e., data coherence). Such domain fusion significantly reduces the effectiveness of randomization and weakens its security protection.
Kartik Ramkrishnan, Stephen McCamant, Antonia Zhai, Pen-Chung Yew
AsiaCCS3
2024 Interleaved Function Stream Execution Model for Cache-Aware High-Speed Stateful Packet Processing
abstract
The evolving network infrastructure, particularly the 5G core network, is increasingly adopting cloud technologies. This shift brings to the forefront the challenge of meeting the demanding per-packet processing requirements posed by multi-hundred Gbps Ethernet NICs (network interface cards). While traditional NFV (network function virtualization) platforms are effective on older hardware, the per-packet run-to-completion (RTC) execution model for per-packet processing suffers from stalling on state access due to L1/L2 cache misses. Although previous work applying software prefetching can mitigate the issues, their applications are fundamentally limited by the nature of a single execution stream, hence limiting them to batch lookups, suffering from control-flow divergence, and requiring manual tuning. To address the limitations, we introduce a novel interleaved function stream execution model that exploits the function-level parallelism through memory-level parallelism, targeting feature-rich network functions such as 5G Core. To provide the visibility into network functions, we introduce a novel programming model based on the principle of Granular Decomposition, which provides deep visibility into the state access by decoupling the state in a more fine-grained manner compared to traditional modular approaches. We integrate these two innovative designs into a new open-source NF platform, which we refer to as GuNFu. We have tested GuNFu on widely deployed network functions such as 5G UPF (User Plane Function), 5G AMF (Access Management Function), NAT (Network Address Translator) and others. Extensive evaluations reveal that GuNFu can achieve throughput ranging from 1.5 to 6 times over the traditional modular approach.
Ziyan Wu 0002, Yang Zhang 0074, Antonia Zhai, Zhi-Li Zhang
ICDCS5
2020 Efficient and scalable cross-ISA virtualization of hardware transactional memory
abstract
System virtualization is a key enabling technology. However, existing virtualization techniques suffer from a significant limitation due to their limited cross-ISA support for emerging architecture-specific hardware extensions. To address this issue, we make the first attempt at hardware transactional memory (HTM), which has been supported by modern multi-core processors and used by more and more applications to simplify concurrent programming. In particular, we propose an efficient and scalable mechanism to support cross-ISA virtualization of HTMs. The mechanism emulates guest HTMs using host HTMs, and tries to preserve as much as possible the performance and the scalability of guest applications. Experimental results on STAMP benchmarks show that an average of 2.3X and 12.6X performance speedup can be achieved respectively for x86_64 and PowerPC64 guest applications on an x86_64 host machine. Moreover, it can attain similar scalability to the native execution of the applications.
Wenwen Wang 0001, Pen-Chung Yew, Antonia Zhai, Stephen McCamant
CGO3
2020 First Time Miss : Low Overhead Mitigation for Shared Memory Cache Side Channels
abstract
Cache hit or miss is an important source of information leakage in cache side channel attacks. An attacker observes a much faster cache access time if the cache line has previously been filled in by the victim, and a much slower memory access time if the victim has not accessed this cache line, thus revealing to the attacker whether the victim has accessed the cache line or not.
Kartik Ramkrishnan, Stephen McCamant, Pen-Chung Yew, Antonia Zhai
ICPP4
2020 In-Network Memory Access Ordering for Heterogeneous Multicore Systems
abstract
In heterogeneous multicore systems, implementing a programmer-friendly memory consistency model while maximizing memory-level parallelism is challenging. Ideally, memory accesses can be performed out of order as long as program order is not violated. But enforcing memory access order at the end-point (e.g., a core) prohibits a number of architecture optimizations and limits memory-level parallelism. In this work, we explore the opportunity of preserving memory access order inside the on-chip interconnection network. We propose a hybrid switching networks-on-chip (NoC) attached with a light-weight token ring network to guarantee global memory access order. The hybrid switching NoC that supports both packet and circuit switching serves as the underlying communication infrastructure, while the token ring network is used to preserve memory order among multiple ordering points. Our proposed design enables strong memory consistency models and deterministic program execution, with negligible performance overhead compared to an un-ordered packet switching network.
Jieming Yin, Antonia Zhai
NOCS2
2019 Unleashing the Power of Learning: An Enhanced Learning-Based Approach for Dynamic Binary Translation
Changheng Song, Wenwen Wang 0001, Pen-Chung Yew, Antonia Zhai
USENIX ATC4
2018 Enhancing Cross-ISA DBT Through Automatically Learned Translation Rules
abstract
This paper presents a novel approach for dynamic binary translation (DBT) to automatically learn translation rules from guest and host binaries compiled from the same source code. The learned translation rules are then verified via binary symbolic execution and used in an existing DBT system, QEMU, to generate more efficient host binary code. Experimental results on SPEC CINT2006 show that the average time of learning a translation rule is less than two seconds. With the rules learned from a collection of benchmark programs excluding the targeted program itself, an average 1.25X performance speedup over QEMU can be achieved for SPEC CINT2006. Moreover, the translation overhead introduced by this rule-based approach is very small even for short-running workloads.
Wenwen Wang 0001, Stephen McCamant, Antonia Zhai, Pen-Chung Yew
ASPLOS3
2017 Enabling Cross-ISA Offloading for COTS Binaries
abstract
Work offloading allows a mobile device, i.e., the client, to execute its computation-intensive code remotely on a more powerful server to improve its performance and to extend its battery life. However, the difference in instruction set architectures (ISAs) between the client and the server poses a great challenge to work offloading. Most of the existing solutions rely on language-level virtual machines to hide such differences. Therefore, they have to tie closely to the specific programming languages. Other approaches try to recompile the mobile applications to achieve the specific goal of offloading, so their applicability is limited to the availability of the source code. To overcome the above limitations, we propose to extend the capability of dynamic binary translation across clients and servers to offload the identified computation-intensive binary code regions automatically to the server at runtime. With this approach, the native binaries on the client can be offloaded to the server seamlessly without the limitations mentioned above. A prototype has been implemented using an existing retargetable dynamic binary translator. Experimental results show that our system achieves 1.93X speedup with 48.66% reduction in energy consumption for six real-world applications, and 1.62X speedup with 42.4% reduction in energy consumption for SPEC CINT2006 benchmarks.
Wenwen Wang 0001, Pen-Chung Yew, Antonia Zhai, Stephen McCamant, Youfeng Wu, Jayaram Bobba
MobiSys3
2016 A General Persistent Code Caching Framework for Dynamic Binary Translation (DBT)
Wenwen Wang 0001, Pen-Chung Yew, Antonia Zhai, Stephen McCamant
USENIX ATC3
2015 Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor
abstract
Heterogeneous multicore processors that integrate CPU cores and data-parallel accelerators such as graphic processing unit (GPU) cores onto the same die raise several new issues for sharing various on-chip resources. The shared last-level cache (LLC) is one of the most important shared resources due to its impact on performance. Accesses to the shared LLC in heterogeneous multicore processors can be dominated by the GPU due to the significantly higher number of concurrent threads supported by the architecture. Under current cache management policies, the CPU applications’ share of the LLC can be significantly reduced in the presence of competing GPU applications. For many CPU applications, a reduced share of the LLC could lead to significant performance degradation. On the contrary, GPU applications can tolerate increase in memory access latency when there is sufficient thread-level parallelism (TLP). In addition to the performance challenge, introduction of diverse cores onto the same die changes the energy consumption profile and, in turn, affects the energy efficiency of the processor. In this work, we propose heterogeneous LLC management (HeLM), a novel shared LLC management policy that takes advantage of the GPU’s tolerance for memory access latency. HeLM is able to throttle GPU LLC accesses and yield LLC space to cache-sensitive CPU applications. This throttling is achieved by allowing GPU accesses to bypass the LLC when an increase in memory access latency can be tolerated. The latency tolerance of a GPU application is determined by the availability of TLP, which is measured at runtime as the average number of threads that are available for issuing. For a baseline configuration with two CPU cores and four GPU cores, modeled after existing heterogeneous processor designs, HeLM outperforms least recently used (LRU) policy by 10.4%. Additionally, HeLM also outperforms competing policies. Our evaluations show that HeLM is able to sustain performance with varying core mix. In addition to the performance benefit, bypassing also reduces total accesses to the LLC, leading to a reduction in the energy consumption of the LLC module. However, LLC bypassing has the potential to increase off-chip bandwidth utilization and DRAM energy consumption. Our experiments show that HeLM exhibits better energy efficiency by reducing the ED 2 by 18% over LRU while impacting only a 7% increase in off-chip bandwidth utilization.
Anup Holey, Vineeth Mekkat, Pen-Chung Yew, Antonia Zhai
ACM Trans. Archit. Code Optim.4
2015 Leveraging Transactional Execution for Memory Consistency Model Emulation
abstract
System emulation is widely used in today’s computer systems. This technology opens new opportunities for resource sharing as well as enhancing system security and reliability. System emulation across different instruction set architectures (ISA) can enable further opportunities. For example, cross-ISA emulation can enable workload consolidation over a wide range of microprocessors and potentially facilitate the seamless deployment of new processor architectures. As multicore and manycore processors become pervasive, it is important to address the challenges toward supporting system emulation on these platforms. A key challenge in cross-ISA emulation on multicore systems is ensuring the correctness of emulation when the guest and the host memory consistency models differ. Many existing cross-ISA system emulators are sequential, thus they are able to avoid this problem at the cost of significant performance degradation. Recently proposed parallel emulators are able to address the performance limitation; however, they provide limited support for memory consistency model emulation. When the host system has a weaker memory consistency model compared to the guest system, the emulator can insert memory fences at appropriate locations in the translated code to enforce the guest memory ordering constraints. These memory fences can significantly degrade the performance of the translated code. Transactional execution support available on certain recent microprocessors provides an alternative approach. Transactional execution of the translated code enforces sequential consistency (SC) at the coarse-grained transaction level, which in turn ensures that all memory accesses made on the host machine conform to SC. Enforcing SC on the host machine guarantees that the emulated execution will be correct for any guest memory model. In this article, we compare and evaluate the overheads associated with using transactions and fences for memory consistency model emulation on the Intel Haswell processor. Our experience of implementing these two approaches on a state-of-the-art parallel emulator, COREMU, demonstrates that memory consistency model emulation using transactions performs better when the transaction sizes are large enough to amortize the transaction overhead and the transaction conflict rate is low, whereas inserting memory fences is better for applications in which the transaction overhead is high. A hybrid implementation that dynamically determines which approach to invoke can outperform both approaches. Our results, based on the SPLASH-2 and the PARSEC benchmark suites, demonstrate that the proposed hybrid approach is able to outperform the fence insertion mechanism by 4.9% and the transactional execution approach by 24.9% for two-thread applications, and outperform them by 4.5% and 44.7%, respectively, for four-threaded execution.
Ragavendra Natarajan, Antonia Zhai
ACM Trans. Archit. Code Optim.2
2015 Efficient Control and Communication Paradigms for Coarse-Grained Spatial Architectures
abstract
There has been recent interest in exploring the acceleration of nonvectorizable workloads with spatially programmed architectures that are designed to efficiently exploit pipeline parallelism. Such an architecture faces two main problems: how to efficiently control each processing element (PE) in the system, and how to facilitate inter-PE communication without the overheads of traditional shared-memory coherent memory. In this article, we explore solving these problems using triggered instructions and latency-insensitive channels. Triggered instructions completely eliminate the program counter (PC) and allow programs to transition concisely between states without explicit branch instructions. Latency-insensitive channels allow efficient communication of inter-PE control information while simultaneously enabling flexible code placement and improving tolerance for variable events such as cache accesses. Together, these approaches provide a unified mechanism to avoid overserialized execution, essentially achieving the effect of techniques such as dynamic instruction reordering and multithreading. Our analysis shows that a spatial accelerator using triggered instructions and latency-insensitive channels can achieve 8 × greater area-normalized performance than a traditional general-purpose processor. Further analysis shows that triggered control reduces the number of static and dynamic instructions in the critical paths by 62% and 64%, respectively, over a PC-style baseline, increasing the performance of the spatial programming approach by 2.0 ×.
Michael Pellauer, Angshuman Parashar, Michael Adler, Bushra Ahsan, Randy L. Allmon, Neal Clayton Crago, Kermin Fleming, Mohit Gambhir, Aamer Jaleel, Tushar Krishna, Daniel Lustig, Stephen Maresh, Vladimir Pavlov, Rachid Rayess, Antonia Zhai, Joel S. Emer
ACM Trans. Comput. Syst.15
2014 Lightweight Software Transactions on GPUs
abstract
Graphics Processing Units (GPUs) provide an attractive option for extracting data-level parallelism from diverse applications. However, some applications, although possess abundant data-level parallelism, exhibit irregular memory access patterns to the shared data structures. Porting such applications to GPUs requires synchronization mechanisms such as locks, which significantly increase the programming complexity. Coarse-grained locking, where a single lock controls all the shared resources, although reduces programming efforts, can substantially serialize GPU threads. On the other hand, fine-grained locking, where each data element is protected by an independent lock, although facilitates maximum parallelism, requires significant programming efforts. To overcome these challenges, we propose to support software transactional memory (STM) on GPU that is able to achieve performance comparable to fine-grained locking, while requiring minimal programming efforts. Software-based transactional execution can incur significant runtime overheads due to activities such as detecting conflicts across thousands of GPU threads and managing a consistent memory state. Thus, in this paper we illustrate three lightweight STM designs that are capable of scaling to a large number of GPU threads. In our system, programmers simply mark the critical sections in the applications, and the underlying STM support is able to achieve performance comparable to fine-grained locking.
Anup Holey, Antonia Zhai
ICPP2
2014 Multi-stage coordinated prefetching for present-day processors
abstract
Data prefetching is an important technique for hiding memory latency. Latest microarchitectures provide support for both hardware and software prefetching. However, the architectural features supporting either are different. In addition, these features can vary from one architecture to another. As a result, the choice of the right prefetching strategy is non-trivial for both the programmers and compiler-writers.
Sanyam Mehta, Zhenman Fang, Antonia Zhai, Pen-Chung Yew
ICS3
2014 Energy-Efficient Time-Division Multiplexed Hybrid-Switched NoC for Heterogeneous Multicore Systems
abstract
NoCs are an integral part of modern multicore processors, they must continuously support high-throughput low-latency on-chip data communication under a stringent energy budget when system size scales up. Heterogeneous multicore systems further push the limit of NoC design by integrating cores with diverse performance requirements onto the same die. Traditional packet-switched NoCs, which have the flexibility of connecting diverse computation and storage devices, are facing great challenges to meet the performance requirements within the energy budget due to latency and energy consumption associated with buffering and routing at each router. In this paper, we take advantage of the diversity in performance requirements of on-chip heterogeneous computing devices by designing, implementing, and evaluating a hybrid-switched network that allows the packet-switched and circuit-switched messages to share the same communication fabric by partitioning the network through time-division multiplexing (TDM). In the proposed hybrid-switched network, circuit-switched paths are established along frequently communicating nodes. Our experiments show that utilizing these paths can improve system performance by reducing communication latency and alleviating network congestion. Furthermore, better energy efficiency is achieved by reducing buffering in routers and in turn enabling aggressive power gating.
Jieming Yin, Pingqiang Zhou, Sachin S. Sapatnekar, Antonia Zhai
IPDPS4
2014 Measuring Microarchitectural Details of Multi- and Many-Core Memory Systems through Microbenchmarking
abstract
As multicore and many-core architectures evolve, their memory systems are becoming increasingly more complex. To bridge the latency and bandwidth gap between the processor and memory, they often use a mix of multilevel private/shared caches that are either blocking or nonblocking and are connected by high-speed network-on-chip. Moreover, they also incorporate hardware and software prefetching and simultaneous multithreading (SMT) to hide memory latency. On such multi- and many-core systems, to incorporate various memory optimization schemes using compiler optimizations and performance tuning techniques, it is crucial to have microarchitectural details of the target memory system. Unfortunately, such details are often unavailable from vendors, especially for newly released processors. In this article, we propose a novel microbenchmarking methodology based on short elapsed-time events (SETEs) to obtain comprehensive memory microarchitectural details in multi- and many-core processors. This approach requires detailed analysis of potential interfering factors that could affect the intended behavior of such memory systems. We lay out effective guidelines to control and mitigate those interfering factors. Taking the impact of SMT into consideration, our proposed methodology not only can measure traditional cache/memory latency and off-chip bandwidth but also can uncover the details of software and hardware prefetching units not attempted in previous studies. Using the newly released Intel Xeon Phi many-core processor (with in-order cores) as an example, we show how we can use a set of microbenchmarks to determine various microarchitectural features of its memory system (many are undocumented from vendors). To demonstrate the portability and validate the correctness of such a methodology, we use the well-documented Intel Sandy Bridge multicore processor (with out-of-order cores) as another example, where most data are available and can be validated. Moreover, to illustrate the usefulness of the measured data, we do a multistage coordinated data prefetching case study on both Xeon Phi and Sandy Bridge and show that by using the measured data, we can achieve 1.3X and 1.08X performance speedup, respectively, compared to the state-of-the-art Intel ICC compiler. We believe that these measurements also provide useful insights into memory optimization, analysis, and modeling of such multicore and many-core architectures.
Zhenman Fang, Sanyam Mehta, Pen-Chung Yew, Antonia Zhai, James B. S. G. Greensky, Gautham Beeraka, Binyu Zang
ACM Trans. Archit. Code Optim.4
2013 Managing shared last-level cache in a heterogeneous multicore processor
abstract
Heterogeneous multicore processors that integrate CPU cores and data-parallel accelerators such as GPU cores onto the same die raise several new issues for sharing various on-chip resources. The shared last-level cache (LLC) is one of the most important shared resources due to its impact on performance. Accesses to the shared LLC in heterogeneous multicore processors can be dominated by the GPU due to the significantly higher number of threads supported. Under current cache management policies, the CPU applications' share of the LLC can be significantly reduced in the presence of competing GPU applications. For cache sensitive CPU applications, a reduced share of the LLC could lead to significant performance degradation. On the contrary, GPU applications can often tolerate increased memory access latency in the presence of LLC misses when there is sufficient thread-level parallelism. In this work, we propose Heterogeneous LLC Management (HeLM), a novel shared LLC management policy that takes advantage of the GPU's tolerance for memory access latency. HeLM is able to throttle GPU LLC accesses and yield LLC space to cache sensitive CPU applications. GPU LLC access throttling is achieved by allowing GPU threads that can tolerate longer memory access latencies to bypass the LLC. The latency tolerance of a GPU application is determined by the availability of thread-level parallelism, which can be measured at runtime as the average number of threads that are available for issuing. Our heterogeneous LLC management scheme outperforms LRU policy by 12.5% and TAP-RRIP by 5.6% for a processor with 4 CPU and 4 GPU cores.
Vineeth Mekkat, Anup Holey, Pen-Chung Yew, Antonia Zhai
PACT4
2013 HAccRG: Hardware-Accelerated Data Race Detection in GPUs
abstract
Modern Graphics Processing Units (GPUs) are capable of supporting thousands of concurrent threads. However, they provide relatively little guarantee with respect to the coherence and consistency of the memory system. Thus, GPUs are prone to multitude of concurrency bugs related to inconsistent memory states. Many such bugs manifest as some form of data races at runtime, and being able to identify these data races can help programmers improve software reliability. Mechanisms that enable efficient and effective data race detection at runtime can form the basis of powerful tools for enhancing GPU software correctness. Most prior works in data race detection for GPU focus on the software-based approaches that incur significant performance overhead. Furthermore, they often focus on the smaller shared memory, while neglecting the larger global memory. We believe that adequate hardware support can enable efficient data race detection in all levels of the memory system for GPUs. In this paper, we propose a hardware-accelerated data race detection mechanism, HAccRG, for efficient data race detection in GPUs. HAccRG provides hardware support for tracking data dependencies across a large number of threads and detects various forms of data races. We incorporate HAccRG on both the shared and global memory spaces in GPU. Our evaluation shows that, with moderate hardware support, HAccRG can detect data races in GPU kernels with a small overhead: 1% for the shared memory and 27% for combined shared and global memory data race detection.
Anup Holey, Vineeth Mekkat, Antonia Zhai
ICPP3
2013 Triggered instructions: a control paradigm for spatially-programmed architectures
abstract
In this paper, we present triggered instructions, a novel control paradigm for arrays of processing elements (PEs) aimed at exploiting spatial parallelism. Triggered instructions completely eliminate the program counter and allow programs to transition concisely between states without explicit branch instructions. They also allow efficient reactivity to inter-PE communication traffic. The approach provides a unified mechanism to avoid over-serialized execution, essentially achieving the effect of techniques such as dynamic instruction reordering and multithreading, which each require distinct hardware mechanisms in a traditional sequential architecture.
Angshuman Parashar, Michael Pellauer, Michael Adler, Bushra Ahsan, Neal Clayton Crago, Daniel Lustig, Vladimir Pavlov, Antonia Zhai, Mohit Gambhir, Aamer Jaleel, Randy L. Allmon, Rachid Rayess, Stephen Maresh, Joel S. Emer
ISCA8
2013 Accelerating Data Race Detection Utilizing On-Chip Data-Parallel Cores
Vineeth Mekkat, Anup Holey, Antonia Zhai
RV3
2013 The design and implementation of heterogeneous multicore systems for energy-efficient speculative thread execution
abstract
With the emergence of multicore processors, various aggressive execution models have been proposed to exploit fine-grained thread-level parallelism, taking advantage of the fast on-chip interconnection communication. However, the aggressive nature of these execution models often leads to excessive energy consumption incommensurate to execution time reduction. In the context of Thread-Level Speculation, we demonstrated that on a same-ISA heterogeneous multicore system, by dynamically deciding how on-chip resources are utilized, speculative threads can achieve performance gain in an energy-efficient way. Through a systematic design space exploration, we built a multicore architecture that integrates heterogeneous components of processing cores and first-level caches. To cope with processor reconfiguration overheads, we introduced runtime mechanisms to mitigate their impacts. To match program execution with the most energy-efficient processor configuration, the system was equipped with a dynamic resource allocation scheme that characterizes program behaviors using novel processor counters. We evaluated the proposed heterogeneous system with a diverse set of benchmark programs from SPEC CPU2000 and CPU20006 suites. Compared to the most efficient homogeneous TLS implementation, we achieved similar performance but consumed 18% less energy. Compared to the most efficient homogeneous uniprocessor running sequential programs, we improved performance by 29% and reduced energy consumption by 3.6%, which is a 42% improvement in energy-delay-squared product.
Yangchun Luo, Wei-Chung Hsu, Antonia Zhai
ACM Trans. Archit. Code Optim.3
2012 Energy-efficient non-minimal path on-chip interconnection network for heterogeneous systems
abstract
Network-on-Chips (NoCs) in heterogeneous systems containing both CPU and GPU cores must be designed to satisfy the performance requirements of both latency-sensitive CPU traffic and throughput-intensive GPU traffic. DVFS and adaptive routing can potentially improve NoC energy and performance efficiency. We further notice that GPU traffic can sometimes tolerate a slack defined as the number of cycles a packet can be delayed without causing performance penalty. In this work, we take advantage of the slack in GPU packets to route packets through non-minimal path, so that routers can operate at a lower frequency without suffering performance penalty.
Jieming Yin, Pingqiang Zhou, Anup Holey, Sachin S. Sapatnekar, Antonia Zhai
ISLPED5
2012 Dynamically dispatching speculative threads to improve sequential execution
abstract
Efficiently utilizing multicore processors to improve their performance potentials demands extracting thread-level parallelism from the applications. Various novel and sophisticated execution models have been proposed to extract thread-level parallelism from sequential programs. One such execution model, Thread-Level Speculation (TLS), allows potentially dependent threads to execute speculatively in parallel. However, TLS execution is inherently unpredictable, and consequently incorrect speculation could degrade performance for the multicore systems. Existing approaches have focused on using the compilers to select sequential program regions to apply TLS. Our research shows that even the state-of-the-art compiler makes suboptimal decisions, due to the unpredictability of TLS execution. Thus, we propose to dynamically optimize TLS performance. This article describes the design, implementation, and evaluation of a runtime thread dispatching mechanism that adjusts the behaviors of speculative threads based on their efficiency. In the proposed system, speculative threads are monitored by hardware-based performance counters and their performance impact is evaluated with a novel methodology that takes into account various unique TLS characteristics. Thread dispatching policies are devised to adjust the behaviors of speculative threads accordingly. With the help of the runtime evaluation, where and how to create speculative threads is better determined. Evaluated with all the SPEC CPU2000 benchmark programs written in C, the dynamic dispatching system outperforms the state-of-the-art compiler-based thread management techniques by 9.4% on average. Comparing to sequential execution, we achieve 1.37X performance improvement on a four-core CMP-based system.
Yangchun Luo, Antonia Zhai
ACM Trans. Archit. Code Optim.2
2011 Enabling improved power management in multicore processors through clustered DVFS
abstract
In recent years, chip multiprocessors (CMP) have emerged as a solution for high-speed computing demands. However, power dissipation in CMPs can be high if numerous cores are simultaneously active. Dynamic voltage and frequency scaling (DVFS) is widely used to reduce the active power, but its effectiveness and cost depends on the granularity at which it is applied. Per-core DVFS allows the greatest flexibility in controlling power, but incurs the expense of an unrealistically large number of on-chip voltage regulators. Per-chip DVFS, where all cores are controlled by a single regulator overcomes this problem at the expense of greatly reduced flexibility. This work considers the problem of building an intermediate solution, clustering the cores of a multicore processor into DVFS domains and implementing DVFS on a per-cluster basis. Based on a typical workload, we propose a scheme to find similarity among the cores and cluster them based on this similarity. We also provide an algorithm to implement DVFS for the clusters, and evaluate the effectiveness of per-cluster DVFS in power reduction.
T. Kolpe, Antonia Zhai, Sachin S. Sapatnekar
DATE2
2011 NoC frequency scaling with flexible-pipeline routers
Pingqiang Zhou, Jieming Yin, Antonia Zhai, Sachin S. Sapatnekar
ISLPED3
2011 Efficient dynamic program monitoring on multi-core systems
Guojin He, Antonia Zhai
J. Syst. Archit.2
2010 Energy efficient speculative threads: dynamic thread allocation in Same-ISA heterogeneous multicore systems
abstract
Thread-level parallelism at the chip level is critical in overcoming some of the challenges that have been ushered in through the advent of modern multicore processors (CMP). Extracting speculatively parallel threads from sequential applications and executing these threads on multicore processors is a promising technique to speed up these applications on multicore systems. However, the potential degradation in energy efficiency associated is an important factor that hinders the deployment of this technique. For multicore systems that integrate same-ISA heterogeneous cores, it is possible to judiciously allocate speculative threads to achieve energy-efficient performance improvement.
Yangchun Luo, Venkatesan Packirisamy, Wei-Chung Hsu, Antonia Zhai
PACT4
2010 Improving the performance of program monitors with compiler support in multi-core environment
abstract
Dynamic program execution monitors allow programmers to observe and verify an application while it is running. Instrumentation-based dynamic program monitors often incur significant performance overhead due to instrumentation. Special hardware supports have been proposed to reduce this overhead. However, these supports mostly target specific monitoring requirements and thus have limited applicability. Recently, with multi-core processors becoming mainstream, executing the monitored program and the monitor simultaneously on separate cores has emerged as an attractive option. However, communication between the two often becomes the new performance bottleneck due to large amounts of information forwarded to the monitor. In this paper, we present compiler techniques that aim to minimize the communication overhead. Our proposal is based on the observations that a monitor only requires specific information from the monitored programs and some information can be easily computed by the monitor from data that have already been communicated. We developed a code generator and optimization techniques to decide the set of data items to forward and the set to compute, so that the total execution time of the monitor is minimized. Our compiler can optimize a variety of monitors with diverse monitoring requirements, taking as input the control flow graph of the monitored program and the set of data that needs verification. Using a static binary rewriter, we evaluate the performance impact of the proposed compiler techniques on the SPEC2006 integer benchmarks for two intensive monitoring tasks: taint-propagation and memory bug detection. Comparing to instrumentation-based monitors, the proposed techniques can bring down the performance overhead of the two monitors from 10.6× and 9.0× to 2.36× and 2.17×, respectively.
Guojin He, Antonia Zhai
IPDPS2
2009 Exploiting TLS Parallelism at Multiple Loop-Nest Levels
abstract
As the number of cores integrated onto a single chip increases, architecture and compiler designers are challenged with the difficulty of utilizing these cores to improve the performance of a single application. Thread-level speculation (TLS) can potentially help by allowing possibly dependent threads to speculatively execute in parallel. Extracting speculative thread from sequential applications is key to efficient TLS execution. Previous work on thread extraction has focused on parallelizing iterations from a single loop-nest level or function continuation. However, the amount of parallelism available at a single loop-nest level is sometimes limited, and we are forced to look for parallelism across multiple loop-nest levels. In this paper we propose SpecOPTAL - a compiler algorithm that statically allocates cores to threads extracted from different levels of loop-nests. We show that, a subset of SPEC 2006 benchmarks are able to benefit from the proposed technique.
Venkatesan Packirisamy, Antonia Zhai
ICPADS2
2009 Dynamic performance tuning for speculative threads
abstract
In response to the emergence of multicore processors, various novel and sophisticated execution models have been introduced to fully utilize these processors. One such execution model is Thread-Level Speculation (TLS), which allows potentially dependent threads to execute speculatively in parallel. While TLS offers significant performance potential for applications that are otherwise non-parallel, extracting efficient speculative threads in the presence of complex control flow and ambiguous data dependences is a real challenge. This task is further complicated by the fact that the performance of speculative threads is often architecture-dependent, input-sensitive, and exhibits phase behaviors. Thus we propose dynamic performance tuning mechanisms that determine where and how to create speculative threads at runtime.
Yangchun Luo, Venkatesan Packirisamy, Wei-Chung Hsu, Antonia Zhai, Nikhil Mungre, Ankit Tarkas
ISCA4
2009 Exploring speculative parallelism in SPEC2006
abstract
The computer industry has adopted multi-threaded and multi-core architectures as the clock rate increase stalled in early 2000's. It was hoped that the continuous improvement of single-program performance could be achieved through these architectures. However, traditional parallelizing compilers often fail to effectively parallelize general-purpose applications which typically have complex control flow and excessive pointer usage. Recently hardware techniques such as Transactional Memory (TM) and Thread-Level Speculation (TLS) have been proposed to simplify the task of parallelization by using speculative threads. Potential of speculative parallelism in general-purpose applications like SPEC CPU 2000 have been well studied and shown to be moderately successful. Preliminary work examining the potential parallelism in SPEC2006 deployed parallel threads with a restrictive TLS execution model and limited compiler support, and thus only showed limited performance potential. In this paper, we first analyze the cross-iteration dependence behavior of SPEC 2006 benchmarks and show that more parallelism potential is available in SPEC 2006 benchmarks, comparing to SPEC2000. We further use a state-of-the-art profile-driven TLS compiler to identify loops that can be speculatively parallelized. Overall, we found that with optimal loop selection we can potentially achieve an average speedup of 60% on four cores over what could be achieved by a traditional parallelizing compiler such as Intel's ICC compiler.We also found that an additional 11% improvement can be potentially obtained on selected benchmarks using 8 cores when we extend TLS on multiple loop levels as opposed to restricting to a single loop level.
Venkatesan Packirisamy, Antonia Zhai, Wei-Chung Hsu, Pen-Chung Yew, Tin-Fook Ngai
ISPASS2
2009 Hardware Supported Flexible Monitoring: Early Results
Antonia Zhai, Guojin He, Mats P. E. Heimdahl
RV1
2008 Efficiency of thread-level speculation in SMT and CMP architectures - performance, power and thermal perspective
abstract
Computer industry has adopted multi-threaded and multi-core architectures as the clock rate increase stalled in early 2000psilas. However, because of the lack of compilers and other related software technologies, most of the general-purpose applications today still cannot take advantage of such architectures to improve their performance. Thread-level speculation (TLS) has been proposed as a way of using these multi-threaded architectures to parallelize general-purpose applications. Both simultaneous multithreading (SMT) and chip multiprocessors (CMP) have been extended to implement TLS. While the characteristics of SMT and CMP have been widely studied under multi-programmed and parallel workloads, their behavior under TLS workload is not well understood. The TLS workload due to speculative nature of the threads which could potentially be rollbacked and due to variable degree of parallelism available in applications, exhibits unique characteristics which makes it different from other workloads. In this paper, we present a detailed study of the performance, power consumption and thermal effect of these multithreaded architectures against that of a Superscalar with equal chip area. A wide spectrum of design choices and tradeoffs are also studied using commonly used simulation techniques. We show that the SMT based TLS architecture performs about 21% better than the best CMP based configuration while it suffers about 16% power overhead. In terms of Energy-Delay-Squared product (ED2), SMT based TLS performs about 26% better than the best CMP based TLS configuration and 11% better than the superscalar architecture. But the SMT based TLS configuration, causes more thermal stress than the CMP based TLS architectures.
Venkatesan Packirisamy, Yangchun Luo, Wei-Lung Hung, Antonia Zhai, Pen-Chung Yew, Tin-Fook Ngai
ICCD4
2008 Compiler optimizations for parallelizing general-purpose applications under thread-level speculation
abstract
No abstract available.
Antonia Zhai, Shengyue Wang, Pen-Chung Yew, Guojin He
PPoPP1
2008 Compiler and hardware support for reducing the synchronization of speculative threads
abstract
Thread-level speculation (TLS) allows us to automatically parallelize general-purpose programs by supporting parallel execution of threads that might not actually be independent. In this article, we focus on one important limitation of program performance under TLS, which stalls as a result of synchronizing and forwarding scalar values between speculative threads that would otherwise cause frequent data dependences and, hence, failed speculation. Using SPECint benchmarks that have been automatically transformed by our compiler to exploit TLS, we present, evaluate in detail, and compare both compiler and hardware techniques for improving the communication of scalar values. We find that through our dataflow algorithms for three increasingly aggressive instruction scheduling techniques, the compiler can drastically reduce thecritical forwarding pathintroduced by the synchronization and forwarding of scalar values. We also show that hardware techniques for reducing synchronization can be complementary to compiler scheduling, but that the additional performance benefits are minimal and are generally not worth the cost.
Antonia Zhai, J. Gregory Steffan, Christopher B. Colohan, Todd C. Mowry
ACM Trans. Archit. Code Optim.1
2006 Supporting Speculative Multithreading on Simultaneous Multithreaded Processors
Venkatesan Packirisamy, Shengyue Wang, Antonia Zhai, Wei-Chung Hsu, Pen-Chung Yew
HiPC3
2005 A General Compiler Framework for Speculative Optimizations Using Data Speculative Code Motion
abstract
Data speculative optimization refers to code transformations that allow load and store instructions to be moved across potentially dependent memory operations. Existing research work on data speculative optimizations has mainly focused on individual code transformation. The required speculative analysis that identifies data speculative optimization opportunities and the required recovery code generation that guarantees the correctness of their execution are handled separately for each optimization. This paper proposes a new compiler framework to facilitate the design and implementation of general data speculative optimizations such as dead store elimination, redundancy elimination, copy propagation, and code scheduling. This framework allows different data speculative optimizations to share the followings: (i) a speculative analysis mechanism to identify data speculative optimization opportunities by ignoring low probability data dependences from optimizations, and (ii) a recovery code generation mechanism to guarantee the correctness of the data speculative optimizations. The proposed recovery code generation is based on data speculative code motion (DSCM) that uses code motion to facilitate a desired transformation. Based on the position of the moved instruction, recovery code can be generated accordingly. The proposed framework greatly simplifies the task of incorporating data speculation into non-speculative optimizations by sharing the recovery code generation and the speculative analysis. We have implemented the proposed framework in the ORC 2.1 compiler and demonstrated its effectiveness on SPEC2000 benchmark programs.
Xiaoru Dai, Antonia Zhai, Wei-Chung Hsu, Pen-Chung Yew
CGO2
2005 The STAMPede approach to thread-level speculation
abstract
Multithreaded processor architectures are becoming increasingly commonplace: many current and upcoming designs support chip multiprocessing, simultaneous multithreading, or both. While it is relatively straightforward to use these architectures to improve the throughput of a multithreaded or multiprogrammed workload, the real challenge is how to easily create parallel software to allow single programs to effectively exploit all of this raw performance potential. One promising technique for overcoming this problem is Thread-Level Speculation (TLS) , which enables the compiler to optimistically create parallel threads despite uncertainty as to whether those threads are actually independent. In this article, we propose and evaluate a design for supporting TLS that seamlessly scales both within a chip and beyond because it is a straightforward extension of write-back invalidation-based cache coherence (which itself scales both up and down). Our experimental results demonstrate that our scheme performs well on single-chip multiprocessors where the first level caches are either private or shared. For our private-cache design, the program performance of two of 13 general purpose applications studied improves by 86% and 56%, four others by more than 8%, and an average across all applications of 16%---confirming that TLS is a promising way to exploit the naturally-multithreaded processing resources of future computer systems.
J. Gregory Steffan, Christopher B. Colohan, Antonia Zhai, Todd C. Mowry
ACM Trans. Comput. Syst.3
2004 Compiler Optimization of Memory-Resident Value Communication Between Speculative Threads
abstract
Efficient inter-thread value communication is essential for improving performance in thread-level speculation (TLS). Although several mechanisms for improving value communication using hardware support have been proposed, there is relatively little work on exploiting the potential of compiler optimization. Building on recent research on compiler optimization of scalar value communication between speculative threads, we propose compiler techniques for the optimization of memory-resident values. In TLS, data dependences through memory-resident values are tracked by the underlying hardware and preserved by re-executing any speculative thread that violates a dependence; however, re-execution incurs a large performance penalty and should be used only to resolve data dependences that are infrequent. In contrast, value communication for frequently-occurring data dependences must be very efficient. We propose using the compiler to first identify frequently-occurring memory-resident data dependences, then insert synchronization for communicating values to preserve these dependences. We find that by synchronizing frequently-occurring data dependences we can significantly improve the efficiency of parallel execution. A comparison between compiler-inserted and hardware-inserted memory synchronization reveals that the two techniques are complementary, with each technique benefitting different benchmarks.
Antonia Zhai, Christopher B. Colohan, J. Gregory Steffan, Todd C. Mowry
CGO1
2002 Compiler optimization of scalar value communication between speculative threads
abstract
While there have been many recent proposals for hardware that supports Thread-Level Speculation (TLS), there has been relatively little work on compiler optimizations to fully exploit this potential for parallelizing programs optimistically. In this paper, we focus on one important limitation of program performance under TLS, which is stalls due to forwarding scalar values between threads that would otherwise cause frequent data dependences. We present and evaluate dataflow algorithms for three increasingly-aggressive instruction scheduling techniques that reduce the critical forwarding path introduced by the synchronization associated with this data forwarding. In addition, we contrast our compiler techniques with related hardware-only approaches. With our most aggressive compiler and hardware techniques, we improve performance under TLS by 6.2-28.5% for 6 of 14 applications, and by at least 2.7% for half of the other applications.
Antonia Zhai, Christopher B. Colohan, J. Gregory Steffan, Todd C. Mowry
ASPLOS1
2002 Improving Value Communication for Thread-Level Speculation
abstract
Thread-level speculation (TLS) allows us to automatically parallelize general-purpose programs by supporting parallel execution of threads that might not actually be independent. In this paper, we show that the key to good performance ties in the three different ways to communicate a value between speculative threads: speculation, synchronization and prediction. The difficult part is deciding how and when to apply each method. This paper shows how we can apply value prediction, dynamic synchronization and hardware instruction prioritization to improve value communication and hence performance in several SPECint benchmarks that have been automatically transformed by our compiler to exploit TLS. We find that value prediction can be effective when properly throttled to avoid the high costs of mis-prediction, while most of the gains of value prediction can be more easily achieved by exploiting silent stores. We also show that dynamic synchronization is quite effective for most benchmarks, while hardware instruction prioritization is not. Overall, we find that these techniques have great potential for improving the performance of TLS.
J. Gregory Steffan, Christopher B. Colohan, Antonia Zhai, Todd C. Mowry
HPCA3
2000 A scalable approach to thread-level speculation
abstract
While architects understand how to build cost-effective parallel machines across a wide spectrum of machine sizes (ranging from within a single chip to large-scale servers), the real challenge is how to easily create parallel software to effectively exploit all of this raw performance potential. One promising technique for overcoming this problem is Thread-Level Speculation (TLS), which enables the compiler to optimistically create parallel threads despite uncertainty as to whether those threads are actually independent. In this paper, we propose and evaluate a design for supporting TLS that seamlessly scales to any machine size because it is a straightforward extension of writeback invalidation-based cache coherence (which itself scales both up and down). Our experimental results demonstrate that our scheme performs well on both single-chip multiprocessors and on larger-scale machines where communication latencies are twenty times larger.
J. Gregory Steffan, Christopher B. Colohan, Antonia Zhai, Todd C. Mowry
ISCA3