Yuan Chou

dblp:13/5626 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 58% Processor architecture and microarchitecture · 28% Performance modeling and evaluation · 15%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory access optimization
memory-level parallelism
0.122005
Store Memory-Level Parallelism Optimizations for Commercial Applications · MICRO 2005
Microarchitecture Optimizations for Exploiting Memory-Level Parallelism · ISCA 2004
Memory systems
cache
0.112007
Low-Cost Epoch-Based Correlation Prefetching for Commercial Applications · MICRO 2007
Memory systems › cache › prefetching
correlation prefetching
0.112007
Low-Cost Epoch-Based Correlation Prefetching for Commercial Applications · MICRO 2007
Memory systems › cache
prefetching
0.112007
Low-Cost Epoch-Based Correlation Prefetching for Commercial Applications · MICRO 2007
Processor architecture and microarchitecture
chip multiprocessor
0.112005
Effective Instruction Prefetching in Chip Multiprocessors for Modern Commercial Applications · HPCA 2005
Memory systems › cache › CPU cache
instruction cache
0.112005
Effective Instruction Prefetching in Chip Multiprocessors for Modern Commercial Applications · HPCA 2005
Processor architecture and microarchitecture › instruction fetch
instruction prefetching
0.112005
Effective Instruction Prefetching in Chip Multiprocessors for Modern Commercial Applications · HPCA 2005
Memory systems › memory consistency
memory consistency model
0.112005
Store Memory-Level Parallelism Optimizations for Commercial Applications · MICRO 2005
Performance modeling and evaluation › workload characterization
commercial workloads
0.132007
Low-Cost Epoch-Based Correlation Prefetching for Commercial Applications · MICRO 2007
Store Memory-Level Parallelism Optimizations for Commercial Applications · MICRO 2005
Effective Instruction Prefetching in Chip Multiprocessors for Modern Commercial Applications · HPCA 2005
Performance modeling and evaluation
workload characterization
0.132007
Low-Cost Epoch-Based Correlation Prefetching for Commercial Applications · MICRO 2007
Store Memory-Level Parallelism Optimizations for Commercial Applications · MICRO 2005
Effective Instruction Prefetching in Chip Multiprocessors for Modern Commercial Applications · HPCA 2005
Processor architecture and microarchitecture
out-of-order execution
0.012004
Microarchitecture Optimizations for Exploiting Memory-Level Parallelism · ISCA 2004
Processor architecture and microarchitecture › latency hiding
runahead execution
0.012004
Microarchitecture Optimizations for Exploiting Memory-Level Parallelism · ISCA 2004

Methods — techniques the papers use, named apart from their topics

epoch-based hiding · 0.1correlation table in main memory · 0.1store prefetching · 0.1store miss accelerator · 0.1speculative lock elision · 0.1sequential prefetching · 0.1discontinuity prefetching · 0.1simulation · 0.0
YearPublicationVenuePosition
2007 Low-Cost Epoch-Based Correlation Prefetching for Commercial Applications
abstract
The performance of many important commercial workloads, such as on-line transaction processing, is limited by the frequent stalls due to off-chip instruction and data accesses. These applications are characterized by irregular control flow and complex data access patterns that render many low-cost prefetching schemes, such as stream-based and stride-based prefetching, ineffective. For such applications, correlation-based prefetching, which is capable of capturing complex data access patterns, has been shown to be a more promising approach. However, the large instruction and data working sets of these applications require extremely large correlation tables, making these tables impractical to be implemented on-chip. This paper proposes the epoch-based correlation prefetcher, which cost-effectively stores its correlation table in main memory and exploits the concept of epochs to hide the long latency of its correlation table access, and which attempts to eliminate entire epochs instead of individual instruction and data misses. Experimental results demonstrate that the epoch-based correlation prefetches which requires minimal on-chip real estate to implement, improves the performance of a suite of important commercial benchmarks by 13% to 31% and significantly outperforms previously proposed correlation prefetchers.
Yuan Chou
MICRO1
2005 Effective Instruction Prefetching in Chip Multiprocessors for Modern Commercial Applications
abstract
In this paper, we study the instruction cache miss behavior of four modern commercial applications (a database workload, TPC-W, SPECjAppServer2002 and SPECweb99). These applications exhibit high instruction cache miss rates for both the L1 and L2 caches, and a sizable performance improvement can be achieved by eliminating these misses. We show that it is important, not only to address sequential misses, but also misses due to branches and function calls. As a result, we propose an efficient discontinuity prefetching scheme that can be effectively combined with traditional sequential prefetching to address all forms of instruction cache misses. Additionally, with the emergence of chip multiprocessors (CMPs), instruction prefetching schemes must take into account their effect on the shared L2 cache. Specifically aggressive instruction cache prefetching can result in an increase in the number of L2 cache data misses. As a solution, we propose a scheme that does not install prefetches into the L2 cache unless they are proven to be useful. Overall, we demonstrate that the combination of our proposed schemes is successful in reducing the instruction miss rate to only 10%-16% of the original miss rate and results in a 1.08X-1.37X performance improvement for the applications studied.
Lawrence Spracklen, Yuan Chou, Santosh G. Abraham
HPCA2
2005 Accurate Modeling of Aggressive Speculation in Modern Microprocessor Architectures
abstract
Computer architects utilize cycle simulators to evaluate microprocessor chip design tradeoffs and estimate performance metrics. Traditionally, cycle simulators are either trace-driven or execution-driven. In this paper, we describe ValueSim, a software layer that is interposed between a cycle simulators and either a functional simulator or a value-enhanced trace. By writing to the ValueSim API, the cycle simulator can run in either trace-driven mode or execution-driven mode, allowing it to exploit the advantages of both approaches. The ValueSim API allows a cycle simulator to accurately model a complete range of aggressive speculative mechanisms developed by computer architects, even in the trace-driven mode. Using ValueSim, we illustrate, for three key commercial applications, the significant underestimation of off-chip bandwidth, queuing delays and cache pollution when modern speculative mechanisms are not accurately modeled, highlighting the importance of accurately modeling these mechanisms in chip multiprocessor designs.
Harit Modi, Lawrence Spracklen, Yuan Chou, Santosh G. Abraham
MASCOTS3
2005 Store Memory-Level Parallelism Optimizations for Commercial Applications
abstract
This paper studies the impact of off-chip store misses on processor performance for modern commercial applications. The performance impact of off-chip store misses is largely determined by the extent of their overlap with other off-chip cache misses. The epoch MLP model is used to explain and quantify how these overlaps are affected by various store handling optimizations and by the memory consistency model implemented by the processor. The extent of these overlaps is then translated to off-chip CPI. Experimental results show that store handling optimizations are crucial for mitigating the substantial performance impact of stores in commercial applications. While some previously proposed optimizations, such as store prefetching, are highly effective, they are unable to fully mitigate the performance impact of off-chip store misses and they also leave a performance gap between the stronger and weaker memory consistency models. New optimizations, such as the store miss accelerator, an optimization of hardware scout and a new application of speculative lock elision, are demonstrated to virtually eliminate the impact of off-chip store misses.
Yuan Chou, Lawrence Spracklen, Santosh G. Abraham
MICRO1
2004 Effective stream-based and execution-based data prefetching
abstract
With processor speeds continuing to outpace the memory subsystem, cache missing memory operations continue to become increasingly important to application performance. In response to this continuing trend, most modern processors now support hardware (HW) prefetchers, which act to reduce the missing loads observed by an application.This paper analyzes the behavior of cache-missing loads in SPEC CPU2000 and highlights the inability of unit and single non-unit stride prefetchers to correctly prefetch for some commonly occurring streams. In response to this analysis, a novel multi-stride prefetcher, that supports streams with up to four distinct strides, is proposed. Performance analysis for SPEC CPU2000 illustrates that the proposed multi-stride prefetcher can outperform current stride prefetchers on several benchmarks; most notably on mcf, lucas and facerec, where it achieves an additional performance gain of up to 57%. Performance of the strided HW prefetchers is also contrasted with another recently proposed prefetch scheme, runahead execution (RAE), and the synergy between the schemes is investigated.
Sorin Iacobovici, Lawrence Spracklen, Sudarshan Kadambi, Yuan Chou, Santosh G. Abraham
ICS4
2004 Microarchitecture Optimizations for Exploiting Memory-Level Parallelism
abstract
The performance of memory-bound commercial applications such as databases is limited by increasing memory latencies. In this paper, we show that exploiting memory-level parallelism (MLP) is an effective approach for improving the performance of these applications and that microarchitecture has a profound impact on achievable MLP. Using the epoch model of MLP, we reason how traditional microarchitecture features such as out-of-order issue and state-of-the-art microarchitecture techniques such as runahead execution affect MLP. Simulation results show that a moderately aggressive out-of-order issue processor improves MLP over an in-order issue processor by 12-30%, and that aggressive handling of loads, branches and serializing instructions is needed to attain the full benefits of large out-of-order instruction windows. The results also show that a processor's issue window and reorder buffer should be decoupled to exploit MLP more efficiently. In addition, we demonstrate that runahead execution is highly effective in enhancing MLP, potentially improving the MLP of the database workload by 82% and its overall performance by 60%. Finally, our limit study shows that there is considerable headroom in improving MLP and overall performance by implementing effective instruction prefetching, more accurate branch prediction and better value prediction in addition to runahead execution.
Yuan Chou, Brian Fahs, Santosh G. Abraham
ISCA1