An-Chow Lai

dblp:69/6971 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
0since 2021 · last 2002
0000-0002-5086-9631ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-authorSoftware engineering, systems software and programming languages · 3 · 3 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.122000
Selective, accurate, and timely self-invalidation using last-touch prediction · ISCA 2000
Memory Sharing Predictor: The Key to a Speculative Coherent DSM · ISCA 1999
Memory systems › shared memory
distributed shared memory
0.022000
Memory Sharing Predictor: The Key to a Speculative Coherent DSM · ISCA 1999
Selective, accurate, and timely self-invalidation using last-touch prediction · ISCA 2000
Memory systems
cache
0.012001
Dead-block prediction & dead-block correlating prefetchers · ISCA 2001
Memory systems › cache › prefetching
data prefetching
0.012001
Dead-block prediction & dead-block correlating prefetchers · ISCA 2001
Memory systems › cache management
dead block prediction
0.012001
Dead-block prediction & dead-block correlating prefetchers · ISCA 2001
Memory systems › cache coherence › invalidation
self-invalidation
0.012000
Selective, accurate, and timely self-invalidation using last-touch prediction · ISCA 2000

Methods — techniques the papers use, named apart from their topics

simulation · 0.1trace-based prediction · 0.0address correlation · 0.0trace-based correlation · 0.0pattern-based prediction · 0.0
YearPublicationVenuePosition
2002 Optimizing Traffic in DSM Clusters: Fine-Grain Memory Caching versus Page Migration/Replication
An-Chow Lai, Babak Falsafi
Theory Comput. Syst.1
2001 Dead-block prediction & dead-block correlating prefetchers
abstract
Effective data prefetching requires accurate mechanisms to predict both “which” cache blocks to prefetch and “when” to prefetch them. This paper proposes the Dead-Block Predictors (DBPs), trace-based predictors that accurately identify “when” an Ll data cache block becomes evictable or “dead”. Predicting a dead block significantly enhances prefetching lookahead and opportunity, and enables placing data directly into Ll, obviating the need for auxiliary prefetch buffers. This paper also proposes Dead-Block Correlating Prefetchers (DBCPs), that use address correlation to predict “which” subsequent block to prefetch when a block becomes evictable. A DBCP enables effective data prefetching in a wide spectrum of pointer-intensive, integer, and floating-point applications.
An-Chow Lai, Cem Fide, Babak Falsafi
ISCA1
2000 Selective, accurate, and timely self-invalidation using last-touch prediction
abstract
Communication in cache-coherent distributed shared memory (DSM) often requires invalidating (or writing back) cached copies of a memory block, incurring high overheads. This paper proposes Last-Touch Predictors (LTPs) that learn and predict the "last touch" to a memory block by one processor before the block is accessed and subsequently invalidated by another. By predicting a last-touch and (self-)invalidating the block in advance, an LTP hides the invalidation time, significantly reducing the coherence overhead. The key behind accurate last-touch prediction is trace-based correlation, associating a last-touch with the sequence of instructions (i.e. a trace) touching the block from a coherence miss until the block is invalidated. Correlating instructions enables an LTP to identify a last-touch to a memory block uniquely throughout an application's execution. In this paper we use results from running shared-memory applications on a simulated DSM to evaluate LTPs. The results indicate that: (1) our base case LTP design maintaining trace signatures on a per-block basis, substantially improves prediction accuracy over previous self-invalidation schemes to an average of 79%; (2) our alternative LTP design, maintaining a global trace signature table, reduces storage overhead but only achieves an average accuracy of 58%; (3) last-touch prediction based on a single instruction only achieves an average accuracy of 41% due to instruction reuse within and across computation; and (4) LTP enables selective, accurate, and timely self-invalidation in DSM, speeding up program execution on average by 11%.
An-Chow Lai, Babak Falsafi
ISCA1
2000 Comparing the effectiveness of fine-grain memory caching against page migration/replication in reducing traffic in DSM clusters
abstract
In this paper, we compare and contrast two techniques to improve capacity/conflict miss traffic in CC-NUMA DSM clusters. Page migration/replication optimizes read-write accesses to a page used by a single processor by migrating the page to that processor and replicates all read-shared pages in the sharers' local memories. R-NUMA optimizes read-write accesses to any page by allowing a processor to cache that page in its main memory. Page migration/replication requires less hardware complexity as compared to R-NUMA, but has limited applicability and incurs much higher overheads even with tuned hardware/software support.
An-Chow Lai, Babak Falsafi
SPAA1
1999 Memory Sharing Predictor: The Key to a Speculative Coherent DSM
abstract
Recent research advocates using general message predictors to learn and predict the coherence activity in distributed shared memory (DSM). By accurately predicting a message and timely invoking the necessary coherence actions, a DSM can hide much of the remote access latency. This paper proposes the Memory Sharing Predictors (MSPs), pattern-based predictors that significantly improve prediction accuracy and implementation cost over general message predictors. An MSP is based on the key observation that to hide the remote access latency, a predictor must accurately predict only the remote memory accesses (i.e., request messages) and not the subsequent coherence messages invoked by an access. Simulation results indicate that MSPs improve prediction accuracy over general message predictors from 81% to 93% while requiring less storage overhead. This paper also presents the first design and evaluation for a speculative coherent DSM using pattern-based predictors. We identify simple techniques and mechanisms to trigger prediction timely and perform speculation for remote read accesses. Our speculation hardware readily works with a conventional full-map write-invalidate coherence protocol without any modifications. Simulation results indicate that performing speculative read requests alone reduces execution times by 12% in our shared-memory applications.
An-Chow Lai, Babak Falsafi
ISCA1