Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yongxin Lyu

dblp:441/0979 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0002-2997-5285ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 93% High-performance computing · 7%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
1.012026
Thoth: Uncovering Data-Dependent Memory Access Patterns via Annotation-Directed Load Sampling · ACM Trans. Archit. Code Optim. 2026
Memory systems › memory access patterns
data-dependent memory access
1.012026
Thoth: Uncovering Data-Dependent Memory Access Patterns via Annotation-Directed Load Sampling · ACM Trans. Archit. Code Optim. 2026
Memory systems › cache › prefetching
hardware prefetching
1.012026
Thoth: Uncovering Data-Dependent Memory Access Patterns via Annotation-Directed Load Sampling · ACM Trans. Archit. Code Optim. 2026
Memory systems
memory access patterns
1.012026
Thoth: Uncovering Data-Dependent Memory Access Patterns via Annotation-Directed Load Sampling · ACM Trans. Archit. Code Optim. 2026

Methods — techniques the papers use, named apart from their topics

register-level dependency tracking · 1.0annotation-directed load sampling · 1.0
YearPublicationVenuePosition
2026 Thoth: Uncovering Data-Dependent Memory Access Patterns via Annotation-Directed Load Sampling
abstract
Sparse data structures are ubiquitous in graph analytics, machine learning, and high-performance computing. Algorithms operating on these structures typically exhibit highly irregular data-dependent memory access (DDMA) patterns, leading to frequent cache misses and degraded memory performance. Prior work on hardware prefetching to mitigate DDMA-induced misses falls into two categories: address-based methods that sample correlated sequences of load data and addresses from cache miss streams, and instruction-based methods that record instruction-level dependency chains. Although both learn single relations effectively, they struggle with multi-level range relations prevalent in DDMA-intensive workloads, leaving substantial prefetching opportunities unexploited. In address-based schemes, misses from deeper-level consumers are often miscorrelated with the producer. Moreover, out-of-order execution and the range relations themselves perturb sampling, yielding mismatched load instances. In instruction-based schemes, chain-structured representations and suboptimal learning strategies prevent the construction of complete dependency chains for these relations. To overcome these limitations, we present Thoth, a hardware prefetcher that operates at the granularity of explicit producer-consumer load pairs rather than constructing dependency chains. Thoth detects such pairs via register-level dependency tracking. It adopts an annotation-directed load sampling strategy that annotates matched producer-consumer load instances and samples only those annotated instances, thereby robustly uncovering DDMA patterns—including multi-level range relations—while avoiding mismatches. To maintain annotation correctness across pipeline flushes, Thoth employs precise load annotation, which leverages reorder identifiers to resume or terminate annotation precisely. On a suite of DDMA-intensive benchmarks, Thoth delivers a 51.1% speedup over a no-prefetching baseline and outperforms two state-of-the-art DDMA prefetchers by 14.7% and 8.2%, respectively.
Kanheng Jiang, Yongxin Lyu, Zengshi Wang, Jun Han 0003
ACM Trans. Archit. Code Optim.2