Quang Duong 0002

dblp:39/5770-2 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2026
0009-0003-5247-171XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 95% Processor architecture and microarchitecture · 5%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › cache › prefetching
temporal prefetching
1.822026
Streamlined on-Chip Temporal Prefetching · HPCA 2026
A New Formulation of Neural Data Prefetching · ISCA 2024
Memory systems › cache
prefetching
1.012026
Streamlined on-Chip Temporal Prefetching · HPCA 2026
Memory systems › cache › prefetching
data prefetching
0.812024
A New Formulation of Neural Data Prefetching · ISCA 2024
Memory systems › memory access optimization
neural prefetching
0.812024
A New Formulation of Neural Data Prefetching · ISCA 2024
Memory systems
memory-bound computation
0.312026
Streamlined on-Chip Temporal Prefetching · HPCA 2026
Processor architecture and microarchitecture › memory system microarchitecture
hardware prefetcher
0.212024
A New Formulation of Neural Data Prefetching · ISCA 2024

Methods — techniques the papers use, named apart from their topics

metadata compression · 1.0neural network · 0.8
YearPublicationVenuePosition
2026 Streamlined on-Chip Temporal Prefetching
abstract
In this paper, we present the Streamline temporal prefetcher, which introduces a stream-based metadata representation that produces three significant benefits over Triangel, the previous state-of-the-art in temporal prefetching. First, it removes redundancy present in Triangel's metadata. Second, it prioritizes the storage of those metadata entries that have higher prefetch utility. Third, it eliminates the need for the untenable LLC traffic that is induced when Triangel dynamically adjusts the size of its metadata store. The end result is that for memory-intensive SPEC 2006, SPEC 2017, and GAP benchmarks, Streamline outperforms Triangel by 6.7 percentage points on an 8-core system with a stride prefetcher. This performance benefit comes from Streamline's improved storage efficiency: It holds 33% more correlations than Triangel, which translates to$\mathbf{1 2. 5}$percentage points better prefetch coverage. Significantly, Streamline matches the performance of Triangel even when Triangel is given twice the metadata storage.
Quang Duong 0002, Calvin Lin
HPCA1
2024 A New Formulation of Neural Data Prefetching
abstract
Temporal data prefetchers have the potential to produce significant performance gains by prefetching irregular data streams. Recent work has introduced a neural model for temporal prefetching that outperforms practical table-based temporal prefetchers, but the large storage and latency costs, along with the inability to generalize to memory addresses outside of the training dataset, prevent such a neural network from seeing any practical use in hardware. In this paper, we reformulate the temporal prefetching prediction problem so that neural solutions to it are more amenable for hardware deployment. Our key insight is that while temporal prefetchers typically assume that each address can be followed by any possible successor, there are empirically only a few successors for each address. Utilizing this insight, we introduce a new abstraction of memory addresses, and we show how this abstraction enables the design of a much more efficient neural prefetcher. Our new prefetcher, Twilight, improves upon the previous state-of-the-art neural prefetcher, Voyager, in multiple dimensions: It reduces latency by $988 \times$, shrinks storage by $10.8 \times$, achieves 4% more speedup on a mix of irregular SPEC 2006, SPEC 2017, and GAP benchmarks, and is capable of predicting new temporal correlations not present in the training data. Twilight outperforms idealized versions of the non-neural temporal prefetchers STMS by 12.2% and Domino by 8.5%. While Twilight is still not practical, T-LITE, a slimmed-down version of Twilight that can prefetch across different program runs, further reduces latency and storage ($1421 \times$ faster and $142 \times$ smaller than Voyager), matches Voyager’s performance and outperforms the practical non-neural Triage prefetcher by 5.9%.
Quang Duong 0002, Akanksha Jain, Calvin Lin
ISCA1