Junliang Wu

dblp:63/8874 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0003-6118-3541ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 84% Processor architecture and microarchitecture · 16%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › cache
prefetching
1.622025
Augur: Semantics-Aware Temporal Prefetching for Linked Data Structure · ACM Trans. Archit. Code Optim. 2025
Tyche: An Efficient and General Prefetcher for Indirect Memory Accesses · ACM Trans. Archit. Code Optim. 2024
Memory systems › cache › prefetching
linked data structure prefetching
0.912025
Augur: Semantics-Aware Temporal Prefetching for Linked Data Structure · ACM Trans. Archit. Code Optim. 2025
Memory systems › cache › prefetching
temporal prefetching
0.912025
Augur: Semantics-Aware Temporal Prefetching for Linked Data Structure · ACM Trans. Archit. Code Optim. 2025
Memory systems › cache › prefetching
indirect memory access prefetching
0.812024
Tyche: An Efficient and General Prefetcher for Indirect Memory Accesses · ACM Trans. Archit. Code Optim. 2024

Methods — techniques the papers use, named apart from their topics

semantic information extraction · 0.9metadata pruning · 0.9hardware prefetching · 0.8bilateral propagation · 0.8
YearPublicationVenuePosition
2025 AdaTP: Enhancing Temporal Prefetching with Adaptive Metadata Filtering
abstract
Temporal prefetching is a promising technology to predict the memory addresses of irregular memory accesses.It retains correlations of cache miss addresses in metadata, which can be stored either on-chip or off-chip.Recent advancements have favoured on-chip metadata storage within portions of the last level cache, making the optimization of metadata storage effectiveness crucial, because the benefits brought by temporal prefetching can be easily offset by the reduced capacity for data in the last level cache.However, current state-of-the-art temporal prefetchers employ static strategies to filter the metadata, which often results in suboptimal performance gains.In this study, we introduce AdaTP, a novel method that dynamically adjusts metadata filtering strategy based on the runtime measurement of data and metadata demands.This adaptive filtering strategy leverages criticality and reuse conditions of load instructions.Specifically, it permits only the most critical loads with repetitive access patterns to store correlations when metadata storage is limited, and allows all loads to store correlations when sufficient storage is available.Our evaluations show that AdaTP achieves a 22.1% speedup compared to baseline stride prefetch in irregular memory intensive benchmarks in SPEC CPU2006 and SPEC CPU2017, and outperforms state-of-the-art temporal prefetcher Triage and Triangel by 4.0% and 6.5% respectively.
Junliang Wu, Fuxin Zhang
CF1
2025 Augur: Semantics-Aware Temporal Prefetching for Linked Data Structure
abstract
Linked data structures (LDS), such as lists and trees, are widely used in modern applications. Traversing LDS typically involves a significant amount of pointer chasing. Due to the serial nature of memory access in pointer chasing, the incurred long memory latency of traversing LDS has become a critical performance bottleneck. Furthermore, the poor spatial locality in LDS makes it difficult for spatial prefetchers to predict access addresses. Although temporal prefetchers can handle irregular memory access patterns, hindered by the challenges of collecting semantic information, current state-of-the-art temporal prefetchers suffer from significant metadata redundancy and frequent metadata conflicts. Consequently, there remain substantial opportunities to enhance the LDS prefetching. To solve this problem, we propose Augur, a semantics-aware temporal prefetcher to enhance LDS performance. Augur utilizes a novel pruning method to obtain semantic information and effectively extracts node address correlations from the perspective of nodes in LDS, thereby diminishing the metadata redundancy and conflicts. Additionally, Augur employs efficient metadata management strategies that guarantee a minimal storage overhead. Evaluated on LDS workloads, Augur achieves an average performance speedup of 17.8% and 11.7% over the baseline stride prefetcher and state-of-the-art spatial prefetcher Berti, respectively. Furthermore, Augur outperforms the state-of-the-art temporal prefetcher MISB, Triage, and Triangel, by 17.4%, 12.8%, and 6.3%, respectively, with a significantly lower storage overhead of only 1.26 KB.
Junliang Wu, Chenji Han, Xinyu Li 0010, Fuxin Zhang
ACM Trans. Archit. Code Optim.2
2024 Tyche: An Efficient and General Prefetcher for Indirect Memory Accesses
abstract
Indirect memory accesses (IMAs, i.e., A [ f ( B [ i ])]) are typical memory access patterns in applications such as graph analysis, machine learning, and database. IMAs are composed of producer-consumer pairs, where the consumers’ memory addresses are derived from the producers’ memory data. Due to the built-in value-dependent feature, IMAs exhibit poor locality, making prefetching ineffective. Hindered by the challenges of recording the potentially complex graphs of instruction dependencies among IMA producers and consumers, current state-of-the-art hardware prefetchers either (a) exhibit inadequate IMA identification abilities or (b) rely on the run-ahead mechanism to prefetch IMAs intermittently and insufficiently. To solve this problem, we propose Tyche, 1 an efficient and general hardware prefetcher to enhance IMA performance. Tyche adopts a bilateral propagation mechanism to precisely excavate the instruction dependencies in simple chains with moderate length (rather than complex graphs). Based on the exact instruction dependencies, Tyche can accurately identify various IMA patterns, including nonlinear ones, and generate accurate prefetching requests continuously. Evaluated on broad benchmarks, Tyche achieves an average performance speedup of 16.2% over the state-of-the-art spatial prefetcher Berti. More importantly, Tyche outperforms the state-of-the-art IMA prefetchers IMP, Gretch, and Vector Runahead, by 15.9%, 12.8%, and 10.7%, respectively, with a lower storage overhead of only 0.57 KB.
Chenji Han, Xinyu Li 0010, Junliang Wu, Yifan Hao 0001, Zidong Du, Qi Guo 0001, Fuxin Zhang
ACM Trans. Archit. Code Optim.4