EDBT 2026 Demo / reviewers in the wild / expert
Yongqing Ren
dblp:41/6276
· DBLP profile ↗
5ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0007-3570-965XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 80% Electronic design automation · 15% Performance modeling and evaluation · 4% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache |
1.7 | 2 | 2025 | Profile-Guided Temporal Prefetching · ISCA 2025 Integrating Prefetcher Selection with Dynamic Request Allocation Improves Prefetching Efficiency · HPCA 2025 |
Memory systems › cache › prefetching
hardware prefetching |
0.9 | 1 | 2025 | Integrating Prefetcher Selection with Dynamic Request Allocation Improves Prefetching Efficiency · HPCA 2025 |
Electronic design automation
hardware/software co-design |
0.9 | 1 | 2025 | Profile-Guided Temporal Prefetching · ISCA 2025 |
Memory systems › cache
prefetching |
0.9 | 1 | 2025 | Profile-Guided Temporal Prefetching · ISCA 2025 |
Memory systems › cache › prefetching
temporal prefetching |
0.9 | 1 | 2025 | Profile-Guided Temporal Prefetching · ISCA 2025 |
Memory systems
memory hierarchy |
0.3 | 1 | 2025 | Integrating Prefetcher Selection with Dynamic Request Allocation Improves Prefetching Efficiency · HPCA 2025 |
Performance modeling and evaluation
profiling |
0.3 | 1 | 2025 | Profile-Guided Temporal Prefetching · ISCA 2025 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 0.9profile-guided optimization · 0.9dynamic request allocation · 0.9counter-based profiling · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Integrating Prefetcher Selection with Dynamic Request Allocation Improves Prefetching EfficiencyabstractHardware prefetching plays a critical role in hiding the off-chip DRAM latency. The complexity of applications results in a wide variety of memory access patterns, prompting the development of numerous cache-prefetching algorithms. Consequently, commercial processors often employ a hybrid of these algorithms to enhance the overall prefetching performance. Nonetheless, since these prefetchers share hardware resources, conflicts arising from competing prefetching requests can negate the benefits of hardware prefetching. Under such circumstances, several prefetcher selection algorithms have been proposed to mitigate conflicts between prefetchers. However, these prior solutions suffer from two limitations. First, the input demand request allocation is inaccurate. Second, the prefetcher selection criteria are coarse-grained. In this paper, we address both limitations by introducing an efficient and widely applicable prefetcher selection algorithm—Alecto1, which tailors the demand requests for each prefetcher. Every demand request is first sent to Alecto to identify suitable prefetchers before being routed to prefetchers for training and prefetching. Our analysis shows that Alecto is adept at not only harmonizing prefetching accuracy, coverage, and timeliness but also significantly enhancing the utilization of the prefetcher table, which is vital for temporal prefetching. Alecto outperforms the state-of-the-art RL-based prefetcher selection algorithm—Bandit by $2.76 \%$ in single-core, and $\mathbf{7. 5 6 \%}$ in eight-core. For memory-intensive benchmarks, Alecto outperforms Bandit by $\mathbf{5. 2 5 \%}$. Alecto consistently delivers state-of-the-art performance in scheduling various types of cache prefetchers. In addition to the performance improvement, Alecto can reduce the energy consumption associated with accessing the prefetchers’ table by $48 \%$ ($7 \%$ energy reduction on the entire memory hierarchy), while only adding less than 1 KB of storage overhead.1The name Alecto stands for the combination of selection and allocation. Mengming Li, Qijun Zhang, Yongqing Ren, Zhiyao Xie |
HPCA | 3 |
| 2025 | Profile-Guided Temporal PrefetchingabstractTemporal prefetching shows promise for handling irregular memory access patterns, which are common in data-dependent and pointer-based data structures.Recent studies introduced on-chip metadata storage to reduce the memory traffic caused by accessing metadata from off-chip DRAM.However, existing prefetching schemes struggle to efficiently utilize the limited on-chip storage.An alternative solution, software indirect access prefetching, remains ineffective for optimizing temporal prefetching.In this work, we propose Prophet-a hardware-software codesigned framework that leverages profile-guided methods to optimize metadata storage management.Prophet profiles programs using counters instead of traces, injects hints into programs to guide metadata storage management, and dynamically tunes these hints to enable the optimized binary to adapt to different program inputs.Prophet is designed to coexist with existing hardware temporal prefetchers, delivering efficient, high-performance solutions for frequently executed workloads while preserving the original runtime scheme for less frequently executed workloads.Prophet outperforms the state-of-the-art temporal prefetcher, Triangel, by 14.23%, effectively addressing complex temporal patterns where prior profile-guided solutions fall short (only achieving 0.1% performance gain).Prophet delivers superior performance across all evaluated workload inputs, introducing negligible profiling, analysis, and instruction overhead. Mengming Li, Qijun Zhang, Yichuan Gao, Wenji Fang, Yao Lu 0031, Yongqing Ren, Zhiyao Xie |
ISCA | 6 |
| 2010 | Dynamic Resource Tuning for Flexible Core Chip Multiprocessors
Yongqing Ren, Hong An, Yaobin Wang |
ICA3PP (2) | 1 |
| 2010 | FACRA: Flexible-Core Architecture Chip Resource AbstractorabstractA family of flexible-core chip multiprocessors (FCMPs) has been recently proposed to allow simple, identical physical cores to be aggregated dynamically to form larger and more powerful logical processors. However, such flexible-core architecture faces a new significant scheduling problem in the operating system, which traditionally assumes only fixed-number and fixed-granularity processors. This paper proposes a framework, called FACRA, that employs low-level runtime software to simplify OS resource allocation and process scheduling on FCMPs. Through exporting a simple, uniform processor abstraction on flexible-core chip resource, FACRA provides a set of functions with uniform interface for system-level scheduling on FCMPs. To verify the design, FACRA is built on TFlex (a typical FCMP) in our experiments, and two well known process schedulers, round-robin and dynamic-priority scheduler of Linux 2.6.11, are modified to schedule on TFlex. The evaluation results demonstrate that FACRA can efficiently simplify OS resource allocation and process scheduling on FCMPs with negligible performance loss. Hong An, Yongqing Ren, Mengjie Mao, Mu Xu, Qi Li 0034 |
PDCAT | 3 |
| 2007 | Balancing Thread Partition for Efficiently Exploiting Speculative Thread-Level Parallelism
Yaobin Wang, Hong An, Yongqing Ren |
APPT | 6 |