Yichuan Gao

dblp:366/5575 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0004-8351-314XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 70% Electronic design automation · 23% Performance modeling and evaluation · 7%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
0.912025
Profile-Guided Temporal Prefetching · ISCA 2025
Electronic design automation
hardware/software co-design
0.912025
Profile-Guided Temporal Prefetching · ISCA 2025
Memory systems › cache
prefetching
0.912025
Profile-Guided Temporal Prefetching · ISCA 2025
Memory systems › cache › prefetching
temporal prefetching
0.912025
Profile-Guided Temporal Prefetching · ISCA 2025
Performance modeling and evaluation
profiling
0.312025
Profile-Guided Temporal Prefetching · ISCA 2025

Methods — techniques the papers use, named apart from their topics

profile-guided optimization · 0.9counter-based profiling · 0.9
YearPublicationVenuePosition
2025 EARTH: Efficient Architecture for RISC-V Vector Memory Access
abstract
Vector processors frequently suffer from inefficient memory accesses, particularly for strided and segment memory access patterns. While coalescing strided accesses is a natural solution, effectively gathering or scattering elements at fixed strides remains a significant challenge. Naive approaches typically rely on high-overhead crossbars that remap any byte in memory or registers to any position in registers or memory, leading to physical design issues. Meanwhile, segment operations require row-column transpositions, which are often handled using either element-level in-place transposition (degrading performance) or large buffer-based bulk transposition (incurring high area overhead). Both options are undesirable, highlighting a need for more efficient solutions. In this paper, we present EARTH, a novel vector memory access architecture designed to overcome these challenges through shifting-based optimizations. For strided accesses, EARTH integrates specialized shift networks for gathering and scattering strided elements. After coalescing multiple accesses into one request within the same cache line, data can be routed between memory and registers through the shifting network with minimal overhead. For segment operations, EARTH employs a shifted register bank that enables direct column-wise access, eliminating the need for dedicated segment buffers while providing highperformance, in-place bulk transposition at acceptable overhead. We implemented the entire EARTH design on FPGA with Chisel HDL based on an open-source RISC-V vector unit Saturn. Our evaluation demonstrates that EARTH enhances performance for strided memory accesses proportionally to their prevalence in workloads, achieving $\mathbf{4 x}-\mathbf{8 x}$ speedups in benchmarks dominated by strided operations. The architecture also delivers area-efficient segment handling. Compared to conventional designs, EARTH reducing hardware area by 9% and power consumption by 41%. By optimizing these necessary memory access patterns, EARTH significantly advances both the performance and efficiency of vector processors.
Hongyi Guan, Yichuan Gao, Chenlu Miao, Mingfeng Lin, Huayue Liang
PACT2
2025 Profile-Guided Temporal Prefetching
abstract
Temporal prefetching shows promise for handling irregular memory access patterns, which are common in data-dependent and pointer-based data structures.Recent studies introduced on-chip metadata storage to reduce the memory traffic caused by accessing metadata from off-chip DRAM.However, existing prefetching schemes struggle to efficiently utilize the limited on-chip storage.An alternative solution, software indirect access prefetching, remains ineffective for optimizing temporal prefetching.In this work, we propose Prophet-a hardware-software codesigned framework that leverages profile-guided methods to optimize metadata storage management.Prophet profiles programs using counters instead of traces, injects hints into programs to guide metadata storage management, and dynamically tunes these hints to enable the optimized binary to adapt to different program inputs.Prophet is designed to coexist with existing hardware temporal prefetchers, delivering efficient, high-performance solutions for frequently executed workloads while preserving the original runtime scheme for less frequently executed workloads.Prophet outperforms the state-of-the-art temporal prefetcher, Triangel, by 14.23%, effectively addressing complex temporal patterns where prior profile-guided solutions fall short (only achieving 0.1% performance gain).Prophet delivers superior performance across all evaluated workload inputs, introducing negligible profiling, analysis, and instruction overhead.
Mengming Li, Qijun Zhang, Yichuan Gao, Wenji Fang, Yao Lu 0031, Yongqing Ren, Zhiyao Xie
ISCA3
2023 Enhancing Evaluation and Feedback in Computer Organization Labs with an Automated RISC-V Processor Verification Framework
abstract
This Practice Work-in-Progress paper presents an innovative approach for providing better evaluation and feedback for student-built Central Processing Units (CPUs) in Computer Organization course settings. A common lab project in such courses is the implementation of a simple CPU in Hardware Description Language (HDL). For example, students major in Computer Science at Tsinghua university are required to build a pipelined RISC-V processor during the autumn semester of their third year. However, due to the complexity of CPU internal states, debugging can be time-consuming for students, while assessing their designs and providing efficient feedback can be challenging for instructors. To overcome this challenge, we propose an automated verification framework using directed assembly code generation and trace comparison. Our framework uses two types of assembly test cases: one set specifically designed to trigger pipeline data and control hazards, and another set randomly generated with controlled instruction types and counts. In order to apply this industry-proven method to a wide variety of student designs, we also developed a technique for identifying common processor structures, such as register files and memory buses, that is essential for locating critical signals required for generating trace data via simulation. To evaluate the correctness of the student-designed CPUs, the recorded register and memory writing events are compared with the log generated by the Spike RISC-V ISA simulator, and the matching percentage between the two records is calculated. To provide feedback to students, we also present detailed information on the mismatches, which can help them identify bugs or areas for improvement in their designs. Tests on real student code demonstrate that our framework is more effective in detecting hidden logic errors, compared to traditional evaluation methods used in our teaching practice. The proposed method is beneficial for instructors as it saves time and improves the feedback loop in computer organization labs. Moreover, it provides additional data for assessing students' independent thinking and practical skills. With the growing popularity of the RISC-VISA in education, this approach is potentially applicable to the teaching of computer systems on a broader scale.
Yichuan Gao, Ziang Liu 0015, Weidong Liu 0001
FIE1