Mengyue Xi

dblp:398/9675 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0003-5711-3110ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › instruction scheduling
instruction-level parallelism
0.912025
GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weaving · DAC 2025
Compilers and program optimization
instruction scheduling
0.912025
GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weaving · DAC 2025
GPUs and heterogeneous computing
GPU kernel optimization
0.912025
GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weaving · DAC 2025
GPUs and heterogeneous computing › GPU kernel optimization
kernel fusion
0.912025
GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weaving · DAC 2025

Methods — techniques the papers use, named apart from their topics

control flow graph · 1.7PTX-level instruction weaving · 1.7
YearPublicationVenuePosition
2025 Mpache: Interaction Aware Multi-level Cache Bypassing on GPUs
abstract
Graphics Processing Units (GPUs) are essential for general-purpose applications and are commonly leveraging multi-level caches to alleviate memory access pressure. However, the default cache management may lose opportunities for optimal performance in different applications. Although existing cache bypassing techniques tend to address this challenge, these methods predominantly concentrate on single-level cache, thus restricting their potential for further enhancements. To mitigate this issue, we propose Mpache, a novel software-based mechanism designed to bypass multi-level caches based on the characterization of load instructions. Mpache constructs an interaction graph and analyzes the cooperation and contention among instructions. Then, the profiling data of bypassing effectiveness guides Mpache to select the appropriate cache levels to bypass for each instruction. Finally, the design is integrated into the compiler to enable automatic bypassing for existing workloads. Evaluations on off-the-shelf GPUs show that Mpache achieves an average 1.15× speedup over the default cache policy, and effectively outperforms prior arts.
Mengyue Xi, Tianyu Guo 0009, Xuanteng Huang, Zejia Lin 0001, Xianwei Zhang 0001
ASP-DAC1
2025 GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weaving
abstract
GPUs have been heavily utilized in diverse applications, and numerous approaches, including kernel fusion, have been proposed to boost GPU efficiency through concurrent kernel execution. However, these approaches generally overlook the opportunities to mitigate warp stalls and improve instruction level parallelism (ILP) in inter-kernel resource sharing. To address this issue, we introduce GOPTX, a novel design for kernel fusion that improves ILP through deliberate weaving instructions at the PTX level. GOPTX establishes a merged control flow graph (CFG) from original kernels, enabling to interleaving of instructions that were sequentially executed by default and minimizing pipeline stalls on data hazards. We further propose a latency-aware instruction weaving algorithm for more efficient instruction scheduling and an adaptive code slicing method to enlarge the scheduling space. Experimental evaluation demonstrates that GOPTX achieves an average speedup of $\mathbf{1 1. 2 \%}$ over the baseline concurrent execution, with a maximum improvement of 23%. The hardware resource utilization statistics show significant enhancements in eligible warps per cycle and resource use.
Zejia Lin 0001, Mengyue Xi, Zhongchun Zheng, Wenxuan Pan, Xianwei Zhang 0001, Yutong Lu
DAC3
2025 CacheC: LLM-Based GPU Cache Management to Enhance Kernel Concurrency
Mengyue Xi, Xianwei Zhang 0001
Euro-Par (2)1