EDBT 2026 Demo / reviewers in the wild / expert
Mengyue Xi
dblp:398/9675
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0003-5711-3110ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization › instruction scheduling
instruction-level parallelism |
0.9 | 1 | 2025 | GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weaving · DAC 2025 |
Compilers and program optimization
instruction scheduling |
0.9 | 1 | 2025 | GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weaving · DAC 2025 |
GPUs and heterogeneous computing
GPU kernel optimization |
0.9 | 1 | 2025 | GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weaving · DAC 2025 |
GPUs and heterogeneous computing › GPU kernel optimization
kernel fusion |
0.9 | 1 | 2025 | GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weaving · DAC 2025 |
Methods — techniques the papers use, named apart from their topics
control flow graph · 1.7PTX-level instruction weaving · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mpache: Interaction Aware Multi-level Cache Bypassing on GPUsabstractGraphics Processing Units (GPUs) are essential for general-purpose applications and are commonly leveraging multi-level caches to alleviate memory access pressure. However, the default cache management may lose opportunities for optimal performance in different applications. Although existing cache bypassing techniques tend to address this challenge, these methods predominantly concentrate on single-level cache, thus restricting their potential for further enhancements. To mitigate this issue, we propose Mpache, a novel software-based mechanism designed to bypass multi-level caches based on the characterization of load instructions. Mpache constructs an interaction graph and analyzes the cooperation and contention among instructions. Then, the profiling data of bypassing effectiveness guides Mpache to select the appropriate cache levels to bypass for each instruction. Finally, the design is integrated into the compiler to enable automatic bypassing for existing workloads. Evaluations on off-the-shelf GPUs show that Mpache achieves an average 1.15× speedup over the default cache policy, and effectively outperforms prior arts. Mengyue Xi, Tianyu Guo 0009, Xuanteng Huang, Zejia Lin 0001, Xianwei Zhang 0001 |
ASP-DAC | 1 |
| 2025 | GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow WeavingabstractGPUs have been heavily utilized in diverse applications, and numerous approaches, including kernel fusion, have been proposed to boost GPU efficiency through concurrent kernel execution. However, these approaches generally overlook the opportunities to mitigate warp stalls and improve instruction level parallelism (ILP) in inter-kernel resource sharing. To address this issue, we introduce GOPTX, a novel design for kernel fusion that improves ILP through deliberate weaving instructions at the PTX level. GOPTX establishes a merged control flow graph (CFG) from original kernels, enabling to interleaving of instructions that were sequentially executed by default and minimizing pipeline stalls on data hazards. We further propose a latency-aware instruction weaving algorithm for more efficient instruction scheduling and an adaptive code slicing method to enlarge the scheduling space. Experimental evaluation demonstrates that GOPTX achieves an average speedup of $\mathbf{1 1. 2 \%}$ over the baseline concurrent execution, with a maximum improvement of 23%. The hardware resource utilization statistics show significant enhancements in eligible warps per cycle and resource use. Zejia Lin 0001, Mengyue Xi, Zhongchun Zheng, Wenxuan Pan, Xianwei Zhang 0001, Yutong Lu |
DAC | 3 |
| 2025 | CacheC: LLM-Based GPU Cache Management to Enhance Kernel Concurrency
Mengyue Xi, Xianwei Zhang 0001 |
Euro-Par (2) | 1 |