EDBT 2026 Demo / reviewers in the wild / expert
Zejia Lin 0001
dblp:281/1733-1
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-7205-4062ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bullet: Boosting GPU Utilization for LLM Serving via Dynamic Spatial-Temporal OrchestrationabstractModern large language model (LLM) serving systems confront inefficient GPU utilization due to the fundamental mismatch between compute-intensive prefill phase and memory-bound decode phase. While current practices attempt to address this by organizing these phases into hybrid batches, such solutions create an inefficient tradeoff that sacrifices either throughput or latency, leaving substantial GPU resources underutilized. For this, we identify two key root causes: 1) the prefill phase suffers from suboptimal compute utilization due to wave quantization and attention bottlenecks, and 2) hybrid batching disproportionately prioritizes latency over throughput, wasting both compute resources and memory bandwidth. To mitigate the issues, we present Bullet, a novel spatial-temporal orchestration system that eliminates these inefficiencies through fine-grained phase coordination. Bullet enables concurrent execution of prefill and decode requests, while dynamically provisioning GPU resources based on real-time performance modeling. By integrating SLO-aware scheduling and adaptive resource allocation, Bullet maximizes GPU utilization without compromising latency targets. Experimental evaluations on real-world workloads demonstrate that Bullet delivers 1.26× average throughput gains (up to 1.55×) over state-of-the-arts, while consistently meeting latency constraints. Zejia Lin 0001, Hongxin Xu, Guanyi Chen, Zhiguang Chen 0001, Yutong Lu, Xianwei Zhang 0001 |
ASPLOS (2) | 1 |
| 2025 | Mpache: Interaction Aware Multi-level Cache Bypassing on GPUsabstractGraphics Processing Units (GPUs) are essential for general-purpose applications and are commonly leveraging multi-level caches to alleviate memory access pressure. However, the default cache management may lose opportunities for optimal performance in different applications. Although existing cache bypassing techniques tend to address this challenge, these methods predominantly concentrate on single-level cache, thus restricting their potential for further enhancements. To mitigate this issue, we propose Mpache, a novel software-based mechanism designed to bypass multi-level caches based on the characterization of load instructions. Mpache constructs an interaction graph and analyzes the cooperation and contention among instructions. Then, the profiling data of bypassing effectiveness guides Mpache to select the appropriate cache levels to bypass for each instruction. Finally, the design is integrated into the compiler to enable automatic bypassing for existing workloads. Evaluations on off-the-shelf GPUs show that Mpache achieves an average 1.15× speedup over the default cache policy, and effectively outperforms prior arts. Mengyue Xi, Tianyu Guo 0009, Xuanteng Huang, Zejia Lin 0001, Xianwei Zhang 0001 |
ASP-DAC | 4 |
| 2025 | GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow WeavingabstractGPUs have been heavily utilized in diverse applications, and numerous approaches, including kernel fusion, have been proposed to boost GPU efficiency through concurrent kernel execution. However, these approaches generally overlook the opportunities to mitigate warp stalls and improve instruction level parallelism (ILP) in inter-kernel resource sharing. To address this issue, we introduce GOPTX, a novel design for kernel fusion that improves ILP through deliberate weaving instructions at the PTX level. GOPTX establishes a merged control flow graph (CFG) from original kernels, enabling to interleaving of instructions that were sequentially executed by default and minimizing pipeline stalls on data hazards. We further propose a latency-aware instruction weaving algorithm for more efficient instruction scheduling and an adaptive code slicing method to enlarge the scheduling space. Experimental evaluation demonstrates that GOPTX achieves an average speedup of $\mathbf{1 1. 2 \%}$ over the baseline concurrent execution, with a maximum improvement of 23%. The hardware resource utilization statistics show significant enhancements in eligible warps per cycle and resource use. Zejia Lin 0001, Mengyue Xi, Zhongchun Zheng, Wenxuan Pan, Xianwei Zhang 0001, Yutong Lu |
DAC | 2 |
| 2024 | MixPert: Optimizing Mixed-Precision Floating-Point Emulation on GPU Integer Tensor CoresabstractFeaturing mixed-precision tensor operations, accelerators significantly enhance performance for many error-tolerant computing tasks, but their applicability is limited in scenarios demanding high precision. While emulating higher-precision data types from lower-precision ones can bridge this gap, existing techniques either struggle to achieve sufficient accuracy or incur excessive overhead, inevitably negating performance gains. To mitigate the issue, we propose MixPert, a novel system that balances performance and accuracy via optimizing single-precision emulation on GPU Integer Tensor Cores. MixPert devises an efficient data layout and augments the computation pipeline on Tensor Cores. By deeply analyzing performance-precision trade-offs, MixPert provides users with multiple configurations based on accuracy requirements. Furthermore, MixPert can seamlessly integrate with compilers, facilitating automatic adaptation and tuning of mixed-precision parameters. Evaluations on real-world scientific computing and deep learning applications demonstrate that MixPert achieves an average speedup of 1.72× compared to cuBLAS on general-purpose cores. Beyond maintaining improved accuracy, MixPert outperforms state-of-the-art approaches APE and CUTLASS by 1.22× and 1.21×, respectively. Zejia Lin 0001, Aoyuan Sun, Xianwei Zhang 0001, Yutong Lu |
LCTES | 1 |
| 2023 | KeSCo: Compiler-based Kernel Scheduling for Multi-task GPU ApplicationsabstractNowadays, Graphics Processing Units (GPUs) dominate in a wide spectrum of computing realms and multi-task is increasingly applied in various complicated applications. To gain higher performance, multi-task programs require cumbersome programming efforts to take advantage of inter-kernel concurrency at source-code level. Although there exist works automatically scheduling kernels to enable inter-kernel concurrency, they all inevitably introduce new programming frameworks and some even bring significant performance downgrade compared to the expertise-based optimizations. To address this issue, we propose KeSCo, a compiler-based scheduler to expose kernel level concurrency in multi-task programs with trivial code modification. In compilation, KeSCo applies a strategy to schedule kernels in task queues, accounting for both load balance and synchronization cost. Also, KeSCo utilizes a customized algorithm designed for computational flow to remove redundant synchronizations. The design is further extended to support multi-process scenario, where multiple GPU processes are sharing a single context. Evaluations on representative benchmarks show that the proposed approach gains a 1.28× average speedup for multi-task scenario (1.22× for multi-process). Even with lessened programming efforts, our proposed design outperforms two state-of-the-arts GrSched and Taskflow by 1.31× and 1.16× on average, respectively. Zejia Lin 0001, Zewei Mo, Xuanteng Huang, Xianwei Zhang 0001, Yutong Lu |
ICCD | 1 |
| 2023 | Hay: Enhancing GPU Sharing Performance With Two-Level Scheduling for RayabstractGraphics Processing Units (GPUs) are extensively adopted in many clusters, providing computational services concurrently for applications from a wide spectrum of domains, especially deep learning (DL). To simplify DL training, a unified framework like Ray has been developed to deploy models on scaled clusters. Nevertheless, existing frameworks commonly choose to allocate GPUs to DL training tasks in an exclusive fashion to maximize performance. Inevitably, the dispatched tasks are incapable of occupying the ample GPU resources fully, and even worse the regular jobs are disallowed to co-locate to guarantee exclusiveness. Towards the issue, this paper proposes Hay, a resource-aware dynamic scheduler to cooperatively dispatch DL training tasks and regular workloads in GPU clusters. The design tracks the resource of all GPUs in the cluster, and models node capacity by a busyness score computed from corresponding GPUs’ utilization. Incoming tasks are processed by a two-phase heuristic policy to select the best node and GPU. Experiment results demonstrate that Hay remarkably reduces the interference between GPU-sharing tasks and achieves an average of 1.18x (up to 1.43x) performance improvement compared to the Ray scheduler. Lianghong Huang, Zejia Lin 0001, Xianwei Zhang 0001 |
ICPADS | 2 |
| 2022 | moTuner: a compiler-based auto-tuning approach for mixed-precision operatorsabstractArithmetic operators are now used in a wide spectrum of domains, including artificial intelligence, data analytics and scientific computing. Meanwhile, specialized hardware components to enable low-precision computing are increasingly deployed in GPUs and accelerators. Whereas promising to boost performance, accelerating the operators on the hardware necessitates manually tuning the mixed-precision knobs to balance the performance and accuracy, which can be extremely challenging in real practices. Zewei Mo, Zejia Lin 0001, Xianwei Zhang 0001, Yutong Lu |
CF | 2 |