EDBT 2026 Demo / reviewers in the wild / expert
Feng Yu 0019
dblp:28/1708-19
· DBLP profile ↗
5ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0004-0974-0512ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoonPoly: Bridging Code Generation and Adaptive Execution via Micro-Kernel Polymerization for Optimizing Dynamic-Shape Tensor OperatorsabstractThe prevalence of dynamic tensor shapes, driven by applications like language model serving with varying sequence lengths, is a defining characteristic of modern deep neural networks. This dynamism poses a fundamental challenge: reconciling the need for intensive, offline code generation to achieve peak performance with the demand for low-latency, adaptive execution to handle unpredictable runtime tensor shapes. Consequently, mainstream strategies are ineffective. Vendor-provided libraries, while highly optimized for a subset of common shapes, suffer performance degradation on unconventional ones. Static tensor compilers are hamstrung by prohibitive just-in-time compilation overheads for each new shape. While recent dynamic-shape compilers offer an alternative, they rely on predefined shape ranges, making them brittle when inputs fall outside these bounds. To resolve this tension, we present MoonPoly , a dynamic-shape tensor compiler that introduces micro-kernel polymerization . Our approach decouples these conflicting requirements through a two-stage process. In the offline stage, it performs intensive auto-tuning to generate a set of micro-kernels and corresponding performance models. The online stage then performs adaptive execution, rapidly assembling a near-optimal tensor operator on-the-fly, guided by a lightweight cost model. Evaluated on an NVIDIA A100 GPU, MoonPoly achieves an average operator-level speedup of 1.27× over the cuBLAS library across a diverse set of operators and data types, which in turn yields end-to-end inference acceleration for a variety of models, including BERT, the Vision Transformer, and large language models. Yangyu Zhang, Guangli Li, Feng Yu 0019, Fan Luo 0003, Qianqi Sun, Xueying Wang 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | TopServe: Task-Operator Co-scheduling for Efficient Multi-DNN Inference Serving on GPUs
Guangli Li, Feng Yu 0019, Xueying Wang 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
Euro-Par (2) | 3 |
| 2024 | Optimizing Dynamic-Shape Neural Networks on Accelerators via On-the-Fly Micro-Kernel PolymerizationabstractIn recent times, dynamic-shape neural networks have gained widespread usage in intelligent applications to address complex tasks, introducing challenges in optimizing tensor programs due to their dynamic nature. As the operators' shapes are determined at runtime in dynamic scenarios, the compilation process becomes expensive, limiting the practicality of existing static-shape tensor compilers. To address the need for effective and efficient optimization of dynamic-shape neural networks, this paper introduces MikPoly, a novel dynamic-shape tensor compiler based on micro-kernel polymerization. MikPoly employs a two-stage optimization approach, dynamically combining multiple statically generated micro-kernels using a lightweight cost model based on the shape of a tensor operator known at runtime. We evaluate the effectiveness of MikPoly by employing popular dynamic-shape operators and neural networks on two representative accelerators, namely GPU Tensor Cores and Ascend NPUs. Our experimental results demonstrate that MikPoly effectively optimizes dynamic-shape workloads, yielding an average performance improvement of 1.49× over state-of-the-art vendor libraries. Feng Yu 0019, Guangli Li, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
ASPLOS (2) | 1 |
| 2023 | VTensor: Using Virtual Tensors to Build a Layout-Oblivious AI Programming Framework
Feng Yu 0019, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
J. Comput. Sci. Technol. | 1 |
| 2020 | VTensor: Using Virtual Tensors to Build a Layout-oblivious AI Programming FrameworkabstractTensors are a popular programming interface for developing AI algorithms. Representative AI programming frameworks require developers to be always aware of tensor layouts, thereby reducing their productivity in integrating an existing operation with a new library and/or writing a new operation. We propose VTensor, a layout-oblivious virtual tensor programming interface, together with a global layout inference mechanism to resolve the layout required by virtual tensors. Furthermore, VTensor leverages a layout-oriented optimization to globally minimize the number of layout conversion operations, together with a straggler-ware scheduling algorithm and a pool-based memory allocation scheme to globally allocate resources. VTensor yields significant speedup and LOC (Lines of Codes) reduction compared to TensorFlow. Feng Yu 0019, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
PACT | 1 |