VLDB 2026 Research / reviewers in the wild / expert
Tianyu Feng
dblp:311/3316
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlashPS: Efficient Generative Image Editing with Mask-aware Caching and SchedulingabstractGenerative image editing using diffusion models has become a prevalent application in today's AI cloud services. In production environments, image editing typically involves a mask that specifies the regions of an image template to be edited. The use of mask provides direct control over the editing process and introduces sparsity in the model inference. In this paper, we present FlashPS, a system that efficiently serves image editing requests. The key insight behind FlashPS is that image editing only modifies the masked regions of image templates, while preserving the original content in the unmasked areas. Driven by this insight, FlashPS judiciously skips redundant computations associated with the unmask areas by reusing cached intermediate activations from previous inferences. To mitigate the high cache loading overhead, FlashPS employs a bubble-free pipeline scheme that overlaps computation with cache loading. Additionally, to reduce queuing latency in online serving while improving the GPU utilization, FlashPS proposes a novel continuous batching strategy for diffusion model serving, allowing newly arrived requests to join the running batch in just one step of denoising computation, without waiting for the entire batch to complete. As heterogenous masks induce imbalanced load, FlashPS also develops a load balancing strategy that takes into account the loads of both computation and cache loading. Collectively, FlashPS outperforms state-of-the-art diffusion serving systems for image editing, achieving up to 3× higher throughput and reducing average request latency by up to 14.7× while ensuring image quality. Xiaoxiao Jiang, Suyi Li 0002, Lingyun Yang, Tianyu Feng, Zhipeng Di, Weiyi Lu, Guoxuan Zhu, Xiu Lin, Yinghao Yu, Tao Lan, Lin Qu, Liping Zhang 0013, Wei Wang 0030 |
EuroSys | 4 |
| 2026 | Exploiting Efficient Mapping and Pipelined Execution for Accelerating SpMV on Tensor CoresabstractSparse matrix-vector multiplication (SpMV) is a fundamental operation in scientific computing, machine learning, and graph analytics, demanding efficient execution on modern hardware. Recent advances in hardware accelerators, such as Tensor Cores, have significantly improved the performance of many compute-intensive workloads. However, effectively utilizing Tensor Cores for SpMV remains challenging due to its irregular sparsity patterns and the mismatch between SpMV’s computational characteristics and constrained architecture design, leading to suboptimal performance and underutilization of Tensor Cores. In this paper, we systematically analyze the state-of-the-art SpMV optimizations on Tensor Cores, identify key performance bottlenecks, and propose Drawloom, a Tensor-Core-aware framework for SpMV with efficient Tensor Core mapping and optimized pipeline execution. Drawloom leverages a redesigned Tensor Core mapping strategy with a zig-zag chained sparse storage format, as well as a multi-stage register pipeline to better exploit hardware parallelism. Our evaluation on SuiteSparse dataset demonstrates that Drawloom outperforms cuSPARSE by 2.71×/1.90× (in FP16), 2.95×/2.39× (in FP32), and 2.47×/1.54× (in FP64) on A100 and H100 GPUs, respectively. Compared to the state-of-the-art SpMV implementations, Drawloom achieves a performance speedup of 1.26×/1.18× (in FP16) and 1.49×/1.56× (in FP64) on A100 and H100 GPUs, respectively. Kaige Zhang 0002, Hailong Yang 0002, Xin You 0001, Tianyu Feng, Yufan Xu 0001, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
PPoPP | 4 |
| 2025 | Towards Efficient LLM Inference via Collective and Adaptive Speculative DecodingabstractLarge language models (LLMs) have gained considerable attention for their remarkable performance across a wide range of tasks. However, efficient LLM inference remains challenging because of the autoregressive decoding process, which generates only one token at a time. Speculative decoding has been introduced to address the limitation by using small speculative models (SSMs) to speed up LLM inference. However, the low acceptance rate of SSMs and the high verification cost of LLM prohibit further performance improvement. In this paper, we present Smurfs, an LLM inference system designed to accelerate LLM inference through collective and adaptive speculative decoding. Smurfs adopts a majority-voted mechanism that harnesses multiple SSMs to collaboratively predict LLM outputs in multi-task scenarios, while avoiding high verification cost. It also decouples SSM speculation from LLM verification and uses a pipelined execution to hide the latency of SSM speculation. Additionally, Smurfs proposes a mechanism to dynamically determine the optimal speculation length of SSM at runtime, balancing the performance impact of accepted tokens and verification cost. The experimental results demonstrate the superiority of Smurfs in terms of inference throughput and latency compared to the state-of-the-art LLM inference systems. Hailong Yang 0002, Tongxuan Liu, Yufan Xu 0001, Xuning Liang, Kejie Ma, Tianyu Feng, Xin You 0001, Ruihao Gong, Rui Wang 0014, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
SC | 9 |
| 2025 | SimTrace: Exploiting Spatial and Temporal Sampling for Large-Scale Performance AnalysisabstractMPI tracing tools is essential to collect the communication events and performance metrics of large-scale programs for further performance analysis and optimization. However, toward the exascale era, the performance and storage overhead for tracing becomes extremely prohibitive that significantly disturbs the original execution of MPI programs, leading to distorted tracing data and thus mislead analysis results. Although process sampling can effectively reduce the tracing overhead, it can easily miss important execution information that is necessary for subsequent performance analysis. In this article, we propose SimTrace , a scalable MPI tracing tool with novel spatial and temporal sampling strategies that exploits the similarity among MPI processes to achieve both low tracing overhead as well as obtain sufficient tracing information. The experimental results demonstrate that SimTrace can significantly reduce the MPI tracing overhead compared to the state-of-the-art tracing tools, meanwhile enabling effective analysis to guide performance optimization of large-scale programs. Zhibo Xuan, Xin You 0001, Tianyu Feng, Hailong Yang 0002, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | Jigsaw: Accelerating SpMM with Vector Sparsity on Sparse Tensor CoreabstractAs deep learning models continue to grow larger, model pruning is employed to reduce memory footprint and computation complexity, which generates a large number of sparse matrix-matrix multiplication (SpMM) with unstructured sparsity (e.g., vector sparsity). However, leveraging GPU especially the newly integrated sparse tensor core (SpTC) to accelerate SpMM is quite challenging due to the unstructured sparsity. Unfortunately, existing works fail to fully exploit the SpTC on GPU due to the difficulty of satisfying the stringent requirement for restricted sparsity (e.g., 2:4 sparsity). In this paper, we propose Jigsaw, a novel method to utilize SpTC for accelerating SpMM with vector sparsity. Specifically, we propose the multi-granularity sparsity reorder method to transform the sparse data for satisfying the sparse pattern supported on SpTC. In addition, we propose a reorder-aware storage format for the transformed sparse data to better adapt to the parallelism of SpTC. Moreover, we propose corresponding optimizations to better exploit the SpTC for further accelerating SpMM. The experiment results demonstrate that Jigsaw outperforms state-of-the-art SpMM implementations and achieves promising speedup over cuBLAS. Kaige Zhang 0002, Hailong Yang 0002, Tianyu Feng, Yi Liu 0013, Zhongzhi Luan, Depei Qian 0001 |
ICPP | 4 |
| 2024 | AtRec: Accelerating Recommendation Model Training on CPUsabstractThe popularity of recommendation models and the enhanced AI processing capability of CPUs have provided massive performance opportunities to deliver satisfactory experiences to a large number of users. Unfortunately, existing recommendation model training methods fail to achieve high efficiency due to unique challenges such as dynamic shape and high parallelism. To address the above limitations, we comprehensively study the distinctive characteristics of recommendation models and discover several unexploited optimization opportunities. To exploit such opportunities, we proposeAtRec, a high-performant recommendation model training engine that significantly accelerates the training process on CPUs. Specifically,AtRecpresents comprehensive approach of training that employs operator-level and graph-level joint optimizations and runtime optimization. At the operator-level,AtRecidentifies and optimizes the time-consuming operators, which enables further efficient graph-level optimizations. At the graph-level,AtRecconducts an in-depth analysis of the inefficiencies in several frequently used subgraphs, enables further performance improvement via eliminating redundant computations and memory accesses. In addition, to achieve better runtime performance,AtRecalso identifies inefficiencies prevalent in the current scheduling and proposes runtime batching. The experiment results demonstrate thatAtReccan significantly outperform state-of-the-art recommendation model training engines. We have open sourced the implementation and corresponding data ofAtRecto boost research in this direction. Tianyu Feng, Hailong Yang 0002, Xin You 0001, Bangduo Chen, Tongxuan Liu, Zhongzhi Luan, Depei Qian 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Exploiting Input Tensor Dynamics in Activation Checkpointing for Efficient Training on GPUabstractLarger deep learning models usually lead to higher model quality, however with an ever-increasing GPU memory footprint. Although several tensor checkpointing techniques have been proposed to enable training under a restricted GPU memory budget, they fail to exploit the input tensor dynamics due to diverse datasets and subsequent data augmentation, and thus leave the training optimization on table. In this paper, we propose Mimose, an input-aware tensor checkpointing planner respecting the memory budget while enabling efficient model training on GPU. Mimose builds a lightweight but accurate prediction model of GPU memory usage online, without pre-analyzing the model. It generates a tensor checkpointing plan based on per-layer memory prediction and applies it to the training process on the fly. Our experiments show that Mimose achieves superior training throughput compared to state-of-the-art checkpointing frameworks under the same GPU memory budgets. Jianjin Liao, Mingzhen Li 0001, Hailong Yang 0002, Qingxiao Sun, Biao Sun 0002, Jiwei Hao, Tianyu Feng, Fengwei Yu, Shengdong Chen, Zhongzhi Luan, Depei Qian 0001 |
IPDPS | 7 |
| 2023 | Modularized composite attention network for continuous music emotion recognition
Meixian Zhang, Yonghua Zhu, Wenjun Zhang 0011, Yunwen Zhu, Tianyu Feng |
Multim. Tools Appl. | 5 |
| 2021 | dgQuEST: Accelerating Large Scale Quantum Circuit Simulation through Hybrid CPU-GPU Memory Hierarchies
Tianyu Feng, Xin You 0001, Shuzhang Zhong, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
NPC | 1 |