Runlong Su

dblp:386/3076 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Generative modeling · 50% Deep learning architectures and training · 50%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 58% Hardware accelerators and domain-specific architectures · 33% Performance modeling and evaluation · 9%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
Fast Video Generation with Sliding Tile Attention · ICML 2025
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention
0.912025
Fast Video Generation with Sliding Tile Attention · ICML 2025
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
sliding window attention
0.912025
Fast Video Generation with Sliding Tile Attention · ICML 2025
Machine learning › Generative modeling › diffusion model › diffusion transformer
video diffusion transformer
0.912025
Fast Video Generation with Sliding Tile Attention · ICML 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
attention acceleration
0.912025
Fast Video Generation with Sliding Tile Attention · ICML 2025
Cloud and datacenter computing › inference serving
LLM serving
0.812024
Efficient LLM Scheduling by Learning to Rank · NeurIPS 2024
Cloud and datacenter computing
request scheduling
0.812024
Efficient LLM Scheduling by Learning to Rank · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

sliding tile attention · 1.7flashattention · 1.7shortest-job-first approximation · 1.5learning-to-rank · 0.8learning to rank · 0.8
YearPublicationVenuePosition
2025 Fast Video Generation with Sliding Tile Attention
abstract
Diffusion Transformers (DiTs) with 3D full attention power state-of-the-art video generation, but suffer from prohibitive compute cost -- when generating just a 5-second 720P video, attention alone takes 800 out of 950 seconds of total inference time. This paper introduces sliding tile attention (STA) to address this challenge. STA leverages the observation that attention scores in pretrained video diffusion models predominantly concentrate within localized 3D windows. By sliding and attending over local spatial-temporal region, STA eliminates redundancy from full attention. Unlike traditional token-wise sliding window attention (SWA), STA operates tile-by-tile with a novel hardware-aware sliding window design, preserving expressiveness while being \emph{hardware-efficient}. With careful kernel-level optimizations, STA offers the first efficient 2D/3D sliding-window-like attention implementation, achieving 58.79\% MFU -- 7.17× faster than prior art methods. On the leading video DiT model, Hunyuan, it accelerates attention by 1.6–10x over FlashAttention-3, yielding a 1.36–3.53× end-to-end speedup with no or minimum quality loss.
Peiyuan Zhang, Runlong Su, Hangliang Ding, Ion Stoica, Zhengzhong Liu 0001, Hao Zhang 0025
ICML3
2024 Efficient LLM Scheduling by Learning to Rank
abstract
In Large Language Model (LLM) inference, the output length of an LLM request is typically regarded as not known a priori. Consequently, most LLM serving systems employ a simple First-come-first-serve (FCFS) scheduling strategy, leading to Head-Of-Line (HOL) blocking and reduced throughput and service quality. In this paper, we reexamine this assumption -- we show that, although predicting the exact generation length of each request is infeasible, it is possible to predict the relative ranks of output lengths in a batch of requests, using learning to rank. The ranking information offers valuable guidance for scheduling requests. Building on this insight, we develop a novel scheduler for LLM inference and serving that can approximate the shortest-job-first (SJF) schedule better than existing approaches. We integrate this scheduler with the state-of-the-art LLM serving system and show significant performance improvement in several important applications: 2.8x lower latency in chatbot serving and 6.5x higher throughput in synthetic data generation. Our code is available at https://github.com/hao-ai-lab/vllm-ltr.git
Yichao Fu, Siqi Zhu, Runlong Su, Aurick Qiao, Ion Stoica, Hao Zhang 0025
NeurIPS3