Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xin Tan 0004

dblp:89/6413-4 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-3785-9700ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 56% Efficient and distributed learning · 35% Deep learning architectures and training · 8%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
diffusion transformer
1.012026
Dynamic Sparsity in Large-Scale Video DiT Training · ASPLOS (1) 2026
Machine learning › Efficient and distributed learning › model compression › sparsity
dynamic sparsity
1.012026
Dynamic Sparsity in Large-Scale Video DiT Training · ASPLOS (1) 2026
Machine learning › Generative modeling
video generation
1.012026
Dynamic Sparsity in Large-Scale Video DiT Training · ASPLOS (1) 2026
Cloud and datacenter computing
cluster resource management and scheduling
0.912025
Towards End-to-End Optimization of LLM-based Applications with Ayo · ASPLOS (2) 2025
Cloud and datacenter computing
serverless computing
0.912025
Towards End-to-End Optimization of LLM-based Applications with Ayo · ASPLOS (2) 2025
Machine learning › Deep learning architectures and training
attention mechanism
0.312026
Dynamic Sparsity in Large-Scale Video DiT Training · ASPLOS (1) 2026
Machine learning › Efficient and distributed learning › inference efficiency
LLM inference optimization
0.312025
Towards End-to-End Optimization of LLM-based Applications with Ayo · ASPLOS (2) 2025

Methods — techniques the papers use, named apart from their topics

large language model · 1.7
YearPublicationVenuePosition
2026 Dynamic Sparsity in Large-Scale Video DiT Training
abstract
Diffusion Transformers (DiTs) have shown remarkable performance in generating high-quality videos. However, the quadratic complexity of 3D full attention remains a bottleneck in scaling DiT training, especially with high-definition, lengthy videos, where it can consume up to 95% of processing time and demand specialized context parallelism.
Xin Tan 0004, Yuetao Chen, Xing Chen 0009, Kun Yan 0004, Nan Duan 0001, Yibo Zhu 0001, Daxin Jiang, Hong Xu 0001
ASPLOS (1)1
2026 ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
Xin Tan 0004, Minchen Yu, Jingzong Li, Hong Xu 0001
IWQoS2
2026 Dynamic Compute and Network Orchestration for Disaggregated RL
abstract
Disaggregating the generation and training stages in RL is widely adopted to scale LLM post-training. There are two critical challenges here. First, the generation stage often becomes a bottleneck due to dynamic workload shifts and severe execution imbalances. Second, the decoupled stages result in diverse and dynamic network traffic patterns that strain the conventional static fabric.
Xin Tan 0004, Yicheng Feng, Yu Zhou 0008, Yibo Zhu 0001, Hong Xu 0001
SIGCOMM1
2025 Towards End-to-End Optimization of LLM-based Applications with Ayo
abstract
Large language model (LLM)-based applications consist of both LLM and non-LLM components, each contributing to the end-to-end latency. Despite great efforts to optimize LLM inference, end-to-end workflow optimization has been overlooked. Existing frameworks employ coarse-grained orchestration with task modules, which confines optimizations to within each module and yields suboptimal scheduling decisions.
Xin Tan 0004, Hong Xu 0001
ASPLOS (2)1
2024 Arlo: Serving Transformer-based Language Models with Dynamic Input Lengths
abstract
A prominent challenge in serving requests for NLP tasks is handling the varying length of input texts. Existing solutions, such as uniform zero-padding and compiler support, suffer from either computational inefficiency or suboptimal latency. To address these practical issues, we propose an approach called polymorphing. Polymorphing involves creating and utilizing multiple runtimes of the model, each statically compiled with a different input length, to serve requests accordingly. This fine-grained use of statically-compiled runtimes reduces the overheads of zero-padding while improving latency performance compared to dynamic compilation. To practically realize polymorphing, we have developed an inference scheduling system, Arlo, which leverages the observed input length distribution to periodically allocate compute resources across multiple runtimes by solving an integer linear program. Upon request arrival, Arlo uses a multi-level queue-based heuristic to dispatch requests to the most suitable runtime instances, efficiently adapting to the dynamics of request length and instance load. Extensive testbed evaluations and large-scale simulations using production traces demonstrate Arlo’s promising potential. It achieves 23.7%–98.1% mean latency reductions compared to existing schemes while significantly reducing tail latency.
Xin Tan 0004, Jiamin Li 0002, Jingzong Li, Hong Xu 0001
ICPP1