Yixin Song 0003

dblp:201/5459-3 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0001-4605-7382ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
abstract
Large multimodal models (LMMs) have demonstrated excellent capabilities in both understanding and generation tasks with various modalities. While these models can accept flexible combinations of input data, their training efficiency suffers from two major issues: pipeline stage imbalance caused by heterogeneous model architectures, and training data dynamicity stemming from the diversity of multimodal data.
Zhenliang Xue, Hanpeng Hu, Xing Chen 0009, Yixin Song 0003, Zeyu Mi, Yibo Zhu 0001, Daxin Jiang, Yubin Xia, Haibo Chen 0001
ASPLOS (2)5
2026 SolidAttention: Low-Latency SSD-based Serving on Memory-Constrained PCs
Xinrui Zheng, Dongliang Wei, Jianxiang Gao, Yixin Song 0003, Zeyu Mi, Haibo Chen 0001
FAST4
2024 PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
abstract
This paper introduces PowerInfer, a high-speed Large Language Model (LLM) inference engine on a personal computer (PC) equipped with a single consumer-grade GPU. The key principle underlying the design of PowerInfer is exploiting the high locality inherent in LLM inference, characterized by a power-law distribution in neuron activation. This distribution indicates that a small subset of neurons, termed hot neurons, are consistently activated across inputs, while the majority, cold neurons, vary based on specific inputs. PowerInfer exploits such an insight to design a GPU-CPU hybrid inference engine: hot-activated neurons are preloaded onto the GPU for fast access, while cold-activated neurons are computed on the CPU, thus significantly reducing GPU memory demands and CPU-GPU data transfers. PowerInfer further integrates adaptive predictors and neuron-aware sparse operators, optimizing the efficiency of neuron activation and computational sparsity. The evaluation shows that PowerInfer significantly outperforms llama.cpp by up to 11.69× while retaining model accuracy across various LLMs (including OPT-175B) on a single NVIDIA RTX 4090 GPU. For the OPT-30B model, PowerInfer achieves performance comparable to that of a high-end server-grade A100 GPU, reaching 82% of its token generation rate on a single consumer-grade RTX 4090 GPU.
Yixin Song 0003, Zeyu Mi, Haotong Xie, Haibo Chen 0001
SOSP1