VLDB 2026 Research / reviewers in the wild / expert
Menghan Yu
dblp:329/8443
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SwiftSpec: Disaggregated Speculative Decoding and Fused Kernels for Low-Latency LLM InferenceabstractLow-latency, single-request decoding of large language models is critical for interactive systems with tight SLA demands. Prior work reduces latency through speculative decoding (combining a small draft model with a larger target model), but the draft model remains on the critical path, and communication overhead limits scaling across GPUs due to the small batch size associated with single-request decoding. To address these limitations, this paper introduces SwiftSpec: a system architecture that disaggregates draft and target models across homogeneous GPUs within a single node and utilizes NCCL-low-latency primitives directly to improve the performance of core GEMM and attention kernels. Our implementation includes 3k lines of custom CUDA for fused kernels and an evolving tree cache for KV-cache consistency and maximized reuse between draft and target models. On a single 8×H800 GPU node, SwiftSpec achieves 347 tokens/s for Llama-3-70B---1.3× faster than NVIDIA's own benchmarks on a higher-performance 8×H200 setup---and averages 1.75× faster decoding than state-of-the-art speculative decoding across five model families and six datasets. Specifically, we find that for Llama-3-70B SwiftSpec is significantly faster across all 480 tested queries, showing 1.7× speedup over the best open-source baseline for 95th percentile requests. Code for SwiftSpec will be available at https://github.com/ByteDance-Seed/SwiftSpec Ziheng Jiang, Chengquan Jiang, Menghan Yu, Size Zheng 0001, Haibin Lin, Xin Liu 0086, Henry Hoffmann |
ASPLOS (2) | 4 |
| 2026 | Collaborative Network Agents Enabled 6G Core Network Architecture
Menghan Yu, Yanxia Xing |
IWCMC | 1 |
| 2025 | A network architecture for 5G and 6G satellites based on Integrated access and backhaulabstractWith the rapid evolution of 5G and 6G technologies, satellite networks are playing a crucial role in extending global connectivity. We present a novel multi-orbit satellite network architecture based on Integrated Access and Backhaul (IAB) technology. The proposed architecture addresses key challenges, including high propagation delays in GEO satellites and frequent handovers in LEO networks, by optimizing handover strategies and ensuring seamless transitions between satellite layers. Furthermore, we propose an authorization mechanism for terminals to address the technical requirements. These advancements aim to enhance network efficiency, throughput, and user experience, particularly in remote areas where traditional infrastructure is limited. Menghan Yu |
IWCMC | 4 |
| 2025 | AI Agent Based Autonomous Cognitive Architecture for 6G Core NetworkabstractWith the growing demand for advanced communication systems and the integration of AI technologies, 6G networks are set to provide enhanced performance and enable new applications such as autonomous driving and mixed reality. This paper presents a novel AI agent-based autonomous cognitive architecture for the 6G core network. The proposed architecture leverages AI agents to autonomously perceive, understand, and act based on real-time network data, thus achieving a higher level of network intelligence and responsiveness. The architecture is designed to address the limitations of current passive AI mode in 5G by providing a proactive, adaptive, and personalized approach to network AI services. The paper discusses the system architecture, service flow, and potential benefits of AI agents in the future of 6G networks. Menghan Yu, Yanxia Xing |
IWCMC | 1 |
| 2025 | ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development
Borui Wan, Mingji Han, Yiyao Sheng, Yanghua Peng, Haibin Lin, Mofan Zhang, Zhichao Lai, Menghan Yu, Junda Zhang, Zuquan Song, Xin Liu 0086, Chuan Wu 0001 |
NSDI | 8 |
| 2025 | Understanding Stragglers in Large Model Training Using What-if Analysis
Jinkun Lin, Ziheng Jiang, Zuquan Song, Sida Zhao, Menghan Yu, Zhanghan Wang, Zuocheng Shi, Zherui Liu, Shuguang Wang, Haibin Lin, Xin Liu 0086, Aurojit Panda, Jinyang Li 0001 |
OSDI | 5 |
| 2025 | Robust LLM Training Infrastructure at ByteDanceabstractThe training scale of large language models (LLMs) has reached tens of thousands of GPUs and is still continuously expanding, enabling faster learning of larger models. Accompanying the expansion of the resource scale is the prevalence of failures (CUDA error, NaN values, job hang, etc.), which poses significant challenges to training stability. Any large-scale LLM training infrastructure should strive for minimal training interruption, efficient fault diagnosis, and effective failure tolerance to enable highly efficient continuous training. This paper presents ByteRobust, a large-scale GPU infrastructure management system tailored for robust and stable training of LLMs. It exploits the uniqueness of LLM training process and gives top priorities to detecting and recovering failures in a routine manner. Leveraging parallelisms and characteristics of LLM training, ByteRobust enables high-capacity fault tolerance, prompt fault demarcation, and localization with an effective data-driven approach, comprehensively ensuring continuous and efficient training of LLM tasks. ByteRobust is deployed on a production GPU platform with over 200,000 GPUs and advances the state of the art in training robustness by achieving 97% ETTR for a three-month training job on 9,600 GPUs. Borui Wan, Gaohong Liu, Zuquan Song, Jun Wang 0039, Guangming Sheng, Shuguang Wang, Houmin Wei, Weiqiang Lou, Mofan Zhang, Kaihua Jiang, Cheng Ren, Xiaoyun Zhi, Menghan Yu, Zhe Nan, Zhuolin Zheng, Baoquan Zhong, Qinlong Wang, Jinxin Chi, Wang Zhang 0017, Zixian Du, Sida Zhao, Jingzhe Tang, Zherui Liu, Chuan Wu 0001, Yanghua Peng, Haibin Lin, Wencong Xiao, Xin Liu 0086 |
SOSP | 16 |
| 2024 | Research on Key Technologies of 6G Network ArchitectureabstractThis article analyzes the evolution trends of mobile network technology and new application development trends, and elaborates on the needs for enhancing 6G network functionality and improving data transmission efficiency; From the perspective of network operators, we analyzed the problems caused by overly complex network architecture in 5G networks, and proposed the need to simplify 6G network architecture design and to improve network reliability. Then, this article focuses on four key technical solutions: simplification of 6G network architecture, stateless design, definition of data channels, and service based architecture of Radio Access Network (RAN). Finally, a prospect was given for the design and standardization of 6G network architecture. Yanxia Xing, Menghan Yu |
IWCMC | 4 |