VLDB 2026 Research / reviewers in the wild / expert
Weilin Zhao
dblp:197/5702
· DBLP profile ↗
12ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate AttentionabstractYuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao, Weilin Zhao, Ao Sun, Ziqi Yuan, Hao Zhou, Fandong Meng, Zhiyuan Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Hao Zhou 0012, Fandong Meng, Zhiyuan Liu 0001 |
ACL (1) | 5 |
| 2025 | Fusing Highly Specialized Language Models for Comprehensive ExpertiseabstractNing Ding, Yulin Chen, Ganqu Cui, Xingtai Lv, Weilin Zhao, Kaiyan Zhang, Ruobing Xie, Bowen Zhou, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ning Ding 0002, Yulin Chen 0001, Ganqu Cui, Xingtai Lv, Weilin Zhao, Ruobing Xie, Bowen Zhou 0002, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 5 |
| 2025 | APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUsabstractWhile long-context inference is crucial for advancing large language model (LLM) applications, its prefill speed remains a significant bottleneck. Current approaches, including sequence parallelism strategies and compute reduction through approximate attention mechanisms, still fall short of delivering optimal inference efficiency. This hinders scaling the inputs to longer sequences and processing long-context queries in a timely manner. To address this, we introduce APB, an efficient long-context inference framework that leverages multi-host approximate attention to enhance prefill speed by reducing compute and enhancing parallelism simultaneously. APB introduces a communication mechanism for essential key-value pairs within a sequence parallelism framework, enabling a faster inference speed while maintaining task performance. We implement APB by incorporating a tailored FlashAttn kernel alongside optimized distribution strategies, supporting diverse models and parallelism configurations. APB achieves speedups of up to 9.2\times, 4.2\times, and 1.6\times compared with FlashAttn, RingAttn, and StarAttn, respectively, without any observable task performance degradation. Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Sun Ao, Hao Zhou 0012, Jie Zhou 0016, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 5 |
| 2025 | FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative SamplingabstractWeilin Zhao, Tengyu Pan, Xu Han, Yudi Zhang, Sun Ao, Yuxiang Huang, Kaihuo Zhang, Weilun Zhao, Yuxuan Li, Jie Zhou, Hao Zhou, Jianyong Wang, Maosong Sun, Zhiyuan Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Weilin Zhao, Tengyu Pan, Xu Han 0007, Sun Ao, Yuxiang Huang 0001, Kaihuo Zhang, Wei-Lun Zhao, Jie Zhou 0016, Hao Zhou 0012, Jianyong Wang 0001, Maosong Sun 0001, Zhiyuan Liu 0001 |
ACL (1) | 1 |
| 2025 | Seq1F1B: Efficient Sequence-Level Pipeline Parallelism for Large Language Model TrainingabstractSun Ao, Weilin Zhao, Xu Han, Cheng Yang, Xinrong Zhang, Zhiyuan Liu, Chuan Shi, Maosong Sun. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sun Ao, Weilin Zhao, Xu Han 0007, Cheng Yang 0002, Zhiyuan Liu 0001, Chuan Shi 0001, Maosong Sun 0001 |
NAACL (Long Papers) | 2 |
| 2025 | BurstEngine: An efficient distributed framework for training transformers On extremely Long sequences of over 1M tokensabstractExisting methods for training LLMs on long-sequence data, such as Tensor Parallelism and Context Parallelism, exhibit low Model FLOPs Utilization as sequence lengths and number of GPUs increase, especially when sequence lengths exceed 1M tokens. To address these challenges, we propose BurstEngine, an efficient framework designed to train LLMs on long-sequence data. BurstEngine introduces BurstAttention, an optimized distributed attention with lower communication cost than RingAttention. BurstAttention leverages topology-aware ring communication to fully utilize network bandwidth and incorporates fine-grained communication-computation overlap. Furthermore, BurstEngine introduces sequence-level selective checkpointing and fuses the language modeling head with the loss function to reduce memory cost. Additionally, BurstEngine introduces workload balance optimization for various types of attention masking. By integrating these optimizations, BurstEngine achieves a 1.2 × speedup with much lower memory overhead than the state-of-the-art baselines when training LLMs on extremely long sequences of over 1M tokens. Weilin Zhao, Xu Han 0007, Cheng Yang 0002, Zhiyuan Liu 0001, Chuan Shi 0001, Maosong Sun 0001 |
SC | 2 |
| 2024 | Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex ModelsabstractXinrong Zhang, Yingfa Chen, Shengding Hu, Xu Han, Zihang Xu, Yuanwei Xu, Weilin Zhao, Maosong Sun, Zhiyuan Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yingfa Chen, Shengding Hu, Xu Han 0007, Zihang Xu, Yuanwei Xu, Weilin Zhao, Maosong Sun 0001, Zhiyuan Liu 0001 |
EMNLP | 7 |
| 2024 | Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative DecodingabstractWeilin Zhao, Yuxiang Huang, Xu Han, Wang Xu, Chaojun Xiao, Xinrong Zhang, Yewei Fang, Kaihuo Zhang, Zhiyuan Liu, Maosong Sun. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Weilin Zhao, Yuxiang Huang 0001, Xu Han 0007, Chaojun Xiao, Yewei Fang, Kaihuo Zhang, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP | 1 |
| 2024 | Predicting Emergent Abilities with Infinite Resolution EvaluationabstractThe scientific scale-up of large language models (LLMs) necessitates a comprehensive understanding of their scaling properties. However, the existing literature on the scaling properties only yields an incomplete answer: optimization loss decreases predictably as the model size increases, in line with established scaling law; yet no scaling law for task has been established and the task performances are far from predictable during scaling. Task performances typically show minor gains on small models until they improve dramatically once models exceed a size threshold, exemplifying the ''emergent abilities''. In this study, we discover that small models, although they exhibit minor performance, demonstrate critical and consistent task performance improvements that are not captured by conventional evaluation strategies due to insufficient measurement resolution. To measure such improvements, we introduce PassUntil, an evaluation strategy with theoretically infinite resolution, through massive sampling in the decoding phase. With PassUntil, we conduct a quantitative investigation into the scaling law of task performance. The investigation contains two parts. Firstly, a strict task scaling law that is not conventionally known to exist, is identified, enhancing the predictability of task performances. Remarkably, we are able to predict the performance of the 2.4B model on code generation with merely 0.05\% deviation before training starts, which is the first systematic attempt to verify predictable scaling proposed by GPT-4's report. Secondly, underpinned by PassUntil, we are able to study emergent abilities quantitatively. We identify a kind of accelerated emergence whose scaling curve cannot be fitted by standard scaling law function and has a increasing speed. We then examine two hypothesis and imply that the ``multiple circuits hypothesis'' might be responsible for the accelerated emergence. Shengding Hu, Xin Liu 0086, Xu Han 0007, Chaoqun He, Weilin Zhao, Yankai Lin 0001, Ning Ding 0002, Zebin Ou, Guoyang Zeng, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICLR | 6 |
| 2023 | H3T: Efficient Integration of Memory Optimization and Parallelism for Large-scale Transformer TrainingabstractIn recent years, big models based on Transformers have achieved state-of-the-art performance on many artificial intelligence (AI) tasks.
Despite the success of these Transformer-based models, their huge parameter size poses a serious challenge to their training, both from the storage and computation perspectives.
To this end, memory optimization (e.g., rematerialization and offloading) and parallelism (e.g., data parallelism and model parallelism) are widely explored to make training Transformers more efficient.
In this paper, we propose a framework to automatically find an efficient integration of memory optimization and parallelism for High-Throughput Transformer Training (named H3T), which is rarely considered by existing efforts for training big Transformer-based models.
Specifically, we design search algorithms to combine appropriate memory optimization strategies and parallelism schemes to achieve a balance between memory overhead and training efficiency.
We implement H3T based on an open-source toolkit BMTrain and then use H3T to train the Transformers of different sizes to evaluate the efficiency of H3T.
The experimental results show that H3T outperforms the most popular deep learning (DL) toolkit Megatron-DeepSpeed by $1.2\times \sim 4.3\times$ training speed while reducing $34.6\% \sim 80.5\%$ of memory overhead.
Moreover, H3T can use only 64 NVIDIA A100 GPUs to train GPT-3-175B, which is very difficult for existing DL toolkits. The source code is available at https://github.com/OpenBMB/BMTrain/tree/h3t. Xu Han 0007, Weilin Zhao, Guoyang Zeng, Zhiyuan Liu 0001, Maosong Sun 0001 |
NeurIPS | 3 |
| 2022 | Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language ModelsabstractDespite the great success of pre-trained language models (PLMs) in a large set of natural language processing (NLP) tasks, there has been a growing concern about their security in real-world applications. Backdoor attack, which poisons a small number of training samples by inserting backdoor triggers, is a typical threat to security. Trained on the poisoned dataset, a victim model would perform normally on benign samples but predict the attacker-chosen label on samples containing pre-defined triggers. The vulnerability of PLMs under backdoor attacks has been proved with increasing evidence in the literature. In this paper, we present several simple yet effective training strategies that could effectively defend against such attacks. To the best of our knowledge, this is the first work to explore the possibility of backdoor-free adaptation for PLMs. Our motivation is based on the observation that, when trained on the poisoned dataset, the PLM's adaptation follows a strict order of two stages: (1) a moderate-fitting stage, where the model mainly learns the major features corresponding to the original task instead of subsidiary features of backdoor triggers, and (2) an overfitting stage, where both features are learned adequately. Therefore, if we could properly restrict the PLM's adaptation to the moderate-fitting stage, the model would neglect the backdoor triggers but still achieve satisfying performance on the original task. To this end, we design three methods to defend against backdoor attacks by reducing the model capacity, training epochs, and learning rate, respectively. Experimental results demonstrate the effectiveness of our methods in defending against several representative NLP backdoor attacks. We also perform visualization-based analysis to attain a deeper understanding of how the model learns different features, and explore the effect of the poisoning ratio. Finally, we explore whether our methods could defend against backdoor attacks for the pre-trained CV model. The codes are publicly available at https://github.com/thunlp/Moderate-fitting. Biru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen, Weilin Zhao, Yangdong Deng, Zhiyuan Liu 0001, Jingang Wang, Wei Wu 0014, Maosong Sun 0001, Ming Gu 0001 |
NeurIPS | 5 |
| 2019 | Energy Efficiency Optimization in OFDM based Two-Way DF Relaying Networks with Energy HarvestingabstractIn this paper, we study energy efficiency optimization in OFDM based two-way DF relaying network with energy harvesting. Instead of time switching (TS) and power splitting (PS) schemes, we adopt OFDM modulation method, where subcarriers are divided into two groups to achieve information decoding (ID) and energy harvesting (EH) separately. The formulated EE optimization problem is non-convex constrained by the minimum information rate and the maximum transmission power. By exploiting fractional programming, an iterative resource allocation algorithm is proposed to solve the problem. Then we adopt dual decomposition and sub-gradient method to obtain the optimal variables, where subcarrier grouping and power allocation are jointly optimized to maximize the system energy efficiency. Weidang Lu, Weilin Zhao, Hong Peng 0002, Su Hu, Yuan Gao 0003 |
IWCMC | 2 |