VLDB 2026 Research / reviewers in the wild / expert
Weilin Cai
dblp:250/1212
· DBLP profile ↗
14ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XTree on EquiMesh: Topology and Algorithm Co-Design for Collective CommunicationabstractMesh topology is widely adopted for both on-chip and chiplet-based interconnects due to its placement-friendly physical layout. However, the low-degree nodes at the edges and corners create bandwidth bottlenecks for common collectives such as AllGather and AllReduce. We address this limitation with EquiMesh, an augmented 2D-Mesh with equivalent-degree nodes without incurring switching complexity. To fully exploit EquiMesh, we propose XTree, a topology-aware algorithm that maximizes utilization of available bandwidth, and MirrorXTree, which constructs ReduceScatter and AllReduce on top of XTree through topology mirroring. Our evaluation shows that EquiMesh with XTree and MirrorXTree achieves 2× and 1.2× higher effective bandwidth than state-of-the-art mesh-based topology-algorithm co-designs for AllGather and AllReduce, respectively. Junwei Cui, Weilin Cai, Jiayi Huang 0001 |
DATE | 3 |
| 2026 | Mapping and Communication Optimizations with Fault Tolerance for Wafer-Scale LLM Inference
Junwei Cui, Weilin Cai, Jiayi Huang 0001 |
ISCA | 3 |
| 2026 | Automated emotional design generation for NEV wheel hubs: Integrating StyleGAN2-ADA and WOA-SVR within Kansei engineering
Weilin Cai, Huijuan Zhu 0002 |
Expert Syst. Appl. | 4 |
| 2025 | MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model TrainingabstractAs large language models continue to scale up, distributed training systems have expanded beyond 10k nodes, intensifying the importance of fault tolerance. Checkpoint has emerged as the predominant fault tolerance strategy, with extensive studies dedicated to optimizing its efficiency. However, the advent of the sparse Mixture-of-Experts (MoE) model presents new challenges due to the substantial increase in model size, despite comparable computational demands to dense models. Weilin Cai, Jiayi Huang 0001 |
ASPLOS (2) | 1 |
| 2025 | Shortcut-connected Expert Parallelism for Accelerating Mixture of ExpertsabstractExpert parallelism has emerged as a key strategy for distributing the computational workload of sparsely-gated mixture-of-experts (MoE) models across multiple devices, enabling the processing of increasingly large-scale models. However, the All-to-All communication inherent to expert parallelism poses a significant bottleneck, limiting the efficiency of MoE models. Although existing optimization methods partially mitigate this issue, they remain constrained by the sequential dependency between communication and computation operations.
To address this challenge, we propose ScMoE, a novel shortcut-connected MoE architecture integrated with an overlapping parallelization strategy. ScMoE decouples communication from its conventional sequential ordering, enabling up to 100\% overlap with computation.
Compared to the prevalent top-2 MoE baseline, ScMoE achieves speedups of $1.49\times$ in training and $1.82\times$ in inference.
Moreover, our experiments and analyses indicate that ScMoE not only achieves comparable but in some instances surpasses the model quality of existing approaches. Weilin Cai, Juyong Jiang, Junwei Cui, Sunghun Kim 0001, Jiayi Huang 0001 |
ICML | 1 |
| 2025 | Chimera: Communication Fusion for Hybrid Parallelism in Large Language ModelsabstractLarge Language Models (LLMs), exemplified by ChatGPT, have emerged as a predominant workload in current machine learning systems.To achieve efficient training and inference within the constraints of limited single-NPU memory capacity, deploying LLMs on multi-NPU systems typically adopt a hybrid approach that combines various parallelism patterns.This hybrid parallelism within LLMs introduces a significant amount of diverse collective communications.However, these frequent blocking communications impose a substantial burden on the multi-NPU systems.Overcoming the communication bottleneck is crucial to unlocking the potential of multi-NPU systems for efficient and scalable LLM processing.This paper introduces Chimera, a communication fusion mechanism for hybrid parallelism in LLMs.We comprehensively analyze the communication processes of each LLM parallelism pattern, identify the communication redundancy in hybrid parallelism and eliminate redundancy by fusing adjacent communication operators during parallelism transformation.By reordering operations and generating redundancy-free communication operator, Chimera effectively mitigates communication bottleneck in hybrid LLM parallelism.Our results show that Chimera achieves 1.23-7.06×network bandwidth speedup.Additionally, the end-to-end performance of LLM forward pass and backward pass on different typical multi-NPU systems achieves respective 1.32-1.58×and 1.16-1.36×speedups on average compared with those without communication fusion. Junwei Cui, Weilin Cai, Jiayi Huang 0001 |
ISCA | 3 |
| 2025 | Optimizing All-to-All Collective Communication with Fault Tolerance on Torus Networks
Junwei Cui, Weilin Cai, Meng Niu, Jiayi Huang 0001 |
MICRO | 3 |
| 2025 | A Survey on Mixture of Experts in Large Language ModelsabstractLarge language models (LLMs) have garnered unprecedented advancements across diverse fields, ranging from natural language processing to computer vision and beyond. The prowess of LLMs is underpinned by their substantial model size, extensive and diverse datasets, and the vast computational power harnessed during training, all of which contribute to the emergent abilities of LLMs (e.g., in-context learning) that are not present in small models. Within this context, the mixture of experts (MoE) has emerged as an effective method for substantially scaling up model capacity with minimal computation overhead, gaining significant attention from academia and industry. Despite its growing prevalence, there lacks a systematic and comprehensive review of the literature on MoE. This survey seeks to bridge that gap, serving as an essential resource for researchers delving into the intricacies of MoE. We first briefly introduce the structure of the MoE layer, followed by proposing a new taxonomy of MoE. Next, we overview the core designs for various MoE models including both algorithmic and systemic aspects, alongside collections of available open-source implementations, hyperparameter configurations and empirical evaluations. Furthermore, we delineate the multifaceted applications of MoE in practice, and outline some potential directions for future research. Weilin Cai, Juyong Jiang, Fan Wang 0041, Jing Tang 0004, Sunghun Kim 0001, Jiayi Huang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Reinforcement-Learning-Informed Prescriptive Analytics for Air Traffic Flow ManagementabstractAir Traffic Flow Management (ATFM) is a complex sequential decision-making problem that involves dynamically matching flights with sectors under changing environmental conditions. Finding an optimal solution for ATFM is challenging due to its dynamic nature and operational constraints. Reinforcement learning is a well-suited approach for sequential decision-making problems. However, ATFM poses three potential challenges: 1) large state space, 2) combinatorial action space, and 3) variational feasible action set, resulting from numerous agents with tightly-coupled constraints. These challenges can hinder the effectiveness of direct application of reinforcement learning methods. While prescriptive analytics can readily handle hard constraints via a mathematical optimization model, but it is computationally intractable for online sequential decision-making problems under changing environments. To address these challenges, we propose a novel framework, Reinforcement-Learning-Informed Prescriptive Analytics (RLIPA), in which an “informing” scheme is devised to integrate reinforcement learning and prescriptive analytics and leverage their strengths in predicting future reward and coping with hard constraints respectively. RLIPA is a general framework that can be adapted to other problems beyond ATFM, which typically involves many agents with tightly-coupled hard constraints. We demonstrate the usage and performance of RLIPA using numerical results and a real case study in comparison to two baseline approaches.Note to Practitioners—To improve Air Traffic Flow Management (ATFM) and reduce flight congestion, we propose a new method called reinforcement-learning-informed prescriptive analytics (RLIPA). RLIPA is a general framework that facilitates online sequential decision-making problems with multiple agents coupled with hard constraints. The approach consists of two stages: first, estimating future potential rewards for each agent via reinforcement learning, and second, informing the potential rewards to the following prescriptive analysis and using the information to construct and solve the downstream optimization problem dealing with hard coupling constraints among agents. Numerical experiments demonstrate the efficiency and effectiveness of RLIPA in the application of ATFM. In the most cases, RLIPA can offer more than 10x improvement in computational efficiency while maintaining or improving the level of optimality. The framework of RLIPA can be further extended to problems such as order dispatch in ride-hailing systems and food delivery. Yuan Wang 0069, Weilin Cai, Yilei Tu, Jianfeng Mao |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2022 | Flexible Supervision System: A Fast Fault-Tolerance Strategy for Cloud Applications in Cloud-Edge Collaborative Environments
Weilin Cai, Heng Chen 0002, Zhimin Zhuo, Ziheng Wang 0002, Ninggang An |
NPC | 1 |
| 2022 | LogSC: Model-based one-sided communication performance estimation
Ziheng Wang 0002, Heng Chen 0002, Xiaoshe Dong, Weilin Cai, Xingjun Zhang |
Future Gener. Comput. Syst. | 4 |
| 2022 | Implementation and optimization of ChaCha20 stream cipher on sunway taihuLight supercomputer
Weilin Cai, Heng Chen 0002, Ziheng Wang 0002, Xingjun Zhang |
J. Supercomput. | 1 |
| 2022 | Extending τ-Lop to model MPI blocking primitives on shared memory
Ziheng Wang 0002, Heng Chen 0002, Xiaoshe Dong, Weilin Cai, Yan Kang 0005, Xingjun Zhang |
J. Supercomput. | 4 |
| 2019 | An Exploratory Study on Judicial Image Quality Assessment Based on Deep LearningabstractImages are important judicial materials. With the deepening of intelligent systems in the judicial area, image quality plays a vital role in the result of many judicial applications. This paper firstly introduces deep learning into judicial image quality assessment. Pre-trained convolutional neural network (CNN) models are fine-tuned and then used to extract image features. Based on the features extracted from CNN models, we convert them into specific numbers representing the quality. A preliminary experiment has been designed and conducted on three types of judicial images. The experimental results show that our approach can outperform the existing image processing technique. Images used as investigation materials are more distinctive than the other two types, and they need an independent model for analyzing. Weilin Cai, Shengcheng Yu, Zhenyu Chen 0001 |
QRS | 2 |