VLDB 2026 Research / reviewers in the wild / expert
Juncai Liu
dblp:304/3355
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0001-5783-731XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionabstractWe present MegaScale-MoE, a production system tailored for the efficient training of large-scale mixture-of-experts (MoE) models. MoE emerges as a promising architecture to scale large language models (LLMs) to unprecedented sizes, thereby enhancing model performance. However, existing MoE training systems experience a degradation in training efficiency, exacerbated by the escalating scale of MoE models and the continuous evolution of hardware. Chao Jin 0007, Ziheng Jiang, Zhihao Bai, Juncai Liu, Xiang Li 0067, Ningxin Zheng, Qi Huang 0001, Wen Heng, Yiyuan Ma, Wenlei Bao, Size Zheng 0001, Xuegui Zheng, Yanghua Peng, Haibin Lin, Xuanzhe Liu, Xin Jin 0008, Xin Liu 0086 |
EuroSys | 5 |
| 2025 | DAPO: An Open-Source LLM Reinforcement Learning System at ScaleabstractInference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the **D**ecoupled Clip and **D**ynamic s**A**mpling **P**olicy **O**ptimization (**DAPO**) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL. Qiying Yu, Zheng Zhang 0001, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, Lingjun Liu, Xin Liu 0039, Haibin Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang 0022, Mofan Zhang, Ru Zhang 0006, Wang Zhang 0017, Jiaze Chen, Jiangjie Chen, Hongli Yu, Yuxuan Song 0002, Xiangpeng Wei, Hao Zhou 0012, Wei-Ying Ma, Ya-Qin Zhang, Mingxuan Wang |
NeurIPS | 10 |
| 2025 | Integrating clinical knowledge and imaging for medical report generation
Juncai Liu, Hongyu Shen, Mingtao Pei |
Pattern Recognit. Lett. | 2 |
| 2024 | Automatic Radiology Reports Generation via Memory Alignment NetworkabstractThe automatic generation of radiology reports is of great significance, which can reduce the workload of doctors and improve the accuracy and reliability of medical diagnosis and treatment, and has attracted wide attention in recent years. Cross-modal mapping between images and text, a key component of generating high-quality reports, is challenging due to the lack of corresponding annotations. Despite its importance, previous studies have often overlooked it or lacked adequate designs for this crucial component. In this paper, we propose a method with memory alignment embedding to assist the model in aligning visual and textual features to generate a coherent and informative report. Specifically, we first get the memory alignment embedding by querying the memory matrix, where the query is derived from a combination of the visual features and their corresponding positional embeddings. Then the alignment between the visual and textual features can be guided by the memory alignment embedding during the generation process. The comparison experiments with other alignment methods show that the proposed alignment method is less costly and more effective. The proposed approach achieves better performance than state-of-the-art approaches on two public datasets IU X-Ray and MIMIC-CXR, which further demonstrates the effectiveness of the proposed alignment method. Hongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing Tian |
AAAI | 3 |
| 2023 | Janus: A Unified Distributed Training Framework for Sparse Mixture-of-Experts ModelsabstractScaling models to large sizes to improve performance has led a trend in deep learning, and sparsely activated Mixture-of-Expert (MoE) is a promising architecture to scale models. However, training MoE models in existing systems is expensive, mainly due to the All-to-All communication between layers. Juncai Liu, Hui Wang 0011 |
SIGCOMM | 1 |
| 2022 | PipeCompress: Accelerating Pipelined Communication for Distributed Deep LearningabstractDistributed learning is widely used to accelerate the training of deep learning models, but it is known that communication efficiency limits the scalability of distributed learning systems. Current gradient compression techniques provide promising methods to reduce communication time, but the extra time incurred by compression is not negligible. After compression techniques are applied, the communication time is significantly reduced because the data size needed to communicate becomes much smaller, but compressing gradients is time-consuming and it becomes a new bottleneck. In this paper, we design and implement PipeCompress, a system to decouple compression and backpropagation operations into two processes and pipeline the two processes to hide compression time. We also propose a specialized inter-process communication mechanism based on the characteristics of DNN distributed training to improve the efficiency of passing messages between the two processes, which makes sure that the decoupling does not bring much extra inter-process communication time cost. As far as we know, this is the first work that notices the overhead of compression and pipelines backpropagation and compression operations to hide compression time in distributed learning. Experiments show that PipeCompress can significantly hide compression time, reduce iteration time, and accelerate the training process on various DNN models. Juncai Liu, Hui Wang 0011, Chenghao Rong, Jilong Wang 0001 |
ICC | 1 |
| 2022 | Exploring the Layered Structure of Containers for Design of Video Analytics Application MigrationabstractThe existing solutions to the migration of container-based applications are not suitable for live video analytics applications because these solutions can result in excessive migration time. Intuitively, it is possible to exploit the layered structure of containers to improve the migration performance, but we still need a good understanding of the characteristics of the containers related to video analytics applications to justify the intuitive idea and design a high-performance solution. In this paper, we pull 3735 representative images from Docker Hub. We analyze these images and get three main findings: (1) the images of video analytics applications have more layers and larger sizes than that of general applications; (2) we can cache images in the destination servers and reuse the same layers cached in the destination server to reduce the migration time when an image is migrated from its source server to its destination server; (3) the size of the remaining data to be transferred during migrations is still large and a high-performance migration scheme is still necessary. Based on the above findings, we propose a pipelined migration scheme to optimize migration performance. Evaluations show that pipelined migration performs significantly better than other migration schemes. Chenghao Rong, Hui Wang 0011, Juncai Liu, Jilong Wang 0001 |
WCNC | 3 |
| 2022 | Scheduling Massive Camera Streams to Optimize Large-Scale Live Video AnalyticsabstractIn smart cities, more and more government departments will make use of live analytics of videos from surveillance cameras in their tasks, such as vehicle traffic monitoring and criminal detection. Obviously, it is costly for each individual department to deploy its own infrastructure,i.e., cameras and analytics system. In this paper, we consider a scenario in which a city deploys an infrastructure and departments submit requests to access and analyze videos for their own purposes. The live analytics of massive streams is computation-intensive and the tasks might be latency-critical, which makes scheduling massive streams to optimize all tasks an essential and challenging work. We exploit an end-edge-cloud architecture and propose an adaptive system to schedule the massive camera streams and tasks, which considers all factors affecting the computation and networking resource consumption,e.g., sharing of model computation, video quality, model partition, and task placement. Particularly, the resource consumption ofFaster R-CNN + ResNet101under each partition scheme is profiled for the first time and we notice the partition must be used together with lossless compression techniques to be beneficial. Furthermore, sometimes tasks might be required to migrate because the scheduling decision made by the system changes to adapt to the changing resource supply and demand. In order to avoid the performance degradation during migration, we propose a non-destructive migration scheme and implement it in the system. Simulations demonstrate our system achieves a total utility close to the maximum and our analytics system performs better than state-of-the-art solutions. Chenghao Rong, Hui Wang 0011, Juncai Liu, Jilong Wang 0001, Sharon X. Huang |
IEEE/ACM Trans. Netw. | 3 |
| 2021 | FedPA: An adaptively partial model aggregation strategy in Federated Learning
Juncai Liu, Hui Wang 0011, Chenghao Rong, Yuedong Xu 0001, Jilong Wang 0001 |
Comput. Networks | 1 |