VLDB 2026 Research / reviewers in the wild / expert
Yunqi Gao
dblp:272/0694
· DBLP profile ↗
14ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STAR-GS: Spatio-Temporal Geometry Alignment and Generative Refinement for Sparse-View 4D Gaussian SplattingabstractRecent 4D Gaussian Splatting (4DGS) methods for reconstructing dynamic scenes have achieved unified spatiotemporal modeling under densely captured multi-view inputs. Nevertheless, reconstruction from sparse-view video sequences remains highly challenging due to unreliable Gaussian initialization and insufficient photometric supervision. We present STAR-GS, a novel framework for sparse-view 4D reconstruction that integrates Spatio-Temporal Geometry Alignment and Generative View Refinement. To address unreliable Gaussian initialization, we develop the Geometry Alignment that combines feed-forward geometry prediction with cross-temporal joint camera optimization, resolving inter-frame similarity ambiguities and establishing a unified global Gaussian initialization. To compensate for insufficient supervision, we further incorporate the Generative View Refinement based on a reference-conditioned single-step diffusion model, which synthesizes high-fidelity novel views to provide dense and temporally consistent photometric guidance for 4DGS optimization. Extensive experiments demonstrate that STAR-GS significantly improves reconstruction quality under sparse-view settings. Ablation studies further validate the effectiveness and complementary contributions of the proposed components. Yunqi Gao, Zhanfeng Liao, Dongbo Zhou, Leyuan Liu 0001 |
ICMR | 1 |
| 2026 | DFQ+: Dynamic queuing for approximate fairness in programmable shared memory switches
Minghui Chang, Yunqi Gao, Bing Hu 0002, Pei Xiao 0001, Chunming Wu 0001, Liyan Li |
Comput. Networks | 2 |
| 2025 | Dynamic Queuing for Approximate Fairness in Programmable Shared Memory SwitchesabstractTo ensure fair bandwidth allocation for diverse application flows from data centers, effective bandwidth management in switches is critical. Modern switches often adopt shared memory architectures to enhance efficiency. Fair queuing mechanisms can achieve fair bandwidth allocation in switches. However, the state-of-the-art fair queuing mechanisms in shared memory switches suffer from excessive packet drops, leading to suboptimal network utilization. In this paper, we propose Dynamic Fair Queuing (DFQ), a novel mechanism that leverages a limited number of priority queues to achieve both high network utilization and fair bandwidth allocation. DFQ is based on two key novel ideas. First, DFQ presents dynamic admission thresholds to manage packet enqueuing by monitoring the accumulated arrived packet bits and the remaining buffer of the queues in real time. Second, DFQ employs queue splitting and merging to maximize the utilization of the shared memory pool while guaranteeing fairness. Simulation results demonstrate that DFQ significantly improves throughput, fairness, and network utilization, while reducing flow completion time by up to 44.1%. Minghui Chang, Yunqi Gao, Bing Hu 0002, Pei Xiao 0001, Shicong Zhang, Chenhui Gu, Yisha Liu |
HPSR | 2 |
| 2025 | ClothHMR: 3D Mesh Recovery of Humans in Diverse Clothing from Single ImageabstractWith 3D data rapidly emerging as an important form of multimedia information, 3D human mesh recovery technology has also advanced accordingly. However, current methods mainly focus on handling humans wearing tight clothing and perform poorly when estimating body shapes and poses under diverse clothing, especially loose garments. To this end, we make two key insights: (1) tailoring clothing to fit the human body can mitigate the adverse impact of clothing on 3D human mesh recovery, and (2) utilizing human visual information from large foundational models can enhance the generalization ability of the estimation. Based on these insights, we propose ClothHMR, to accurately recover 3D meshes of humans in diverse clothing. ClothHMR primarily consists of two modules: clothing tailoring (CT) and FHVM-based mesh recovering (MR). The CT module employs body semantic estimation and body edge prediction to tailor the clothing, ensuring it fits the body silhouette. The MR module optimizes the initial parameters of the 3D human mesh by continuously aligning the intermediate representations of the 3D mesh with those inferred from the foundational human visual model (FHVM). ClothHMR can accurately recover 3D meshes of humans wearing diverse clothing, precisely estimating their body shapes and poses. Experimental results demonstrate that ClothHMR significantly outperforms existing state-of-the-art methods across benchmark datasets and in-the-wild images. Additionally, a web application for online fashion and shopping powered by ClothHMR is developed, illustrating that ClothHMR can effectively serve real-world usage scenarios. The code and model for ClothHMR are available at: https://github.com/starVisionTeam/ClothHMR. Yunqi Gao, Leyuan Liu 0001, Yuhan Li 0009, Changxin Gao, Yuanyuan Liu 0004, Jingying Chen 0001 |
ICMR | 1 |
| 2025 | FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts TrainingabstractThe parameter size of modern large language models (LLMs) can be scaled up to the trillion-level via the sparsely-activated Mixture-of-Experts (MoE) technique to avoid excessive increase of the computational costs. To further improve training efficiency, pipelining computation and communication has become a promising solution for distributed MoE training. However, existing work primarily focuses on scheduling tasks within the MoE layer, such as expert computing and all-to-all (A2A) communication, while neglecting other key operations including multi-head attention (MHA) computing, gating, and all-reduce communication. In this paper, we propose FlowMoE, a scalable framework for scheduling multi-type task pipelines. First, FlowMoE constructs a unified pipeline to consistently scheduling MHA computing, gating, expert computing, and A2A communication. Second, FlowMoE introduces a tensor chunk-based priority scheduling mechanism to overlap the all-reduce communication with all computing tasks. We implement FlowMoE as an adaptive and generic framework atop PyTorch. Extensive experiments with 675 typical MoE layers and four real-world MoE models across two GPU clusters demonstrate that our proposed FlowMoE framework outperforms state-of-the-art MoE training frameworks, reducing training time by14%-57%, energy consumption by 10%-39%, and memory usage by 7%-32%. FlowMoE’s code is anonymously available at https://anonymous.4open.science/r/FlowMoE. Yunqi Gao, Bing Hu 0002, Mahdi Boloursaz Mashhadi, A-Long Jin, Yanfeng Zhang 0001, Pei Xiao 0001, Rahim Tafazolli, Mérouane Debbah |
NeurIPS | 1 |
| 2025 | HeaPS: Heterogeneity-aware participant selection for efficient federated learning
Duo Yang 0005, Bing Hu 0002, Yunqi Gao, A-Long Jin, Kwan Lawrence Yeung |
J. Parallel Distributed Comput. | 3 |
| 2025 | PipeSFL: A Fine-Grained Parallelization Framework for Split Federated Learning on Heterogeneous ClientsabstractSplit Federated Learning (SFL) improves scalability of Split Learning (SL) by enabling parallel computing of the learning tasks on multiple clients. However, state-of-the-art SFL schemes neglect the effects of heterogeneity in the clients’ computation and communication performance as well as the computation time for the tasks offloaded to the cloud server. In this paper, we propose a fine-grained parallelization framework, called PipeSFL, to accelerate SFL on heterogeneous clients. PipeSFL is based on two key novel ideas. First, we design a server-side priority scheduling mechanism to minimize per-iteration time. Second, we propose a hybrid training mode to reduce per-round time, which employs asynchronous training within rounds and synchronous training between rounds. We theoretically prove the optimality of the proposed priority scheduling mechanism within one round and analyze the total time per round for PipeSFL, SFL and SL. We implement PipeSFL on PyTorch. Extensive experiments on seven 64-client clusters with different heterogeneity demonstrate that at training speed, PipeSFL achieves up to 1.65x and 1.93x speedup compared to EPSL and SFL, respectively. At energy consumption, PipeSFL saves up to 30.8% and 43.4% of the energy consumed within each training round compared to EPSL and SFL, respectively. Yunqi Gao, Bing Hu 0002, Mahdi Boloursaz Mashhadi, Wei Wang 0021, Mehdi Bennis |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | VS: Reconstructing Clothed 3D Human from Single Image via Vertex ShiftabstractVarious applications require high-fidelity and artifact free 3D human reconstructions. However, current implicit function-based methods inevitably produce artifacts while existing deformation methods are difficult to reconstruct high-fidelity humans wearing loose clothing. In this paper, we propose a two-stage deformation method named Vertex Shift (VS) for reconstructing clothed 3D humans from single images. Specifically, VS first stretches the estimated SMPL-X mesh into a coarse 3D human model using shift fields inferred from normal maps, then refines the coarse 3D human model into a detailed 3D human model via a graph convolutional network embedded with implicit-function-learned features. This “stretch-refine” strategy addresses large deformations required for reconstructing loose clothing and delicate deformations for recovering intricate and detailed surfaces, achieving high-fidelity reconstructions that faithfully convey the pose, clothing, and surface details from the input images. The graph convolutional network's ability to exploit neighborhood vertices coupled with the advantages inherited from the deformation methods ensure VS rarely produces artifacts like distortions and non-human shapes and never produces artifacts like holes, broken parts, and dismembered limbs. As a result, VS can reconstruct highfidelity and artifact-less clothed 3D humans from single images, even under scenarios of challenging poses and loose clothing. Experimental results on three benchmarks and two in-the-wild datasets demonstrate that VS significantly outperforms current state-of-the-art methods. The code and models of VS are available for research purposes at https://github.com/starVisionTeam/VS. Leyuan Liu 0001, Yuhan Li 0009, Yunqi Gao, Changxin Gao, Yuanyuan Liu 0004, Jingying Chen 0001 |
CVPR | 3 |
| 2024 | GWPF: Communication-efficient federated learning with Gradient-Wise Parameter Freezing
Duo Yang 0005, Yunqi Gao, Bing Hu 0002, A-Long Jin, Wei Wang 0021 |
Comput. Networks | 2 |
| 2024 | DGS: An Efficient Delay-Guaranteed Scheduling Framework for Wireless Deterministic NetworkingabstractDeterministic Networking (DetNet) aims to provide an end-to-end ultra-reliable data network with ultra-low latency and jitter. However, implementing DetNet in wireless networks, particularly in the air interface, still faces the challenge of guaranteeing bounded delay. This paper proposes a delay-guaranteed three-layer scheduling framework for DetNet, named Deterministic Guarantee Scheduling (DGS). The top layer calculates the amount of new data entering the queue in each scheduling period and timestamps the data to track its arrival time. Based on the remaining waiting time of each flow’s data volume, the middle layer proposes a scheduling algorithm based on urgency, prioritizing the scheduling of data volumes with the shortest remaining queuing time. The lower layer fine-tunes the scheduling results obtained by the middle layer for actual transmission. We implemented the DGS framework on the 5G-air-simulator platform. Simulation results demonstrate that DGS outperforms all other mechanisms by guaranteeing delay for a larger number of deterministic flows and achieving better throughput performance. Minghui Chang, Haojun Lv, Yunqi Gao, Bing Hu 0002, Wei Wang 0021 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2024 | US-Byte: An Efficient Communication Framework for Scheduling Unequal-Sized Tensor Blocks in Distributed Deep LearningabstractThe communication bottleneck severely constrains the scalability of distributed deep learning, and efficient communication scheduling accelerates distributed DNN training by overlapping computation and communication tasks. However, existing approaches based on tensor partitioning are not efficient and suffer from two challenges: 1) the fixed number of tensor blocks transferred in parallel can not necessarily minimize the communication overheads; 2) although the scheduling order that preferentially transmits tensor blocks close to the input layer can start forward propagation in the next iteration earlier, the shortest per-iteration time is not obtained. In this paper, we propose an efficient communication framework called US-Byte. It can schedule unequal-sized tensor blocks in a near-optimal order to minimize the training time. We build the mathematical model of US-Byte by two phases: 1) the overlap of gradient communication and backward propagation, and 2) the overlap of gradient communication and forward propagation. We theoretically derive the optimal solution for the second phase and efficiently solve the first phase with a low-complexity algorithm. We implement the US-Byte architecture on PyTorch framework. Extensive experiments on two different 8-node GPU clusters demonstrate that US-Byte can achieve up to 1.26x and 1.56x speedup compared to ByteScheduler and WFBP, respectively. We further exploit simulations of 128 GPUs to verify the potential scaling performance of US-Byte. Simulation results show that US-Byte can achieve up to 1.69x speedup compared to the state-of-the-art communication framework. Yunqi Gao, Bing Hu 0002, Mahdi Boloursaz Mashhadi, A-Long Jin, Pei Xiao 0001, Chunming Wu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | Single-image clothed 3D human reconstruction guided by a well-aligned parametric body model
Leyuan Liu 0001, Yunqi Gao, Jianchi Sun, Jingying Chen 0001 |
Multim. Syst. | 2 |
| 2023 | OF-WFBP: A near-optimal communication mechanism for tensor fusion in distributed deep learning
Yunqi Gao, Zechao Zhang, Bing Hu 0002, A-Long Jin, Chunming Wu 0001 |
Parallel Comput. | 1 |
| 2021 | HEI-Human: A Hybrid Explicit and Implicit Method for Single-View 3D Clothed Human Reconstruction
Leyuan Liu 0001, Jianchi Sun, Yunqi Gao, Jingying Chen 0001 |
PRCV (2) | 3 |