Yuhao Zhou 0004

dblp:121/6722-4 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
17since 2021 · last 2025
0000-0001-8074-6416ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Ferret: An Efficient Online Continual Learning Framework under Varying Memory Constraints
abstract
In the realm of high-frequency data streams, achieving real-time learning within varying memory constraints is paramount. This paper presents Ferret, a comprehensive framework designed to enhance online accuracy of Online Continual Learning (OCL) algorithms while dynamically adapting to varying memory budgets. Ferret employs a finegrained pipeline parallelism strategy combined with an iterative gradient compensation algorithm, ensuring seamless handling of high-frequency data with minimal latency, and effectively counteracting the challenge of stale gradients in parallel training. To adapt to varying memory budgets, its automated model partitioning and pipeline planning optimizes performance regardless of memory limitations. Extensive experiments across 20 benchmarks and 5 integrated OCL algorithms show Ferret’s remarkable efficiency, achieving up to 3.7× lower memory overhead to reach the same online accuracy compared to competing methods. Furthermore, Ferret consistently outperforms these methods across diverse memory budgets, underscoring its superior adaptability. These findings position Ferret as a premier solution for efficient and adaptive OCL framework in real-time environments.
Yuhao Zhou 0004, Jindi Lv, Mingjia Shi, Jiancheng Lv 0001
CVPR1
2025 EA-Vit: Efficient Adaptation for Elastic Vision Transformer
abstract
Vision Transformers (ViTs) have emerged as a foundational model in computer vision, excelling in generalization and adaptation to downstream tasks. However, deploying ViTs to support diverse resource constraints typically requires retraining multiple, size-specific ViTs, which is both time-consuming and energy-intensive. To address this issue, we propose an efficient ViT adaptation framework that enables a single adaptation process to generate multiple models of varying sizes for deployment on platforms with various resource constraints. Our approach comprises two stages. In the first stage, we enhance a pre-trained ViT with a nested elastic architecture that enables structural flexibility across MLP expansion ratio, number of attention heads, embedding dimension, and network depth. To preserve pre-trained knowledge and ensure stable adaptation, we adopt a curriculum-based training strategy that progressively increases elasticity. In the second stage, we design a lightweight router to select submodels according to computational budgets and downstream task demands. Initialized with Pareto-optimal configurations derived via a customized NSGA-II algorithm, the router is then jointly optimized with the backbone. Extensive experiments on multiple benchmarks demonstrate the effectiveness and versatility of EA-ViT. The code is available at https://github.com/zcxcf/EA-ViT.
Wangbo Zhao, Yuhao Zhou 0004, Weidong Tang, Shuo Wang 0001, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Kai Wang 0036
ICCV4
2025 FedSH: Tackling Staleness By Scheduling High-order Approximation in Asynchronous Federated Learning
abstract
In recent years, Federated Learning (FL) has made significant progress in utilizing decentralized data and protecting data privacy, but communication bottlenecks remain a key challenge. Asynchronous Federated Learning (AFL) addresses this by enhancing training speed through improved asynchronous capabilities. However, AFL primarily faces the issue of gradient staleness and current methods do not consider when approximate estimations are most effective or when they lead to significant errors. To address these issues, we developed FedSH (Federated Scheduled High-order Approximation), which minimizes the large errors in approximate estimates and stabilizes gradient directions, resulting in improved performance in AFL while maintaining similar training times. The experimental results show that the proposed FedSH are commendable on at least 6 benchmarks.
Haixin Gao, Mingjia Shi, Yuhao Zhou 0004, Deng Xiong, Jiancheng Lv 0001
IJCNN3
2025 Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
abstract
Modern Parameter-Efficient Fine-Tuning (PEFT) methods such as low-rank adaptation (LoRA) reduce the cost of customizing large language models (LLMs), yet still require a separate optimization run for every downstream dataset. We introduce \textbf{Drag-and-Drop LLMs (\textit{DnD})}, a prompt-conditioned parameter generator that eliminates per-task training by mapping a handful of unlabeled task prompts directly to LoRA weight updates. A lightweight text encoder distills each prompt batch into condition embeddings, which are then transformed by a cascaded hyper-convolutional decoder into the full set of LoRA matrices. Once trained in a diverse collection of prompt-checkpoint pairs, DnD produces task-specific parameters in seconds, yielding i) up to \textbf{12,000$\times$} lower overhead than full fine-tuning, ii) average gains up to \textbf{30\%} in performance over the strongest training LoRAs on unseen common-sense reasoning, math, coding, and multimodal benchmarks, and iii) robust cross-domain generalization improving \textbf{40\%} performance without access to the target data or labels. Our results demonstrate that prompt-conditioned parameter generation is a viable alternative to gradient-based adaptation for rapidly specializing LLMs. We open source \href{https://jerryliang24.github.io/DnD}{our project} in support of future research.
Zhiyuan Liang, Dongwen Tang, Yuhao Zhou 0004, Xuanlei Zhao, Mingjia Shi, Wangbo Zhao, Peihao Wang, Konstantin Schürholt, Damian Borth, Michael M. Bronstein, Yang You 0001, Zhangyang Wang, Kai Wang 0036
NeurIPS3
2025 Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths
abstract
Mamba has proven efficient for long-sequence modeling in vision tasks. However, when token reduction techniques are applied to improve efficiency, Mamba-based models exhibit drastic performance degradation compared to Vision Transformers (ViTs). This decline is potentially attributed to Mamba's chain-like scanning mechanism, which we hypothesize not only induces cascading losses in token connectivity but also limits the diversity of spatial receptive fields. In this paper, we propose Asymmetric Multi-scale Vision Mamba (AMVim), a novel architecture designed to enhance pruning robustness. AMVim employs a dual-path structure, integrating a window-aware scanning mechanism into one path while retaining sequential scanning in the other. This asymmetry design promotes token connection diversity and enables multi-scale information flow, reinforcing spatial awareness. Empirical results demonstrate that AMVim achieves state-of-the-art pruning robustness. During token reduction, AMVim-T achieves a substantial 34\% improvement in training-free accuracy with identical model sizes and FLOPs. Meanwhile, AMVim-S exhibits only a 1.5\% accuracy drop, performing comparably to ViT. Notably, AMVim also delivers superior performance during pruning-free settings, further validating its architectural advantages.
Jindi Lv, Yuhao Zhou 0004, Mingjia Shi, Zhiyuan Liang, Xiaojiang Peng, Wangbo Zhao, Jiancheng Lv 0001, Kai Wang 0036
NeurIPS2
2025 REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
abstract
Diffusion Transformers (DiTs) deliver state-of-the-art image quality, yet their training remains notoriously slow. A recent remedy---representation alignment (REPA) that matches DiT hidden features to those of a non-generative teacher (e.g., DINO)---dramatically accelerates the early epochs but plateaus or even degrades performance later. We trace this failure to the capacity mismatch: once the generative student begins modeling the joint data distribution, the teacher's lower-dimensional embeddings and attention patterns become a straitjacket rather than a guide. We then introduce HASTE (Holistic Alignment with Stage-wise Termination for Efficient training), a two-phase schedule that keeps the help and drops the hindrance. Phase I applies a holistic alignment loss that simultaneously distills attention maps (relational priors) and feature projections (semantic anchors) from the teacher into mid-level layers of the DiT, yielding rapid convergence. Phase II then performs one-shot termination that deactivates the alignment loss, once a simple trigger such as a fixed iteration is hit, freeing the DiT to focus on denoising and exploit its generative capacity. HASTE speeds up training of diverse DiTs without architecture changes. On ImageNet 256×256, it reaches the vanilla SiT-XL/2 baseline FID in 50 epochs and matches REPA’s best FID in 500 epochs, amounting to a 28× reduction in optimization steps. HASTE also improves text-to-image DiTs on MS-COCO, proving to be a simple yet principled recipe for efficient diffusion training across various tasks.
Wangbo Zhao, Yuhao Zhou 0004, Zhiyuan Liang, Mingjia Shi, Xuanlei Zhao, Kaipeng Zhang, Zhangyang Wang, Kai Wang 0036, Yang You 0001
NeurIPS3
2025 E-3SFC: Communication-Efficient Federated Learning With Double-Way Features Synthesizing
abstract
The exponential growth in model sizes has significantly increased the communication burden in federated learning (FL). Existing methods to alleviate this burden by transmitting compressed gradients often face high compression errors, which slow down the model's convergence. To simultaneously achieve high compression effectiveness and lower compression errors, we study the gradient compression problem from a novel perspective. Specifically, we propose a systematical algorithm termed extended single-step synthetic features compressing (E-3SFC), which consists of three subcomponents, i.e., the single-step synthetic features compressor (3SFC), a double-way compression (DWC) algorithm, and a communication budget scheduler (BS). First, we regard the process of gradient computation of a model as decompressing gradients from corresponding inputs, while the inverse process is considered as compressing the gradients. Based on this, we introduce a novel gradient compression method termed 3SFC, which utilizes the model itself as a decompressor, leveraging training priors such as model weights and objective functions. The 3SFC compresses raw gradients into tiny synthetic features in a single-step simulation, incorporating error feedback (EF) to minimize overall compression errors. To further reduce communication overhead, 3SFC is extended to E-3SFC, allowing DWC and dynamic communication budget scheduling. Our theoretical analysis under both strongly convex and nonconvex conditions demonstrates that 3SFC achieves linear and sublinear convergence rates with aggregation noise. Extensive experiments across six datasets and six models reveal that 3SFC outperforms the state-of-the-art methods by up to 13.4% while reducing communication costs by 111.6 times. These findings suggest that 3SFC can significantly enhance communication efficiency in FL without compromising model performance.
Yuhao Zhou 0004, Mingjia Shi, Yanan Sun 0001, Jiancheng Lv 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Federated CINN Clustering for Accurate Clustered Federated Learning
abstract
Federated Learning (FL) presents an innovative approach to privacy-preserving distributed machine learning and enables efficient crowd intelligence on a large scale. However, a significant challenge arises when coordinating FL with crowd intelligence which diverse client groups possess disparate objectives due to data heterogeneity or distinct tasks. To address this challenge, we propose the Federated cINN Clustering Algorithm (FCCA) to robustly cluster clients into different groups, avoiding mutual interference between clients with data heterogeneity, and thereby enhancing the performance of the global model. Specifically, FCCA utilizes a global encoder to transform each client’s private data into multivariate Gaussian distributions. It then employs a generative model to learn encoded latent features through maximum likelihood estimation, which eases optimization and avoids mode collapse. Finally, the central server collects converged local models to approximate similarities between clients and thus partition them into distinct clusters. Extensive experimental results demonstrate FCCA’s superiority over other state-of-the-art clustered federated learning algorithms, evaluated on various models and datasets. These results suggest that our approach has substantial potential to enhance the efficiency and accuracy of real-world federated learning tasks.
Yuhao Zhou 0004, Minjia Shi, Jiancheng Lv 0001
ICASSP1
2024 DeFTA: A plug-and-play peer-to-peer decentralized federated learning framework
Yuhao Zhou 0004, Minjia Shi, Jiancheng Lv 0001
Inf. Sci.1
2023 Communication-efficient Federated Learning with Single-Step Synthetic Features Compressor for Faster Convergence
abstract
Reducing communication overhead in federated learning (FL) is challenging but crucial for large-scale distributed privacy-preserving machine learning. While methods utilizing sparsification or other techniques can largely reduce the communication overhead, the convergence rate is also greatly compromised. In this paper, we propose a novel method named Single-Step Synthetic Features Compressor (3SFC) to achieve communication-efficient FL by directly constructing a tiny synthetic dataset containing synthetic features based on raw gradients. Therefore, 3SFC can achieve an extremely low compression rate when the constructed synthetic dataset contains only one data sample. Additionally, the compressing phase of 3SFC utilizes a similarity-based objective function so that it can be optimized with just one step, considerably improving its performance and robustness. To minimize the compressing error, error feedback (EF) is also incorporated into 3SFC. Experiments on multiple datasets and models suggest that 3SFC has significantly better convergence rates compared to competing methods with lower compression rates (i.e., up to 0.02%). Furthermore, ablation studies and visualizations show that 3SFC can carry more information than competing methods for every communication round, further validating its effectiveness.
Yuhao Zhou 0004, Mingjia Shi, Yanan Sun 0001, Jiancheng Lv 0001
ICCV1
2023 Unconstrained Feature Model and Its General Geometric Patterns in Federated Learning: Local Subspace Minority Collapse
Mingjia Shi, Yuhao Zhou 0004, Jiancheng Lv 0001
ICONIP (8)2
2023 Violence-MFAS: Audio-Visual Violence Detection Using Multimodal Fusion Architecture Search
Dan Si, Jindi Lv, Yuhao Zhou 0004, Jiancheng Lv 0001
ICONIP (14)4
2023 PRIOR: Personalized Prior for Reactivating the Information Overlooked in Federated Learning
abstract
Classical federated learning (FL) enables training machine learning models without sharing data for privacy preservation, but heterogeneous data characteristic degrades the performance of the localized model. Personalized FL (PFL) addresses this by synthesizing personalized models from a global model via training on local data. Such a global model may overlook the specific information that the clients have been sampled. In this paper, we propose a novel scheme to inject personalized prior knowledge into the global model in each client, which attempts to mitigate the introduced incomplete information problem in PFL. At the heart of our proposed approach is a framework, the $\textit{PFL with Bregman Divergence}$ (pFedBreD), decoupling the personalized prior from the local objective function regularized by Bregman divergence for greater adaptability in personalized scenarios. We also relax the mirror descent (RMD) to extract the prior explicitly to provide optional strategies. Additionally, our pFedBreD is backed up by a convergence analysis. Sufficient experiments demonstrate that our method reaches the $\textit{state-of-the-art}$ performances on 5 datasets and outperforms other methods by up to 3.5% across 8 benchmarks. Extensive analyses verify the robustness and necessity of proposed designs. The code will be made public.
Mingjia Shi, Yuhao Zhou 0004, Kai Wang 0036, Huaizheng Zhang, Shudong Huang, Jiancheng Lv 0001
NeurIPS2
2022 FLSGD: free local SGD with parallel synchronization
Yuhao Zhou 0004, Mingjia Shi, Jiancheng Lv 0001
J. Supercomput.2
2022 Correction to: FLSGD: free local SGD with parallel synchronization
Yuhao Zhou 0004, Mingjia Shi, Jiancheng Lv 0001
J. Supercomput.2
2022 Communication-Efficient Federated Learning With Compensated Overlap-FedAvg
abstract
While petabytes of data are generated each day by a number of independent computing devices, only a few of them can be finally collected and used for deep learning (DL) due to the apprehension of data security and privacy leakage, thus seriously retarding the extension of DL. In such a circumstance, federated learning (FL) was proposed to perform model training by multiple clients' combined data without the dataset sharing within the cluster. Nevertheless, federated learning with periodic model averaging (FedAvg) introduced massive communication overhead as the synchronized data in each iteration is about the same size as the model, and thereby leading to a low communication efficiency. Consequently, variant proposals focusing on the communication rounds reduction and data compression were proposed to decrease the communication overhead of FL. In this article, we propose Overlap-FedAvg, an innovative framework that loosed the chain-like constraint of federated learning and paralleled the model training phase with the model communication phase (i.e., uploading local models and downloading the global model), so that the latter phase could be totally covered by the former phase. Compared to vanilla FedAvg, Overlap-FedAvg was further developed with a hierarchical computing strategy, a data compensation mechanism, and a nesterov accelerated gradients (NAG) algorithm. In Particular, Overlap-FedAvg is orthogonal to many other compression methods so that they could be applied together to maximize the utilization of the cluster. Besides, the theoretical analysis is provided to prove the convergence of the proposed framework. Extensive experiments conducting on both image classification and natural language processing tasks with multiple models and datasets also demonstrate that the proposed framework substantially reduced the communication overhead and boosted the federated learning process.
Yuhao Zhou 0004, Jiancheng Lv 0001
IEEE Trans. Parallel Distributed Syst.1
2021 LANA: Towards Personalized Deep Knowledge Tracing Through Distinguishable Interactive Sequences
Yuhao Zhou 0004, Xihua Li 0002, Yunbo Cao, Xuemin Zhao, Jiancheng Lv 0001
EDM1
2020 HPSGD: Hierarchical Parallel SGD with Stale Gradients Featuring
Yuhao Zhou 0004, Jiancheng Lv 0001
ICONIP (2)1