Mingjia Shi

dblp:271/0186 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
14since 2021 · last 2025
0000-0002-9988-3741ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2025 A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training
abstract
Training diffusion models is always a computation-intensive task. In this paper, we introduce a novel speed-up method for diffusion model training, called SpeeD, which is based on a closer look at time steps. Our key findings are: i) Time steps can be empirically divided into acceleration, deceleration, and convergence areas based on the process increment. ii) These time steps are imbalanced, with many concentrated in the convergence area. iii) The concentrated steps provide limited benefits for diffusion training. To address this, we design an asymmetric sampling strategy that reduces the frequency of steps from the convergence area while increasing the sampling probability in other areas. Additionally, we propose a weighting strategy to emphasize the importance of time steps with rapid-change process increments. As a plug-and-play and architecture-agnostic approach, SpeeD consistently achieves 3 × acceleration across various diffusion architectures, datasets, and tasks. Notably, due to its simple design, our approach significantly reduces the cost of diffusion model training with minimal overhead. Our research enables more researchers to train diffusion models at a lower cost.
Kai Wang 0036, Mingjia Shi, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Hanwang Zhang, Yang You 0001
CVPR2
2025 Ferret: An Efficient Online Continual Learning Framework under Varying Memory Constraints
abstract
In the realm of high-frequency data streams, achieving real-time learning within varying memory constraints is paramount. This paper presents Ferret, a comprehensive framework designed to enhance online accuracy of Online Continual Learning (OCL) algorithms while dynamically adapting to varying memory budgets. Ferret employs a finegrained pipeline parallelism strategy combined with an iterative gradient compensation algorithm, ensuring seamless handling of high-frequency data with minimal latency, and effectively counteracting the challenge of stale gradients in parallel training. To adapt to varying memory budgets, its automated model partitioning and pipeline planning optimizes performance regardless of memory limitations. Extensive experiments across 20 benchmarks and 5 integrated OCL algorithms show Ferret’s remarkable efficiency, achieving up to 3.7× lower memory overhead to reach the same online accuracy compared to competing methods. Furthermore, Ferret consistently outperforms these methods across diverse memory budgets, underscoring its superior adaptability. These findings position Ferret as a premier solution for efficient and adaptive OCL framework in real-time environments.
Yuhao Zhou 0004, Jindi Lv, Mingjia Shi, Jiancheng Lv 0001
CVPR4
2025 Towards Frame Rate Insensitive Video Object Segmentation
abstract
Video Object Segmentation (VOS) in the scenario of Low Frame Rate Video (LFR-V) is a promising solution for deploying VOS methods on edge devices with limited computation, storage and transmission bandwidth. However, LFR-V leads robustness challenges that cause existing VOS methods to suffer significant performance degradation. Despite their noisy nature on LFR-V, both motion features and appearance features demonstrate remarkable utility in capturing the dynamic information. How to utilize motion and appearance features more effectively is the key to address the challenges of LFR-V. To this end, we propose a Frame Rate Insensitive Video Object Segmentation (FriVOS) model. In detail, we introduce the Interactive Appearance and Motion (IAM) module to explicitly establish the interactions between appearance and motion. Designs of training are proposed for mitigating the gap of training and inference, further capturing the information on LFR-V. Empirically, sufficient experiments show the effectiveness of our methods in training and inference on LFR-V. The proposed methods surpass other outstanding methods across multiple benchmarks 6.2% at most (e.g., DeAOT on DAVIS-2017). More additional experimental analyses validate the robustness of our methods on LFR-V.
Weitu Chong, Kaer Huang, Mingjia Shi
IJCNN3
2025 FedSH: Tackling Staleness By Scheduling High-order Approximation in Asynchronous Federated Learning
abstract
In recent years, Federated Learning (FL) has made significant progress in utilizing decentralized data and protecting data privacy, but communication bottlenecks remain a key challenge. Asynchronous Federated Learning (AFL) addresses this by enhancing training speed through improved asynchronous capabilities. However, AFL primarily faces the issue of gradient staleness and current methods do not consider when approximate estimations are most effective or when they lead to significant errors. To address these issues, we developed FedSH (Federated Scheduled High-order Approximation), which minimizes the large errors in approximate estimates and stabilizes gradient directions, resulting in improved performance in AFL while maintaining similar training times. The experimental results show that the proposed FedSH are commendable on at least 6 benchmarks.
Haixin Gao, Mingjia Shi, Yuhao Zhou 0004, Deng Xiong, Jiancheng Lv 0001
IJCNN2
2025 Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
abstract
Modern Parameter-Efficient Fine-Tuning (PEFT) methods such as low-rank adaptation (LoRA) reduce the cost of customizing large language models (LLMs), yet still require a separate optimization run for every downstream dataset. We introduce \textbf{Drag-and-Drop LLMs (\textit{DnD})}, a prompt-conditioned parameter generator that eliminates per-task training by mapping a handful of unlabeled task prompts directly to LoRA weight updates. A lightweight text encoder distills each prompt batch into condition embeddings, which are then transformed by a cascaded hyper-convolutional decoder into the full set of LoRA matrices. Once trained in a diverse collection of prompt-checkpoint pairs, DnD produces task-specific parameters in seconds, yielding i) up to \textbf{12,000$\times$} lower overhead than full fine-tuning, ii) average gains up to \textbf{30\%} in performance over the strongest training LoRAs on unseen common-sense reasoning, math, coding, and multimodal benchmarks, and iii) robust cross-domain generalization improving \textbf{40\%} performance without access to the target data or labels. Our results demonstrate that prompt-conditioned parameter generation is a viable alternative to gradient-based adaptation for rapidly specializing LLMs. We open source \href{https://jerryliang24.github.io/DnD}{our project} in support of future research.
Zhiyuan Liang, Dongwen Tang, Yuhao Zhou 0004, Xuanlei Zhao, Mingjia Shi, Wangbo Zhao, Peihao Wang, Konstantin Schürholt, Damian Borth, Michael M. Bronstein, Yang You 0001, Zhangyang Wang, Kai Wang 0036
NeurIPS5
2025 Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths
abstract
Mamba has proven efficient for long-sequence modeling in vision tasks. However, when token reduction techniques are applied to improve efficiency, Mamba-based models exhibit drastic performance degradation compared to Vision Transformers (ViTs). This decline is potentially attributed to Mamba's chain-like scanning mechanism, which we hypothesize not only induces cascading losses in token connectivity but also limits the diversity of spatial receptive fields. In this paper, we propose Asymmetric Multi-scale Vision Mamba (AMVim), a novel architecture designed to enhance pruning robustness. AMVim employs a dual-path structure, integrating a window-aware scanning mechanism into one path while retaining sequential scanning in the other. This asymmetry design promotes token connection diversity and enables multi-scale information flow, reinforcing spatial awareness. Empirical results demonstrate that AMVim achieves state-of-the-art pruning robustness. During token reduction, AMVim-T achieves a substantial 34\% improvement in training-free accuracy with identical model sizes and FLOPs. Meanwhile, AMVim-S exhibits only a 1.5\% accuracy drop, performing comparably to ViT. Notably, AMVim also delivers superior performance during pruning-free settings, further validating its architectural advantages.
Jindi Lv, Yuhao Zhou 0004, Mingjia Shi, Zhiyuan Liang, Xiaojiang Peng, Wangbo Zhao, Jiancheng Lv 0001, Kai Wang 0036
NeurIPS3
2025 REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
abstract
Diffusion Transformers (DiTs) deliver state-of-the-art image quality, yet their training remains notoriously slow. A recent remedy---representation alignment (REPA) that matches DiT hidden features to those of a non-generative teacher (e.g., DINO)---dramatically accelerates the early epochs but plateaus or even degrades performance later. We trace this failure to the capacity mismatch: once the generative student begins modeling the joint data distribution, the teacher's lower-dimensional embeddings and attention patterns become a straitjacket rather than a guide. We then introduce HASTE (Holistic Alignment with Stage-wise Termination for Efficient training), a two-phase schedule that keeps the help and drops the hindrance. Phase I applies a holistic alignment loss that simultaneously distills attention maps (relational priors) and feature projections (semantic anchors) from the teacher into mid-level layers of the DiT, yielding rapid convergence. Phase II then performs one-shot termination that deactivates the alignment loss, once a simple trigger such as a fixed iteration is hit, freeing the DiT to focus on denoising and exploit its generative capacity. HASTE speeds up training of diverse DiTs without architecture changes. On ImageNet 256×256, it reaches the vanilla SiT-XL/2 baseline FID in 50 epochs and matches REPA’s best FID in 500 epochs, amounting to a 28× reduction in optimization steps. HASTE also improves text-to-image DiTs on MS-COCO, proving to be a simple yet principled recipe for efficient diffusion training across various tasks.
Wangbo Zhao, Yuhao Zhou 0004, Zhiyuan Liang, Mingjia Shi, Xuanlei Zhao, Kaipeng Zhang, Zhangyang Wang, Kai Wang 0036, Yang You 0001
NeurIPS6
2025 Tackling Feature-Classifier Mismatch in Federated Learning via Prompt-Driven Feature Transformation
abstract
Federated Learning (FL) faces challenges due to data heterogeneity, which limits the global model’s performance across diverse client distributions. Personalized Federated Learning (PFL) addresses this by enabling each client to process an individual model adapted to its local distribution. Many existing methods assume that certain global model parameters are difficult to train effectively in a collaborative manner under heterogeneous data. Consequently, they localize or fine-tune these parameters to obtain personalized models. In this paper, we reveal that both the feature extractor and classifier of the global model are inherently strong, and the primary cause of its suboptimal performance is the mismatch between local features and the global classifier. Although existing methods alleviate this mismatch to some extent and improve performance, we find that they either (1) fail to fully resolve the mismatch while degrading the feature extractor, or (2) address the mismatch only post-training, allowing it to persist during training. This increases inter-client gradient divergence, hinders model aggregation, and ultimately leaves the feature extractor suboptimal for client data. To address this issue, we propose FedPFT, a novel framework that resolves the mismatch during training using personalized prompts. These prompts, along with local features, are processed by a shared self-attention-based transformation module, ensuring alignment with the global classifier. Additionally, this prompt-driven approach offers strong flexibility, enabling task-specific prompts to incorporate additional training objectives (\eg, contrastive learning) to further enhance the feature extractor. Extensive experiments show that FedPFT outperforms state-of-the-art methods by up to 5.07%, with further gains of up to 7.08% when collaborative contrastive learning is incorporated.
Xinghao Wu, Xuefeng Liu 0001, Jianwei Niu 0002, Guogang Zhu, Mingjia Shi, Shaojie Tang 0001, Jing Yuan 0002
NeurIPS5
2025 E-3SFC: Communication-Efficient Federated Learning With Double-Way Features Synthesizing
abstract
The exponential growth in model sizes has significantly increased the communication burden in federated learning (FL). Existing methods to alleviate this burden by transmitting compressed gradients often face high compression errors, which slow down the model's convergence. To simultaneously achieve high compression effectiveness and lower compression errors, we study the gradient compression problem from a novel perspective. Specifically, we propose a systematical algorithm termed extended single-step synthetic features compressing (E-3SFC), which consists of three subcomponents, i.e., the single-step synthetic features compressor (3SFC), a double-way compression (DWC) algorithm, and a communication budget scheduler (BS). First, we regard the process of gradient computation of a model as decompressing gradients from corresponding inputs, while the inverse process is considered as compressing the gradients. Based on this, we introduce a novel gradient compression method termed 3SFC, which utilizes the model itself as a decompressor, leveraging training priors such as model weights and objective functions. The 3SFC compresses raw gradients into tiny synthetic features in a single-step simulation, incorporating error feedback (EF) to minimize overall compression errors. To further reduce communication overhead, 3SFC is extended to E-3SFC, allowing DWC and dynamic communication budget scheduling. Our theoretical analysis under both strongly convex and nonconvex conditions demonstrates that 3SFC achieves linear and sublinear convergence rates with aggregation noise. Extensive experiments across six datasets and six models reveal that 3SFC outperforms the state-of-the-art methods by up to 13.4% while reducing communication costs by 111.6 times. These findings suggest that 3SFC can significantly enhance communication efficiency in FL without compromising model performance.
Yuhao Zhou 0004, Mingjia Shi, Yanan Sun 0001, Jiancheng Lv 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Communication-efficient Federated Learning with Single-Step Synthetic Features Compressor for Faster Convergence
abstract
Reducing communication overhead in federated learning (FL) is challenging but crucial for large-scale distributed privacy-preserving machine learning. While methods utilizing sparsification or other techniques can largely reduce the communication overhead, the convergence rate is also greatly compromised. In this paper, we propose a novel method named Single-Step Synthetic Features Compressor (3SFC) to achieve communication-efficient FL by directly constructing a tiny synthetic dataset containing synthetic features based on raw gradients. Therefore, 3SFC can achieve an extremely low compression rate when the constructed synthetic dataset contains only one data sample. Additionally, the compressing phase of 3SFC utilizes a similarity-based objective function so that it can be optimized with just one step, considerably improving its performance and robustness. To minimize the compressing error, error feedback (EF) is also incorporated into 3SFC. Experiments on multiple datasets and models suggest that 3SFC has significantly better convergence rates compared to competing methods with lower compression rates (i.e., up to 0.02%). Furthermore, ablation studies and visualizations show that 3SFC can carry more information than competing methods for every communication round, further validating its effectiveness.
Yuhao Zhou 0004, Mingjia Shi, Yanan Sun 0001, Jiancheng Lv 0001
ICCV2
2023 Unconstrained Feature Model and Its General Geometric Patterns in Federated Learning: Local Subspace Minority Collapse
Mingjia Shi, Yuhao Zhou 0004, Jiancheng Lv 0001
ICONIP (8)1
2023 PRIOR: Personalized Prior for Reactivating the Information Overlooked in Federated Learning
abstract
Classical federated learning (FL) enables training machine learning models without sharing data for privacy preservation, but heterogeneous data characteristic degrades the performance of the localized model. Personalized FL (PFL) addresses this by synthesizing personalized models from a global model via training on local data. Such a global model may overlook the specific information that the clients have been sampled. In this paper, we propose a novel scheme to inject personalized prior knowledge into the global model in each client, which attempts to mitigate the introduced incomplete information problem in PFL. At the heart of our proposed approach is a framework, the $\textit{PFL with Bregman Divergence}$ (pFedBreD), decoupling the personalized prior from the local objective function regularized by Bregman divergence for greater adaptability in personalized scenarios. We also relax the mirror descent (RMD) to extract the prior explicitly to provide optional strategies. Additionally, our pFedBreD is backed up by a convergence analysis. Sufficient experiments demonstrate that our method reaches the $\textit{state-of-the-art}$ performances on 5 datasets and outperforms other methods by up to 3.5% across 8 benchmarks. Extensive analyses verify the robustness and necessity of proposed designs. The code will be made public.
Mingjia Shi, Yuhao Zhou 0004, Kai Wang 0036, Huaizheng Zhang, Shudong Huang, Jiancheng Lv 0001
NeurIPS1
2022 FLSGD: free local SGD with parallel synchronization
Yuhao Zhou 0004, Mingjia Shi, Jiancheng Lv 0001
J. Supercomput.3
2022 Correction to: FLSGD: free local SGD with parallel synchronization
Yuhao Zhou 0004, Mingjia Shi, Jiancheng Lv 0001
J. Supercomput.3