Tao Feng 0014

dblp:12/4774-14 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0001-5571-8018ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models
abstract
Vision-Language Continual Learning (VLCL) has attracted significant research attention for its robust capabilities, and the adoption of Parameter-Efficient Fine-Tuning (PEFT) strategies is enabling these models to achieve competitive performance with substantially reduced resource consumption. However, dominated First-Order (FO) optimization is prone to trap models in suboptimal local minima, especially in limited exploration subspace within PEFT. To overcome this challenge, this paper pioneers a systematic exploration of adopting Zeroth-Order (ZO) optimization for PEFT-based VLCL. We first identify the incompatibility of naive full-ZO adoption in VLCL due to optimization process instability. We then investigate the application of ZO optimization from a modality branch-wise to a fine-grained layer-wise across various training units to identify an optimal strategy. Besides, a key theoretical insight reveals that vision modality exhibit higher variance than language counterparts in VLCL during the ZO optimization process, and we propose a modality-aware stabilized ZO strategy, which adopts gradient sign normalization in ZO and constrains vision modality perturbation to further improve performance. Benefiting from the adoption of ZO optimization, PEFT-based VLCL fulfills better ability to escape local minima during the optimization process, extensive experiments on four benchmarks demonstrate that our method achieves state-of-the-art results.
Ziwei Liu 0002, Borui Kang, Wei Li 0313, Hangjie Yuan, Yanbing Yang 0001, Yifan Zhu 0001, Tao Feng 0014, Jun Luo 0001
AAAI8
2026 Adapt Before Continual Learning
abstract
Continual Learning (CL) seeks to enable neural networks to incrementally acquire new knowledge (plasticity) while retaining existing knowledge (stability). Although pre-trained models (PTMs) have provided a strong foundation for CL, existing approaches face a fundamental challenge in balancing these two competing objectives. Current methods typically address stability by freezing the PTM backbone, which severely limits the model's plasticity, particularly when incoming data distribution diverges largely from the pre-training data. Alternatively, sequentially fine-tuning the entire PTM can adapt to new knowledge but often leads to catastrophic forgetting, highlighting the critical stability-plasticity trade-off in PTM-based CL. To address this limitation, we propose Adapting PTMs before the core CL process (ACL), a novel framework that introduces a plug-and-play adaptation phase prior to learning each new task. During this phase, ACL refines the PTM backbone by aligning embeddings with their original class prototypes while distancing them from irrelevant classes. This mechanism theoretically and empirically demonstrates desirable balance between stability and plasticity, significantly improving CL performance across benchmarks and integrated methods.
Aojun Lu, Tao Feng 0014, Hangjie Yuan, Chunhui Ding, Yanan Sun 0001
AAAI2
2026 MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems
abstract
Shuhang Chen, Hangjie Yuan, Yunqiu Xu, Pengwei Liu, Tao Feng, Jun Cen, Zeying Huang, Yi Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hangjie Yuan, Yunqiu Xu, Pengwei Liu, Tao Feng 0014, Jun Cen, Zeying Huang, Yi Yang 0001
ACL (1)5
2025 SAMora: Enhancing SAM through Hierarchical Self-Supervised Pre-Training for Medical Images
abstract
The Segment Anything Model (SAM) has demonstrated significant potential in medical image segmentation. Yet, its performance is limited when only a small amount of labeled data is available, while there is abundant valuable yet often overlooked hierarchical information in medical data. To address this limitation, we draw inspiration from self-supervised learning and propose SAMora, an innovative framework that captures hierarchical medical knowledge by applying complementary self-supervised learning objectives at the image, patch, and pixel levels. To fully exploit the complementarity of hierarchical knowledge within LoRAs, we introduce HL-Attn, a hierarchical fusion module that integrates multi-scale features while maintaining their distinct characteristics. SAMora is compatible with various SAM variants, including SAM2, SAMed, and H-SAM. Experimental results on the Synapse, LA, and PROMISE12 datasets demonstrate that SAMora outperforms existing SAM variants. It achieves state-of-the-art performance in both few-shot and fully supervised settings while reducing fine-tuning epochs by 90%. The code is available at https://github.com/ShChen233/SAMora.
Hangjie Yuan, Pengwei Liu, Hanxue Gu, Tao Feng 0014, Dong Ni 0002
ICCV5
2025 ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think
abstract
Backpropagation provides a generalized configuration for overcoming catastrophic forgetting. Optimizers such as SGD and Adam are commonly used for weight updates in continual learning and continual pre-training. However, access to gradient information is not always feasible in practice due to black-box APIs, hardware constraints, or non-differentiable systems, a challenge we refer to as the gradient bans. To bridge this gap, we introduce ZeroFlow, the first benchmark designed to evaluate gradient-free optimization algorithms for overcoming forgetting. ZeroFlow examines a suite of forward pass-based methods across various algorithms, forgetting scenarios, and datasets. Our results show that forward passes alone can be sufficient to mitigate forgetting. We uncover novel optimization principles that highlight the potential of forward pass-based methods in mitigating forgetting, managing task conflicts, and reducing memory demands. Additionally, we propose new enhancements that further improve forgetting resistance using only forward passes. This work provides essential tools and insights to advance the development of forward-pass-based methods for continual learning.
Tao Feng 0014, Didi Zhu, Hangjie Yuan, Wendi Zheng, Jie Tang 0001
ICML1
2025 Rethinking the Stability-Plasticity Trade-off in Continual Learning from an Architectural Perspective
abstract
The quest for Continual Learning (CL) seeks to empower neural networks with the ability to learn and adapt incrementally. Central to this pursuit is addressing the stability-plasticity dilemma, which involves striking a balance between two conflicting objectives: preserving previously learned knowledge and acquiring new knowledge. While numerous CL methods aim to achieve this trade-off, they often overlook the impact of network architecture on stability and plasticity, restricting the trade-off to the parameter level. In this paper, we delve into the conflict between stability and plasticity at the architectural level. We reveal that under an equal parameter constraint, deeper networks exhibit better plasticity, while wider networks are characterized by superior stability. To address this architectural-level dilemma, we introduce a novel framework denoted Dual-Arch, which serves as a plug-in component for CL. This framework leverages the complementary strengths of two distinct and independent networks: one dedicated to plasticity and the other to stability. Each network is designed with a specialized and lightweight architecture, tailored to its respective objective. Extensive experiments demonstrate that Dual-Arch enhances the performance of existing CL methods while being up to 87% more compact in terms of parameters.
Aojun Lu, Hangjie Yuan, Tao Feng 0014, Yanan Sun 0001
ICML3
2025 A Stronger Mixture of Low-Rank Experts for Fine-Tuning Foundation Models
abstract
In order to streamline the fine-tuning of foundation models, Low-Rank Adapters (LoRAs) have been substantially adopted across various fields, including instruction tuning and domain adaptation. The underlying concept of LoRA involves decomposing a full-rank matrix into the product of two lower-rank matrices, which reduces storage consumption and accelerates the training process. Furthermore, to address the limited expressive capacity of LoRA, the Mixture-of-Expert (MoE) has been introduced for incorporating multiple LoRA adapters. The integration of LoRA experts leads to a visible improvement across several downstream scenes. However, the mixture of LoRAs (MoE-LoRA) still exhibits its low robustness during tuning and inferring. Inspired by the Riemannian Preconditioners which train LoRA as a sub-space projector, we propose a new training strategy for MoE-LoRA, to stabilize and boost its feature learning by gate-rescaled multi-space projections. We provide both a theoretical solution as well as an alternative engineering strategy. Examinations on SGD and AdamW optimizers demonstrate the effectiveness of our methodology. Source code is available at https://github.com/THUDM/MoELoRA_Riemannian.
Mengyang Sun, Tao Feng 0014, Yifan Zhu 0001, Jie Tang 0001
ICML3
2025 Multi-label feature selection with feature reconstruction and label correlations
Pengwei Lu, Tao Feng 0014, Guodong Du 0002
Expert Syst. Appl.4
2024 InstructVideo: Instructing Video Diffusion Models with Human Feedback
abstract
Diffusion models have emerged as the de facto paradigm for video generation. However, their reliance on web-scale data of varied quality often yields results that are visually unappealing and misaligned with the textual prompts. To tackle this problem, we propose InstructVideo to instruct text-to-video diffusion models with human feedback by reward fine-tuning. InstructVideo has two key ingredients: 1) To ameliorate the cost of reward fine-tuning induced by generating through the full DDIM sampling chain, we recast reward fine-tuning as editing. By leveraging the diffusion process to corrupt a sampled video, InstructVideo requires only partial inference of the DDIM sampling chain, reducing fine-tuning cost while improving fine-tuning efficiency. 2) To mitigate the absence of a dedicated video reward model for human preferences, we repurpose established image reward models, e.g., HPSv2. To this end, we propose Segmental Video Reward, a mechanism to provide reward signals based on segmental sparse sampling, and Temporally Attenuated Reward, a method that mitigates temporal modeling degradation during fine-tuning. Extensive experiments, both qualitative and quantitative, validate the practicality and efficacy of using image reward models in InstructVideo, significantly enhancing the visual quality of generated videos without compromising generalization capabilities. Code and models can be accessed through our project page https://instructvideo.github.io/.
Hangjie Yuan, Shiwei Zhang 0001, Xiang Wang 0012, Yujie Wei 0001, Tao Feng 0014, Yining Pan, Yingya Zhang, Ziwei Liu 0002, Samuel Albanie, Dong Ni 0002
CVPR5
2024 Revisiting Neural Networks for Continual Learning: An Architectural Perspective
Aojun Lu, Tao Feng 0014, Hangjie Yuan, Xiaotian Song, Yanan Sun 0001
IJCAI2
2024 Make Continual Learning Stronger via C-Flat
abstract
How to balance the learning ’sensitivity-stability’ upon new task training and memory preserving is critical in CL to resolve catastrophic forgetting. Improving model generalization ability within each learning phase is one solution to help CL learning overcome the gap in the joint knowledge space. Zeroth-order loss landscape sharpness-aware minimization is a strong training regime improving model generalization in transfer learning compared with optimizer like SGD. It has also been introduced into CL to improve memory representation or learning efficiency. However, zeroth-order sharpness alone could favors sharper over flatter minima in certain scenarios, leading to a rather sensitive minima rather than a global optima. To further enhance learning stability, we propose a Continual Flatness (C-Flat) method featuring a flatter loss landscape tailored for CL. C-Flat could be easily called with only one line of code and is plug-and-play to any CL methods. A general framework of C-Flat applied to all CL categories and a thorough comparison with loss minima optimizer and flat minima based CL approaches is presented in this paper, showing that our method can boost CL performance in almost all cases. Code is available at https://github.com/WanNaa/C-Flat.
Ang Bian, Wei Li 0313, Hangjie Yuan, Chengrong Yu, Mang Wang 0001, Zixiang Zhao, Aojun Lu, Pengliang Ji, Tao Feng 0014
NeurIPS9
2024 Learning to cluster person via graph convolution networks for video-based person re-identification
abstract
Summary Unsupervised person re‐identification based on video sequences can be applied to surveillance systems and is attracting much more attention. It aims to spot specific person in other scenes captured by different cameras. This work explores an innovative strategy, namely, learning to cluster unlabeled person in the videos through graph convolutional networks. In this article, we find that the possibility of inter‐frame linkage can be inferred from context. Therefore, a pose‐guided topology linkage clustering framework is proposed. Our framework consists of three modules: (i) a pose‐guided representation module; (ii) a pose‐guided embedding module; (iii) a link prediction module. First, the representation coding alone is performed at the level of relational induction bias, embedding the implicit pose structure information in image features. Then, based on the consideration of the topology relationship between adjacent and cross‐frame, graph convolutional network is introduced to infer the likelihood of linkage between frame nodes. Experiments show that the proposed method demonstrates excellent scalability in addition to being an effective response to person clustering in case of changes, and does not need the number of clusters as a prior.
Wei Li 0313, Tao Feng 0014, Guodong Du 0002, Sixin Liang, Ang Bian
Concurr. Comput. Pract. Exp.2
2024 Hospital readmission prediction with hybrid-sampling and self-paced balance learning
abstract
Summary Hospital readmission prediction is defined as an evaluation task to model the historical medical data to predict whether patients will be readmitted after discharge. In the past few years, many feasible and effective prediction methods have been proposed, however, most of them neglect the imbalanced distribution of medical data, which causes great difficulties in modeling. Thus, we proposed a new hospital readmission prediction method, which utilizes hybrid‐sampling and self‐paced balance learning strategies to solve the class‐imbalance problem. To be specifically, we first employ an interference negative sample deletion strategy to reduce the probability of important majority class samples being deleted. Then, we design a hard positive sample generation strategy to generate more positive samples. Meanwhile, we also introduce a self‐paced balance factor during the oversampling process to improve the similarity between newly generated minority class samples and hard positive samples. Finally, we perform the experiments on six real‐world readmission datasets to indicate the superiority of the proposed method.
Tao Feng 0014, Guodong Du 0002
Concurr. Comput. Pract. Exp.4
2024 Revisiting class-incremental object detection: An efficient approach via intrinsic characteristics alignment and task decoupling
Liang Bai 0006, Hong Song 0003, Tao Feng 0014, Tianyu Fu 0003, Qingzhe Yu, Jian Yang 0009
Expert Syst. Appl.3
2024 UniGrad-FS: Unified Gradient Projection With Flatter Sharpness for Continual Learning
abstract
Continual learning (CL) desires that the neural network sequentially perform learning tasks from a dynamic data stream without forgetting learned knowledge. To overcome forgetting, a line of work relies on gradient projection to minimize the influence between gradients during optimization. This article focuses on a challenging problem concerning CL:When, how, and where to implement gradient projection to promote CL.Tackling this problem can be divided into two perspectives, namely the gradient direction (when and how) and the area of gradient conflict (where). First, we propose a plug-and-play method UniGrad to tackle the inconsistency of conflicting and nonconflicting gradients during optimization in CL. Second, we explore the interaction mechanism of gradient projection and loss landscape in CL, and further propose a pluggable method UniGrad-FS to improve the CL performance. In short, this work expects to overcome forgetting through an efficient gradient projection at the area where the gradient conflicts are less intense. In essence, the proposed method is a general and pluggable method that can be used in any gradient-based optimizer. For evaluation, we plug UniGrad and UniGrad-FS into two top-performing baselines (WA and MEMO). Our method shows clear improvements, i.e., boosting WA and MEMO by +2.09% and 1.72% in the 20-step of the CIFAR100 benchmark. In addition, we observe performance enhancement on all settings of CIFAR100 and Tiny-ImageNet datasets. Extensive experiments demonstrate the simplicity and effectiveness of the proposed method.
Wei Li 0313, Tao Feng 0014, Hangjie Yuan, Ang Bian, Guodong Du 0002, Sixin Liang, Jianhong Gan, Ziwei Liu 0002
IEEE Trans. Ind. Informatics2
2023 RLIPv2: Fast Scaling of Relational Language-Image Pre-training
abstract
Relational Language-Image Pre-training (RLIP) aims to align vision representations with relational texts, thereby advancing the capability of relational reasoning in computer vision tasks. However, hindered by the slow convergence of RLIPv11architecture and the limited availability of existing scene graph data, scaling RLIPv1 is challenging. In this paper, we propose RLIPv2, a fast converging model that enables the scaling of relational pre-training to large-scale pseudo-labelled scene graph data. To enable fast scaling, RLIPv2 introduces Asymmetric Language-Image Fusion (ALIF), a mechanism that facilitates earlier and deeper gated cross-modal fusion with sparsified language encoding layers. ALIF leads to comparable or better performance than RLIPv1 in a fraction of the time for pre-training and fine-tuning. To obtain scene graph data at scale, we extend object detection datasets with free-form relation labels by introducing a captioner (e.g., BLIP) and a designed Relation Tagger. The Relation Tagger assigns BLIP-generated relation texts to region pairs, thus enabling larger-scale relational pre-training. Through extensive experiments conducted on Human-Object Interaction Detection and Scene Graph Generation, RLIPv2 shows state-of-the-art performance on three benchmarks under fully-finetuning, few-shot and zero-shot settings. Notably, the largest RLIPv2 achieves 23.29mAP on HICO-DET without any fine-tuning, yields 32.22mAP with just 1% data and yields 45.09mAP with 100% data. Code and models are publicly available at https://github.com/JacobYuan7/RLIPv2.
Hangjie Yuan, Shiwei Zhang 0001, Xiang Wang 0012, Samuel Albanie, Yining Pan, Tao Feng 0014, Jianwen Jiang, Dong Ni 0002, Yingya Zhang, Deli Zhao
ICCV6
2022 Overcoming Catastrophic Forgetting in Incremental Object Detection via Elastic Response Distillation
abstract
Traditional object detectors are ill-equipped for incremental learning. However, fine-tuning directly on a well-trained detection model with only new data will lead to catastrophic forgetting. Knowledge distillation is a flexible way to mitigate catastrophic forgetting. In Incremental Object Detection (IOD), previous work mainly focuses on distilling for the combination of features and responses. However, they under-explore the information that contains in responses. In this paper, we propose a response-based incremental distillation method, dubbed Elastic Response Distillation (ERD), which focuses on elastically learning responses from the classification head and the regression head. Firstly, our method transfers category knowledge while equipping student detector with the ability to retain localization information during incremental learning. In addition, we further evaluate the quality of all locations and provide valuable responses by the Elastic Response Selection (ERS) strategy. Finally, we elucidate that the knowledge from different responses should be assigned with different importance during incremental distillation. Extensive experiments conducted on MS COCO demonstrate our method achieves state-of-the-art result, which substantially narrows the performance gap towards full training. Code is available at https://github.com/Hi-FT/ERD.
Tao Feng 0014, Mang Wang 0001, Hangjie Yuan
CVPR1
2022 RLIP: Relational Language-Image Pre-training for Human-Object Interaction Detection
abstract
The task of Human-Object Interaction (HOI) detection targets fine-grained visual parsing of humans interacting with their environment, enabling a broad range of applications. Prior work has demonstrated the benefits of effective architecture design and integration of relevant cues for more accurate HOI detection. However, the design of an appropriate pre-training strategy for this task remains underexplored by existing approaches. To address this gap, we propose $\textit{Relational Language-Image Pre-training}$ (RLIP), a strategy for contrastive pre-training that leverages both entity and relation descriptions. To make effective use of such pre-training, we make three technical contributions: (1) a new $\textbf{Par}$allel entity detection and $\textbf{Se}$quential relation inference (ParSe) architecture that enables the use of both entity and relation descriptions during holistically optimized pre-training; (2) a synthetic data generation framework, Label Sequence Extension, that expands the scale of language data available within each minibatch; (3) ambiguity-suppression mechanisms, Relation Quality Labels and Relation Pseudo-Labels, to mitigate the influence of ambiguous/noisy samples in the pre-training data. Through extensive experiments, we demonstrate the benefits of these contributions, collectively termed RLIP-ParSe, for improved zero-shot, few-shot and fine-tuning HOI detection performance as well as increased robustness to learning from noisy annotations. Code will be available at https://github.com/JacobYuan7/RLIP.
Hangjie Yuan, Jianwen Jiang, Samuel Albanie, Tao Feng 0014, Ziyuan Huang 0003, Dong Ni 0002, Mingqian Tang
NeurIPS4