EDBT 2026 Demo / reviewers in the wild / expert
Yizhao Gao 0004
dblp:132/7629-4
· DBLP profile ↗
12ranked-venue papers
4as first author
11since 2021 · last 2024
0000-0002-6903-2962ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Unsupervised Continual Learning of Image Representation Via Rememory-Based SimsiamabstractUnsupervised continual learning (UCL) of image representation has garnered attention due to practical need. However, recent UCL methods focus on mitigating the catastrophic forgetting with a replay buffer (i.e., rehearsal-based strategy), which needs much extra storage. To overcome this drawback, we propose a novel rememory-based SimSiam (RM-SimSiam) method to reduce the dependency on replay buffer. The core idea of RM-SimSiam is to store and remember the old knowledge with a data-free historical module instead of replay buffer. Specifically, this historical module is designed to store the historical average model of all previous models (the memory process) and then transfer the knowledge of the historical average model to the new model (the rememory process). To further improve the rememory ability of RMSimSiam, we devise an enhanced SimSiam-based contrastive loss by aligning the representations outputted by the historical and new models. Extensive experiments on three benchmarks demonstrate the effectiveness of our RM-SimSiam. Feifei Fu, Yizhao Gao 0004, Zhiwu Lu 0001 |
ICASSP | 2 |
| 2024 | Enhancing Class-Incremental Learning for Image Classification via Bidirectional Transport and Selective MomentumabstractClass-Incremental Learning (Class-IL) aims to continuously learn new knowledge without forgetting old knowledge from a given data stream in the realm of image classification. Recent Class-IL methods strive to balance old and new knowledge and have achieved excellent results in mitigating the forgetting by mainly employing the rehearsal-based strategy. However, the representation learning on new tasks is often impaired since the trade-off is hard to taken between old and new knowledge. To overcome this challenge, based on the Complementary Learning System (CLS) theory, we propose a novel CLS-based method by focusing on the representation of old and new knowledge under the Class-IL setting, which can acquire more new knowledge from new tasks while consolidating the old knowledge so as to make a better balance between them (i.e., enhancing the overall model performance). Specifically, our proposed method has two novel components: (1) To effectively mitigate the forgetting, we first propose a bidirectional transport (BDT) strategy between old and new models, which can better integrate the old knowledge into the new knowledge and meanwhile enforce the old knowledge to be better consolidated by bidirectionally transferring parameters across old and new models. (2) To ensure that the representation of new knowledge is not impaired by the old knowledge, we further devise a selective momentum (SMT) mechanism to give parameters greater flexibility to learn new knowledge while transferring important old knowledge, which is achieved by selectively (momentum) updating network parameters through parameter importance evaluation. Extensive experiments on five benchmarks show that our proposed method significantly outperforms the state-of-the-arts under the Class-IL setting. Feifei Fu, Yizhao Gao 0004, Zhiwu Lu 0001 |
ICMR | 2 |
| 2023 | R-STAR: Robust Self-Taught Task-Wise Reweighting for Rehearsal-Based Class Incremental LearningabstractClass incremental learning (CIL) requires a model to learn the knowledge of new classes without overwriting that of old classes. The main challenge thus lies in catastrophic forgetting. Among all advances in addressing this challenge, rehearsal-based methods are the most widely-used due to their convenience and effectiveness. However, the (classification) scores bias between the old and new classes, known as the main cause of catastrophic forgetting for rehearsal-based methods, is still not fully addressed. Although some recent strategies are proposed to reduce the scores bias, they either take extra training time or sacrifice too much performance on the current task. In this paper, we propose a novel Robust Self-Taught Task-Wise Reweighting (R-STAR) method, which can act as a flexible and key component for improving existing rehearsal-based methods. Concretely, on top of the standard training process, it measures the forgetting degree of the model over the augmented buffer (for robust evaluation) on each task. Further, following the self-taught paradigm, it directly activates the task-wise forgetting degree into a reweighting ratio for scores bias reduction during the inference stage. Extensive experiments show that our R-STAR can improve most rehearsal-based methods with remarkable margins, but with (almost) no extra training cost or excessive performance sacrifice on the new task. Moreover, it also shows its advantages over existing scores bias correction strategies. Yutian Luo, Yizhao Gao 0004, Ruitao Ma, Zhiwu Lu 0001 |
ECAI | 2 |
| 2023 | Mixup-Inspired Video Class-Incremental LearningabstractContinual learning aims to learn a sequence of tasks without forgetting the previously learned knowledge. Although existing memory-based approaches can be easily deployed for video Class-Incremental Learning (CIL), little efforts have been made to explore how to better exploit the data from the previous work (in the memory) for alleviating the catastrophic forgetting. In this work, we thus propose a simple yet effective framework called Mixup-Inspired Video Class-Incremental Learning (MIV-CIL). The core idea of our MIVCIL framework is to impose mixup on the current video data and the previous video data (from the memory buffer) to mitigate the catastrophic forgetting. By exploring different mixup strategies on the video data, our MIVCIL framework has three instantiations for video class-incremental learning. We further provide a detailed analysis of the performance and computational overhead of the three instantiations on the latest benchmark vCLIMB. Experimental results show that all three instantiations achieve significant improvements over the representative/state-of-the-art methods. Jinqiang Long, Yizhao Gao 0004, Zhiwu Lu 0001 |
ICDM | 2 |
| 2023 | CMMT: Cross-Modal Meta-Transformer for Video-Text RetrievalabstractVideo-text retrieval has drawn great attention due to the prosperity of online video contents. Most existing methods extract the video embeddings by densely sampling abundant (generally dozens of) video clips, which acquires tremendous computational cost. To reduce the resource consumption, recent works propose to sparsely sample fewer clips from each raw video with a narrow time span. However, they still struggle to learn a reliable video representation with such locally sampled video clips, especially when testing on cross-dataset setting. In this work, to overcome this problem, we sparsely and globally (with wide time span) sample a handful of video clips from each raw video, which can be regarded as different samples of a pseudo video class (i.e., each raw video denotes a pseudo video class). From such viewpoint, we propose a novel Cross-Modal Meta-Transformer (CMMT) model that can be trained in a meta-learning paradigm. Concretely, in each training step, we conduct a cross-modal fine-grained classification task where the text queries are classified with pseudo video class prototypes (each has aggregated all sampled video clips per pseudo video class). Since each classification task is defined with different/new videos (by simulating the evaluation setting), this task-based meta-learning process enables our model to generalize well on new tasks and thus learn generalizable video/text representations. To further enhance the generalizability of our model, we induce a token-aware adaptive Transformer module to dynamically update our model (prototypes) for each individual text query. Extensive experiments on three benchmarks show that our model achieves new state-of-the-art results in cross-dataset video-text retrieval, demonstrating that it has more generalizability in video-text retrieval. Importantly, we find that our new meta-learning paradigm indeed brings improvements under both cross-dataset and in-dataset retrieval settings. Yizhao Gao 0004, Zhiwu Lu 0001 |
ICMR | 1 |
| 2023 | Learning with Adaptive Knowledge for Continual Image-Text ModelingabstractIn realistic application scenarios, existing methods for image-text modeling have limitations in dealing with data stream: training on all data needs too much computation/storage resources, and even the full access to previous data is invalid. In this work, we thus propose a new continual image-text modeling (CITM) setting that requires a model to be trained sequentially on a number of diverse image-text datasets. Although recent continual learning methods can be directly applied to the CITM setting, most of them only consider reusing part of previous data or aligning the output distributions of previous and new models, which is a partial or indirect way to acquire the old knowledge. In contrast, we propose a novel dynamic historical adaptation (DHA) method which can holistically and directly review the old knowledge from a historical model. Concretely, the historical model transfers its total parameters to the main/current model to utilize the holistic old knowledge. In turn, the main model dynamically transfers its parameters to the historical model at every five training steps to ensure that the knowledge gap between them is not too large. Extensive experiments show that our proposed DHA outperforms other representative/latest continual learning methods under the CITM setting. Yutian Luo, Yizhao Gao 0004, Zhiwu Lu 0001 |
ICMR | 2 |
| 2022 | SST-VLM: Sparse Sampling-Twice Inspired Video-Language Model
Yizhao Gao 0004, Zhiwu Lu 0001 |
ACCV (4) | 1 |
| 2022 | COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal RetrievalabstractLarge-scale single-stream pre-training has shown dramatic performance in image-text retrieval. Regrettably, it faces low inference efficiency due to heavy attention layers. Recently, two-stream methods like CLIP and ALIGN with high inference efficiency have also shown promising performance, however, they only consider instance-level alignment between the two streams (thus there is still room for improvement). To overcome these limitations, we propose a novel COllaborative Two-Stream vision-language pretraining model termed COTS for image-text retrieval by enhancing cross-modal interaction. In addition to instance-level alignment via momentum contrastive learning, we leverage two extra levels of cross-modal interactions in our COTS: (1) Token-level interaction - a masked vision-language modeling (MVLM) learning objective is devised without using a cross-stream network module, where variational autoencoder is imposed on the visual encoder to generate visual tokens for each image. (2) Task-level interaction - a KL-alignment learning objective is devised between text-to-image and image-to-text retrieval tasks, where the probability distribution per task is computed with the negative queues in momentum contrastive learning. Under a fair comparison setting, our COTS achieves the highest performance among all two-stream methods and comparable performance (but with 10,800× faster in inference) w.r.t. the latest single-stream methods. Importantly, our COTS is also applicable to text-to-video retrieval, yielding new state-of-the-art on the widely-used MSR-VTT dataset. Haoyu Lu, Nanyi Fei, Yuqi Huo, Yizhao Gao 0004, Zhiwu Lu 0001, Ji-Rong Wen |
CVPR | 4 |
| 2022 | BMU-MoCo: Bidirectional Momentum Update for Continual Video-Language ModelingabstractVideo-language models suffer from forgetting old/learned knowledge when trained with streaming data. In this work, we thus propose a continual video-language modeling (CVLM) setting, where models are supposed to be sequentially trained on five widely-used video-text datasets with different data distributions. Although most of existing continual learning methods have achieved great success by exploiting extra information (e.g., memory data of past tasks) or dynamically extended networks, they cause enormous resource consumption when transferred to our CVLM setting. To overcome the challenges (i.e., catastrophic forgetting and heavy resource consumption) in CVLM, we propose a novel cross-modal MoCo-based model with bidirectional momentum update (BMU), termed BMU-MoCo. Concretely, our BMU-MoCo has two core designs: (1) Different from the conventional MoCo, we apply the momentum update to not only momentum encoders but also encoders (i.e., bidirectional) at each training step, which enables the model to review the learned knowledge retained in the momentum encoders. (2) To further enhance our BMU-MoCo by utilizing earlier knowledge, we additionally maintain a pair of global momentum encoders (only initialized at the very beginning) with the same BMU strategy. Extensive results show that our BMU-MoCo remarkably outperforms recent competitors w.r.t. video-text retrieval performance and forgetting rate, even without using any extra data or dynamic networks. Yizhao Gao 0004, Nanyi Fei, Haoyu Lu, Zhiwu Lu 0001, Hao Jiang 0022, Zhao Cao |
NeurIPS | 1 |
| 2021 | Z-Score Normalization, Hubness, and Few-Shot LearningabstractThe goal of few-shot learning (FSL) is to recognize a set of novel classes with only few labeled samples by exploiting a large set of abundant base class samples. Adopting a meta-learning framework, most recent FSL methods meta-learn a deep feature embedding network, and during inference classify novel class samples using nearest neighbor in the learned high-dimensional embedding space. This means that these methods are prone to the hubness problem, that is, a certain class prototype becomes the nearest neighbor of many test instances regardless which classes they belong to. However, this problem is largely ignored in existing FSL studies. In this work, for the first time we show that many FSL methods indeed suffer from the hubness problem. To mitigate its negative effects, we further propose to employ z-score feature normalization, a simple yet effective trans-formation, during meta-training. A theoretical analysis is provided on why it helps. Extensive experiments are then conducted to show that with z-score normalization, the performance of many recent FSL methods can be boosted, resulting in new state-of-the-art on three benchmarks. Nanyi Fei, Yizhao Gao 0004, Zhiwu Lu 0001, Tao Xiang 0002 |
ICCV | 2 |
| 2021 | Contrastive prototype learning with augmented embeddings for few-shot learningabstractMost recent few-shot learning (FSL) methods are based on meta-learning with episodic training. In each meta-training episode, a discriminative feature embedding and/or classifier are first constructed from a support set in an inner loop, and then evaluated in an outer loop using a query set for model updating. This query set sample centered learning objective is however intrinsically limited in addressing the lack of training data problem in the support set. In this paper, a novel contrastive prototype learning with augmented embeddings (CPLAE) model is proposed to overcome this limitation. First, data augmentations are introduced to both the support and query sets with each sample now being represented as an augmented embedding (AE) composed of concatenated embeddings of both the original and augmented versions. Second, a novel support set class prototype centered contrastive loss is proposed for contrastive prototype learning (CPL). With a class prototype as an anchor, CPL aims to pull the query samples of the same class closer and those of different classes further away. This support set sample centered loss is highly complementary to the existing query centered loss, fully exploiting the limited training data in each episode. Extensive experiments on several benchmarks demonstrate that our proposed CPLAE achieves new state-of-the-art. Yizhao Gao 0004, Nanyi Fei, Guangzhen Liu, Zhiwu Lu 0001, Tao Xiang 0002 |
UAI | 1 |
| 2020 | Few-Shot Zero-Shot Learning: Knowledge Transfer with Less Supervision
Nanyi Fei, Jiechao Guan, Zhiwu Lu 0001, Yizhao Gao 0004 |
ACCV (3) | 4 |