Jin Tang 0005

dblp:56/4951-5 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
9since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021
YearPublicationVenuePosition
2024 Visually Guided Audio Source Separation with Meta Consistency Learning
abstract
In this paper, we tackle the problem of visually guided audio source separation in the context of both known and unknown objects (e.g., musical instruments). Recent successful end-to-end deep learning approaches adopt a single network with fixed parameters to generalize across unseen test videos. However, it can be challenging to generalize in cases where the distribution shift between training and test videos is higher as they fail to utilize internal information of unknown test videos. Based on this observation, we introduce a meta-consistency driven test time adaptation scheme that enables the pretrained model to quickly adapt to known and unknown test music videos in order to bring substantial improvements. In particular, we design a self-supervised audio-visual consistency objective as an auxiliary task that learns the synchronization between audio and its corresponding visual embedding. Concretely, we apply a meta-consistency training scheme to further optimize the pretrained model for effective and faster test time adaptation. We obtain substantial performance gains with only a smaller number of gradient updates and without any additional parameters for the task of audio source separation. Extensive experimental results across datasets demonstrate the effectiveness of our proposed method.
Md. Amirul Islam, Seyed Shahabeddin Nabavi, Irina Kezele, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005
WACV6
2023 Meta-Auxiliary Learning for Future Depth Prediction in Videos
abstract
We consider a new problem of future depth prediction in videos. Given a sequence of observed frames in a video, the goal is to predict the depth map of a future frame that has not been observed yet. Depth estimation plays a vital role for scene understanding and decision-making in intelligent systems. Predicting future depth maps can be valuable for autonomous vehicles to anticipate the behaviours of their surrounding objects. Our proposed model for this problem has a two-branch architecture. One branch is for the primary task of future depth prediction. The other branch is for an auxiliary task of image reconstruction. The auxiliary branch can act as a regularization. Inspired by some recent work on test-time adaption, we use the auxiliary task during testing to adapt the model to a specific test video. We also propose a novel meta-auxiliary learning that learns the model specifically for the purpose of effective test-time adaptation. Experimental results demonstrate that our proposed approach outperforms other alternative methods.
Huan Liu 0014, Zhixiang Chi, Yuanhao Yu, Yang Wang 0003, Jun Chen 0005, Jin Tang 0005
WACV6
2023 Test-Time Adaptation for Optical Flow Estimation Using Motion Vectors
abstract
Due to the prohibitive cost as well as technical challenges in annotating ground-truth optical flow for large-scale realistic video datasets, the existing deep learning models for optical flow estimation mostly rely on synthetic data for training, which in turn may lead to significant performance degradation under test-data distribution shift in real-world environments. In this work, we propose the methodology to tackle this important problem. We design a self-supervised learning task for adjusting the optical flow estimation model at test time. We exploit the fact that most videos are stored in compressed formats, from which compact information on motion, in the form of motion vectors and residuals, can be made readily available. We formulate the self-supervised task as motion vector prediction, and link this task to optical flow estimation. To the best of our knowledge, our Test-Time Adaption guided with Motion Vectors (TTA-MV), is the first work to perform such adaptation for optical flow. The experimental results demonstrate that TTA-MV can improve the generalization capability of various well-known deep learning methods for optical flow estimation, such as FlowNet, PWCNet, and RAFT.
Seyed Mehdi Ayyoubzadeh, Irina Kezele, Yuanhao Yu, Xiaolin Wu 0001, Yang Wang 0003, Jin Tang 0005
IEEE Trans. Image Process.7
2023 Fast Human Pose Estimation in Compressed Videos
abstract
Current approaches for human pose estimation in videos can be categorized into per-frame and warping-based methods. Both approaches have their pros and cons. For example, per-frame methods are generally more accurate, but they are often slow. Warping-based approaches are more efficient, but the performance is usually not good. To bridge the gap, in this paper, we propose a novel fast framework for human pose estimation to meet the real-time inference with controllable accuracy degradation in compressed video domain. Our approach takes advantage of the motion representation (called “motion vector”) that is readily available in a compressed video. Pose joints in a frame are obtained by directly warping the pose joints from the previous frame using the motion vectors. We also propose modules to correct possible errors introduced by the pose warping when needed. Extensive experimental results demonstrate the effectiveness of our proposed framework for accelerating the speed of top-down human pose estimation in videos.
Huan Liu 0014, Zhixiang Chi, Yang Wang 0003, Yuanhao Yu, Jun Chen 0005, Jin Tang 0005
IEEE Trans. Multim.7
2022 MetaFSCIL: A Meta-Learning Approach for Few-Shot Class Incremental Learning
abstract
In this paper, we tackle the problem of few-shot class incremental learning (FSCIL). FSCIL aims to incrementally learn new classes with only a few samples in each class. Most existing methods only consider the incremental steps at test time. The learning objective of these methods is often hand-engineered and is not directly tied to the objective (i.e. incrementally learning new classes) during testing. Those methods are sub-optimal due to the misalignment between the training objectives and what the methods are expected to do during evaluation. In this work, we proposed a bi-level optimization based on meta-learning to directly optimize the network to learn how to incrementally learn in the setting of FSCIL. Concretely, we propose to sample sequences of incremental tasks from base classes for training to simulate the evaluation protocol. For each task, the model is learned using a meta-objective such that it is capable to perform fast adaptation without forgetting. Furthermore, we propose a bi-directional guided modulation, which is learned to automatically modulate the activations to reduce catastrophic forgetting. Extensive experimental results demonstrate that the proposed method outperforms the baseline and achieves the state-of-the-art results on CIFARIOO, MiniImageNet, and CUB200 datasets.
Zhixiang Chi, Li Gu, Huan Liu 0014, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005
CVPR6
2022 Few-Shot Class-Incremental Learning via Entropy-Regularized Data-Free Replay
Huan Liu 0014, Li Gu, Zhixiang Chi, Yang Wang 0003, Yuanhao Yu, Jun Chen 0005, Jin Tang 0005
ECCV (24)7
2022 Meta-DMoE: Adapting to Domain Shift by Meta-Distillation from Mixture-of-Experts
abstract
In this paper, we tackle the problem of domain shift. Most existing methods perform training on multiple source domains using a single model, and the same trained model is used on all unseen target domains. Such solutions are sub-optimal as each target domain exhibits its own specialty, which is not adapted. Furthermore, expecting single-model training to learn extensive knowledge from multiple source domains is counterintuitive. The model is more biased toward learning only domain-invariant features and may result in negative knowledge transfer. In this work, we propose a novel framework for unsupervised test-time adaptation, which is formulated as a knowledge distillation process to address domain shift. Specifically, we incorporate Mixture-of-Experts (MoE) as teachers, where each expert is separately trained on different source domains to maximize their specialty. Given a test-time target domain, a small set of unlabeled data is sampled to query the knowledge from MoE. As the source domains are correlated to the target domains, a transformer-based aggregator then combines the domain knowledge by examining the interconnection among them. The output is treated as a supervision signal to adapt a student prediction network toward the target domain. We further employ meta-learning to enforce the aggregator to distill positive knowledge and the student network to achieve fast adaptation. Extensive experiments demonstrate that the proposed method outperforms the state-of-the-art and validates the effectiveness of each proposed component. Our code is available at https://github.com/n3il666/Meta-DMoE.
Tao Zhong 0003, Zhixiang Chi, Li Gu, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005
NeurIPS6
2021 Test-Time Fast Adaptation for Dynamic Scene Deblurring via Meta-Auxiliary Learning
abstract
In this paper, we tackle the problem of dynamic scene deblurring. Most existing deep end-to-end learning approaches adopt the same generic model for all unseen test images. These solutions are sub-optimal, as they fail to utilize the internal information within a specific image. On the other hand, a self-supervised approach, SelfDeblur, enables internal training within a test image from scratch, but it does not fully take advantage of large external datasets. In this work, we propose a novel self-supervised meta-auxiliary learning to improve the performance of deblurring by integrating both external and internal learning. Concretely, we build a self-supervised auxiliary reconstruction task that shares a portion of the network with the primary deblurring task. The two tasks are jointly trained on an external dataset. Furthermore, we propose a meta-auxiliary training scheme to further optimize the pretrained model as a base learner, which is applicable for fast adaptation at test time. During training, the performance of both tasks is coupled. Therefore, we are able to exploit the internal information at test time via the auxiliary task to enhance the performance of deblurring. Extensive experimental results across evaluation datasets demonstrate the effectiveness of test-time adaptation of the proposed method.
Zhixiang Chi, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005
CVPR4
2021 Toward Personalized Emotion Recognition: A Face Recognition Based Attention Method for Facial Emotion Recognition
abstract
This paper aims to address the subject-dependent challenge of the facial emotion recognition (FER) task. To accomplish this, we propose a novel face recognition based attention FER (FRA-FER) framework which propagates subtle face recognition (FR) features through the FER network. Particularly, first a spatial attention map from the feature maps of an FR convolutional neural network (CNN) is created and then it is fused into the FER-CNN. By doing this FR feature propagation, the FER network is personalized as it takes the advantage of the FR features learned from large-scale face recognition datasets. Experiments on the two challenging datasets AffectNet and AFEW demonstrate the superiority of our proposed FRA-FER network to the state-of-the-art work.
Mostafa Shahabinejad, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005
FG4
2020 All at Once: Temporally Adaptive Multi-frame Interpolation with Advanced Motion Modeling
Zhixiang Chi, Rasoul Mohammadi Nasiri, Juwei Lu, Jin Tang 0005, Konstantinos N. Plataniotis
ECCV (27)5