EDBT 2026 Demo / reviewers in the wild / expert
Tao Jin 0004
dblp:88/4850-4
· DBLP profile ↗
68ranked-venue papers
11as first author
65since 2021 · last 2026
0000-0003-3564-1628ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 7 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 5 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scene-Aware Spatiotemporal Generalization: Towards Robust Temporal Action Detection Across DomainsabstractTemporal Action Detection (TAD) aims to identify specific actions in long, untrimmed videos by determining their start, end times and categories, yet existing models suffer from performance degradation under out-of-distribution scenarios due to unrealistic i.i.d. assumptions. While domain generalization (DG) offers a promising solution, image-based DG methods fail to address the unique spatiotemporal challenges in video-based TAD, including the spatiotemporal complexities and significant variations in action instance scales and densities across domains. To bridge this gap, we propose the first DG framework tailored for TAD. We propose Scene-Aware Video Segmentation, which segments videos based on semantic similarity, addressing cross-domain action instance density and scale discrepancies. Additionally, we present Temporal-Aware Normalization Perturbation to generate diverse video features while preserving temporal integrity. We establish the first DG-TAD benchmark, evaluating 11 state-of-the-art DG methods across four datasets. The experiments demonstrate that our framework consistently outperforms existing approaches, achieving superior generalization on unseen domains. The proposed modules are architecture-agnostic, offering plug-and-play compatibility for broader video understanding tasks. Fangming Feng, Sihang Cai, Zequn Xie, Tao Jin 0004 |
AAAI | 5 |
| 2026 | Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-SpeechabstractWhile diffusion and flow-matching models have advanced TTS, generating high-arousal emotions remains a persistent challenge due to the trade-off between stability and expressiveness.Existing systems often suffer from linguistic collapse when pursuing high intensity or fail to meet target emotional levels under stable settings.In this work, we identify that standard Gaussian initialization inevitably introduces a neutral prosody bias, while uniform Classifier-Free Guidance often distorts the acoustic manifold, leading to artifacts.To address this, we propose an inference framework that rectifies the emotional trajectory.An Emotion-Rectified Noise Prior injects a semantic gradient at initialization to align sampling with the target emotional manifold, and Likelihood-Inverse Guidance adaptively schedules guidance via a conditional/unconditional likelihood ratio, strengthening guidance only when the trajectory drifts toward a neutral fallback.Extensive experiments demonstrate that our method effectively resolves the stability bottleneck in high-intensity scenarios, achieving superior linguistic accuracy and emotional fidelity without model retraining.Code is available at https://github.com/MM-Speech/emo-tts. Fangming Feng, Zequn Xie, Yu Zhang 0126, Zhou Zhao 0001, Tao Jin 0004 |
ACL (1) | 7 |
| 2026 | DPDV: Dual-Pathway and Dual-View Representation Learning for Bridging Information Asymmetry in Text-Video RetrievalabstractText-based person anomaly search retrieves specific behavioral events from surveillance archives using natural-language queries.Although recent pose-aware methods align geometric structures well, they face a fundamental Pose-Semantic Gap: semantically different actions can share similar skeletal geometries.While Multimodal Large Language Models (MLLMs) can reduce this ambiguity, using them for large-scale retrieval is computationally prohibitive.We propose the Structure-Semantic Decoupled Cascade (SSDC) framework, which decouples retrieval into two stages:(1) Structure-Aware Coarse Retrieval, where a lightweight model quickly filters candidates by skeletal similarity; and (2) Detective Squad Interaction, a multi-agent semantic verification module.The squad consists of a Detective for fast binary filtering, an Analyst for evidence extraction, and a Writer for semantic synthesis.Finally, we re-rank candidates by fusing the synthesized captions with structural priors.Experiments on the PAB benchmark show that SSDC achieves state-of-theart performance by balancing efficiency and semantic reasoning.Our code is available Zequn Xie, Fangming Feng, Boyun Zhang, Tao Jin 0004 |
ACL (1) | 5 |
| 2026 | SAME: Signer-Aware Mixture-of-Experts for Test-Time Adaptation in Sign Language TranslationabstractLujia Yang, Weicai Yan, Yongbo He, Qifei Zhang, Tao Jin, Jinshan Zhang, Meng Xi, Jianwei Yin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Lujia Yang, Weicai Yan, Yongbo He, Qifei Zhang 0001, Tao Jin 0004, Jinshan Zhang 0001, Meng Xi 0002, Jianwei Yin |
ACL (1) | 5 |
| 2026 | TAG: Triple Alignment With Rationale Generation for Knowledge-Based Visual Question AnsweringabstractKnowledge-based Visual Question Answering (VQA) involves answering questions based not only on the given image, but also on external knowledge. Existing methods for knowledge-based VQA can be classified into two main categories: those that rely on external knowledge bases, and those that use Large Language Models (LLMs) as implicit knowledge engines. However, the former approach heavily relies on the quality of information retrieval, introducing additional information bias to the entire system. And the latter approach suffers from the extremely high computational cost and the loss of image information. To address these issues, we propose a novel framework called TAG that reformulates knowledge-based VQA as a contrastive learning problem. We innovatively propose a triple asymmetric paradigm, which aligns a lightweight text encoder to the image space with an extremely low training cost (0.0152B trainable parameters), and enhance its understanding ability on semantic granularity. TAG is both computation-efficient and effective, and we evaluate it on the knowledge-based VQA datasets, A-OKVQA, OK-VQA and VCR. The results show that TAG (0.387B) achieves the state-of-the-art performance when compared to methods using less than 1B parameters. Besides, TAG still shows competitive performance when compared to methods with LLM. Sihang Cai, Xuan Lin, Jingtong Wu, Tao Jin 0004, Zhou Zhao 0001, Fei Wu 0001, Jun Yu 0002 |
IEEE Trans. Big Data | 5 |
| 2026 | Emphasizing Domain Differences Through Interactive-Augmented Prompts in Continual Audio-Visual Speech RecognitionabstractAudio-Visual Speech Recognition (AVSR) has been studied for a long time in the literature. By leveraging the complementary information from both acoustic and visual modalities, this approach offers a promising solution for robust speech transcription. While recent AVSR models have achieved impressive performance on large-scale, uniformly distributed datasets, they often overlook the challenges posed by real-world scenarios-where data is collected across multiple sessions and environments, leading to significant domain shifts and heterogeneous distributions. Such heterogeneity can result in catastrophic forgetting and hinder the generalization ability of the conventional models. To bridge this gap, we introduce the Continual Audio-Visual Speech Recognition (CL-AVSR) problem, which formulates AVSR as a continual learning task. We establish a dedicated benchmark for CL-AVSR by designing three experimental scenarios that reflect real-world challenges: introducing varying background noise for the audio stream, degrading video quality for the visual stream, and dividing tasks by speaker characteristics to jointly affect both modalities. These scenarios systematically evaluate the model's ability to adapt and retain knowledge across dynamic and non-stationary data streams. To address the unique challenges of CL-AVSR, we propose the Interaction-enhanced Multimodal Prompt learning (IMP) framework. IMP builds upon a pre-trained AV-HuBERT backbone and integrates task-relevant soft prompts with cross-modal and cross-task interactions, enabling efficient knowledge transfer from high-quality source domains to typical low-quality target domains with minimal parameter overhead. The interactive prompts facilitate fine-grained alignment and adaptation between modalities and tasks, while contrastive regularization further mitigates catastrophic forgetting. Furthermore, we devise a multi-modal prompt selection strategy that leverages clustering-based feature analysis, empowering the model to dynamically select optimal prompts for unseen data distributions during inference. Extensive experiments on the LRS2 dataset demonstrate that IMP achieves substantial improvements over strong baselines, setting new state-of-the-art performance in all CL-AVSR scenarios. Our results highlight the effectiveness of IMP in enhancing continual learning capabilities for AVSR, paving the way for more robust and adaptable multi-modal speech recognition systems in real-world applications. Xize Cheng, Jingyuan Chen 0003, Tao Jin 0004, Zhongfei Zhang |
IEEE Trans. Image Process. | 4 |
| 2025 | Bridging the Gap for Test-Time Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) is an emerging research topic that aims to understand and recognize human sentiment or emotions through multiple modalities. However, in real-world dynamic scenarios, the distribution of target data is always changing and different from the source data used to train the model, which leads to performance degradation. Common adaptation methods usually need source data, which could pose privacy issues or storage overheads. Therefore, test-time adaptation (TTA) methods are introduced to improve the performance of the model at inference time. Existing TTA methods are always based on probabilistic models and unimodal learning, and thus can not be applied to MSA which is often considered as a multimodal regression task. In this paper, we propose two strategies: Contrastive Adaptation and Stable Pseudo-label generation (CASP) for test-time adaptation for multimodal sentiment analysis. The two strategies deal with the distribution shifts for MSA by enforcing consistency and minimizing empirical risk, respectively. Extensive experiments show that CASP brings significant and consistent improvements to the performance of the model across various distribution shift settings and with different backbones, demonstrating its effectiveness and versatility. Zirun Guo, Tao Jin 0004 |
AAAI | 2 |
| 2025 | A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal AdapterabstractEfficient transfer learning methods such as adapter-based methods have shown great success in unimodal models and vision-language models. However, existing methods have two main challenges in fine-tuning multimodal models. Firstly, they are designed for vision-language tasks and fail to extend to situations where there are more than two modalities. Secondly, they exhibit limited exploitation of interactions between modalities and lack efficiency. To address these issues, in this paper, we propose the loW-rank sequence multimodal adapter (Wander). We first use the outer product to fuse the information from different modalities in an element-wise way effectively. For efficiency, we use CP decomposition to factorize tensors into rank-one components and achieve substantial parameter reduction. Furthermore, we implement a token-level low-rank decomposition to extract more fine-grained features and sequence relationships between modalities. With these designs, Wander enables token-level interactions between sequences of different modalities in a parameter-efficient way. We conduct extensive experiments on datasets with different numbers of modalities, where Wander outperforms state-of-the-art efficient transfer learning methods consistently. The results fully demonstrate the effectiveness, efficiency and universality of Wander. Zirun Guo, Xize Cheng, Tao Jin 0004 |
AAAI | 4 |
| 2025 | Speech Watermarking with Discrete Intermediate RepresentationsabstractSpeech watermarking techniques can proactively mitigate the potential harmful consequences of instant voice cloning techniques. These techniques involve the insertion of signals into speech that are imperceptible to humans but can be detected by algorithms. Previous approaches typically embed watermark messages into continuous space. However, intuitively, embedding watermark information into robust discrete latent space can significantly improve the robustness of watermarking systems. In this paper, we propose DiscreteWM, a novel speech watermarking framework that injects watermarks into the discrete intermediate representations of speech. Specifically, we map speech into discrete latent space with a vector-quantized autoencoder and inject watermarks by changing the modular arithmetic relation of discrete IDs. To ensure the imperceptibility of watermarks, we also propose a manipulator model to select the candidate tokens for watermark embedding. Experimental results demonstrate that our framework achieves state-of-the-art performance in robustness and imperceptibility, simultaneously. Moreover, our flexible frame-wise approach can serve as an efficient solution for both voice cloning detection and information hiding. Additionally, DiscreteWM can encode 1 to 150 bits of watermark information within a 1-second speech clip, indicating its encoding capacity. Shengpeng Ji, Ziyue Jiang 0001, Jialong Zuo, Minghui Fang 0002, Tao Jin 0004, Zhou Zhao 0001 |
AAAI | 6 |
| 2025 | T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI FeedbackabstractZehan Wang, Ke Lei, Chen Zhu, Jiawei Huang, Sashuai Zhou, Luping Liu, Xize Cheng, Shengpeng Ji, Zhenhui Ye, Tao Jin, Zhou Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zehan Wang 0001, Ke Lei, Jiawei Huang 0008, Sashuai Zhou, Luping Liu, Xize Cheng, Shengpeng Ji, Zhenhui Ye, Tao Jin 0004, Zhou Zhao 0001 |
ACL (1) | 10 |
| 2025 | SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative LanguageabstractContrastive Language-Image Pre-training (CLIP) learns robust visual models through language supervision, making it a crucial visual encoding technique for various applications. However, CLIP struggles with comprehending spatial concepts in images, potentially restricting the spatial intelligence of CLIP-based AI systems. In this work, we propose SpatialCLIP, an enhanced version of CLIP with better spatial understanding capabilities. To capture the intricate 3D spatial relationships in images, we improve both "visual model" and "language supervision" of CLIP. Specifically, we design 3D-inspired ViT to replace the standard ViT in CLIP. By lifting 2D image tokens into 3D space and incorporating design insights from point cloud networks, our visual model gains greater potential for spatial perception. Meanwhile, captions with accurate and detailed spatial information are very rare. To explore better language supervision for spatial understanding, we re-caption images and perturb their spatial phrases as negative descriptions, which compels the visual model to seek spatial cues to distinguish these hard negative captions. With the enhanced visual model, we introduce SpatialLLaVA, following the same LLaVA-1.5 training protocol, to investigate the importance of visual representations for MLLM’s spatial intelligence. Furthermore, we create SpatialBench, a benchmark specifically designed to evaluate CLIP and MLLM in spatial reasoning. Spatial-CLIP and SpatialLLaVA achieve substantial performance improvements, demonstrating stronger capabilities in spatial perception and reasoning, while maintaining comparable results on general-purpose benchmarks. Zehan Wang 0001, Sashuai Zhou, Shaoxuan He, Haifeng Huang 0001, Lihe Yang, Xize Cheng, Shengpeng Ji, Tao Jin 0004, Hengshuang Zhao, Zhou Zhao 0001 |
CVPR | 9 |
| 2025 | ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion MitigationabstractDiffusion customization methods have achieved impressive results with only a minimal number of user-provided images. However, existing approaches customize concepts collectively, whereas real-world applications often require sequential concept integration. This sequential nature can lead to catastrophic forgetting, where previously learned concepts are lost. In this paper, we investigate concept forgetting and concept confusion in the continual customization. To tackle these challenges, we present ConceptGuard, a comprehensive approach that combines shift embedding, concept-binding prompts and memory preservation regularization, supplemented by a priority queue which can adaptively update the importance and occurrence order of different concepts. These strategies can dynamically update, unbind and learn the relationship of the previous concepts, thus alleviating concept forgetting and confusion. Through comprehensive experiments, we show that our approach outperforms all the baseline methods consistently and significantly in both quantitative and qualitative analyses. Zirun Guo, Tao Jin 0004 |
CVPR | 2 |
| 2025 | Non-Natural Image Understanding with Advancing Frequency-based Vision EncodersabstractLarge language models (LLMs) have significantly enhanced cross-modal understanding capabilities by integrating visual encoders with textual embeddings, giving rise to multimodal large language models (MLLMs). However, these models struggle with non-natural images such as geometric and charts, particularly in fields like education and finance. Despite efforts to collect datasets and fine-tune the MLLMs, the gap with natural image understanding is still evident, and the cost of collecting large and diverse non-natural image datasets is high. To address this, we analyzed the limitations of transformer-based vision encoders(ViT) within existing MLLMs from a frequency perspective. Studies have shown that ViT models are less effective at capturing high-frequency information, impairing their ability to capture elements like points, lines, and angles in non-natural images. In response, we introduced FM-ViT, a frequency-modulated vision encoder that utilizes Fourier decomposition to extract high and low frequency components from self-attention features and re-weight them during tuning to non-natural images. In addition, we combine the features of CNN models with FM-ViT and propose EDGE, an MLLM with enhanced graphical encoders tailored for understanding non-natural images. Extensive experiments have confirmed the effectiveness of our FM-ViT and EDGE in 4 types. Project page. Yueying Feng, Shulei Wang, Tao Jin 0004, Zhou Zhao 0001, Fei Wu 0001, Chang Yao 0001, Jingyuan Chen 0003 |
CVPR | 5 |
| 2025 | Towards Transformer-Based Aligned Generation with Self-Coherence GuidanceabstractWe introduce a novel, training-free approach for enhancing alignment in Transformer-based Text-Guided Diffusion Models (TGDMs). Existing TGDMs often struggle to generate semantically aligned images, particularly when dealing with complex text prompts or multi-concept attribute binding challenges. Previous U-Net-based methods primarily optimized the latent space, but their direct application to Transformer-based architectures has shown limited effectiveness. Our method addresses these challenges by directly optimizing cross-attention maps during the generation process. Specifically, we introduce Self-Coherence Guidance, a method that dynamically refines attention maps using masks derived from previous denoising steps, ensuring precise alignment without additional training. To validate our approach, we constructed more challenging benchmarks for evaluating coarse-grained attribute binding, fine-grained attribute binding, and style binding. Experimental results demonstrate the superior performance of our method, significantly surpassing other state-of-the-art methods across all evaluated tasks. Our code is available at https://scg-diffusion.github.io/scg-diffusion. Shulei Wang, Hai Huang 0013, Hanting Wang, Sihang Cai, WenKang Han, Tao Jin 0004, Jingyuan Chen 0003, Jieming Zhu, Zhou Zhao 0001 |
CVPR | 7 |
| 2025 | PACHAT: Persona-Aware Speech Assistant for Multi-party DialogueabstractExtensive research on LLM-based spoken dialogue systems has significantly advanced the development of intelligent voice assistants.However, the integration of role information within speech remains an underexplored area, limiting its application in real-world scenarios, particularly in multi-party dialogue settings.With the growing demand for personalization, voice assistants that can recognize and remember users establish a deeper connection with them.We focus on enabling LLMs with speaker-awareness capabilities and enhancing their understanding of character settings through synthetic data to generate contextually appropriate responses.We introduce Persona-Dialogue, the first large-scale multi-party spoken dialogue dataset that incorporates speaker profiles.Based on this dataset, we propose PAChat, an architecture that simultaneously models both linguistic content and speaker features, allowing LLMs to map character settings to speaker identities in speech.Through extensive experiments, we demonstrate that PAChat successfully achieves speaker-specific responses, character understanding, and the generation of targeted replies in multi-party dialogue scenarios, surpassing existing spoken dialogue systems.For more details, please visit our demo page at https Xize Cheng, Linjun Li, Xiaoda Yang, Lujia Yang, Tao Jin 0004 |
EMNLP | 6 |
| 2025 | Curriculum Learning aided Audio-Visual Speech Recognition with Arbitrary Speaker NumberabstractRecently, audio-visual speech recognition has attracted increasing attention. However, most existing works only focused on scenarios with two speakers. In this work, we study the effect of speaker number in AVSR task and propose an end-to-end audio-visual speech recognition framework under a more realistic condition where the speaker number is arbitrary. Specifically, we adopted curriculum learning to train models from easy scenarios to hard ones and introduce a new training strategy named Challenge-based Curriculum Learning (CBCL) that forces the model to focus on hard, challenging data instead of easy ones during training. Further, to avoid scenario bias from unbalanced sampling during curriculum learning, we propose a Speaker-number Aware Mixture-of-Expert (SA-MoE) mechanism to explicitly model the characteristic difference in scenarios with different speaker numbers. Yuxiao Lin, Tao Jin 0004, Xize Cheng, Zhou Zhao 0001, Fei Wu 0001 |
ICASSP | 2 |
| 2025 | Open-Set Cross Modal Generalization via Multimodal Unified RepresentationabstractThis paper extends Cross Modal Generalization (CMG) to open-set environments by proposing the more challenging Open-set Cross Modal Generalization (OSCMG) task. This task evaluates multimodal unified representations in open-set conditions, addressing the limitations of prior closed-set cross-modal evaluations. OSCMG requires not only cross-modal knowledge transfer but also robust generalization to unseen classes within new modalities, a scenario frequently encountered in real-world applications. Existing multimodal unified representation work lacks consideration for open-set environments. To tackle this, we propose MICU, comprising two key components: Fine-Coarse Masked multimodal InfoNCE (FCMI) and Cross modal Unified Jigsaw Puzzles (CUJP). FCMI enhances multimodal alignment by applying contrastive learning at both holistic semantic and temporal levels, incorporating masking to enhance generalization. CUJP enhances feature diversity and model uncertainty by integrating modality-agnostic feature selection with self-supervised learning, thereby strengthening the model's ability to handle unknown categories in open-set tasks. Extensive experiments on CMG and the newly proposed OSCMG validate the effectiveness of our approach. The code is available at https://github.com/haihuangcode/CMG. Hai Huang 0013, Yan Xia 0006, Shulei Wang, Hanting Wang, Minghui Fang 0002, Shengpeng Ji, Sashuai Zhou, Tao Jin 0004, Zhou Zhao 0001 |
ICCV | 8 |
| 2025 | OmniBind: Large-scale Omni Multimodal Representation via Binding SpacesabstractRecently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Meanwhile, multimodal representation models have emerged as the foundation for these versatile multimodal understanding and generation pipeline. Models like CLIP, CLAP and ImageBind can map their specialized modalities into respective joint spaces. To construct a high-quality omni representation space that can be shared and expert in any modality, we propose to merge these advanced models into a unified space in scale. With this insight, we present \textbf{OmniBind}, advanced multimodal joint representation models via fusing knowledge of 14 pre-trained spaces, which support 3D, audio, image, video and language inputs. To alleviate the interference between different knowledge sources in integrated space, we dynamically assign weights to different spaces by learning routers with two objectives: cross-modal overall alignment and language representation decoupling. Notably, since binding and routing spaces only require lightweight networks, OmniBind is extremely training-efficient. Extensive experiments demonstrate the versatility and superiority of OmniBind as an omni representation model, highlighting its great potential for diverse applications, such as any-query and composable multimodal understanding. Zehan Wang 0001, Minjie Hong, Luping Liu, Rongjie Huang 0001, Xize Cheng, Shengpeng Ji, Tao Jin 0004, Hengshuang Zhao, Zhou Zhao 0001 |
ICLR | 9 |
| 2025 | VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?abstractWith the rapid advancement of large models, voice assistants are gradually acquiring the ability to engage in open-ended daily conversations with humans. However, current spoken dialogue systems often overlook multi-modal information in audio beyond text, such as speech rate, volume, emphasis, and background sounds. Relying solely on Automatic Speech Recognition (ASR) can lead to the loss of valuable auditory cues, thereby weakening the system’s ability to generate contextually appropriate responses. To address this limitation, we propose \textbf{VoxDialogue}, a comprehensive benchmark for evaluating the ability of spoken dialogue systems to understand multi-modal information beyond text. Specifically, we have identified 12 attributes highly correlated with acoustic information beyond words and have meticulously designed corresponding spoken dialogue test sets for each attribute, encompassing a total of 4.5K multi-turn spoken dialogue samples. Finally, we evaluated several existing spoken dialogue models, analyzing their performance on the 12 attribute subsets of VoxDialogue. Experiments have shown that in spoken dialogue scenarios, many acoustic cues cannot be conveyed through textual information and must be directly interpreted from the audio input. In contrast, while direct spoken dialogue systems excel at processing acoustic signals, they still face limitations in handling complex dialogue tasks due to their restricted context understanding capabilities. All data and code will be open source at \url{https://voxdialogue.github.io/}. Xize Cheng, Ruofan Hu 0002, Xiaoda Yang, Jingyu Lu 0001, Zehan Wang 0001, Shengpeng Ji, Rongjie Huang 0001, Tao Jin 0004, Zhou Zhao 0001 |
ICLR | 10 |
| 2025 | OmniSep: Unified Omni-Modality Sound Separation with Query-MixupabstractQuery-based sound separation (QSS) effectively isolate sound signals that match the content of a given query, enhancing the understanding of audio data. However, most existing QSS methods rely on a single modality for separation, lacking the ability to fully leverage homologous but heterogeneous information across multiple modalities for the same sound signal. To address this limitation, we introduce Omni-modal Sound Separation (**OmniSep**), a novel framework capable of isolating clean soundtracks based on omni-modal queries, encompassing both single-modal and multi-modal composed queries. Specifically, we introduce the **Query-Mixup** strategy, which blends query features from different modalities during training. This enables OmniSep to optimize multiple modalities concurrently, effectively bringing all modalities under a unified framework for sound separation. We further enhance this flexibility by allowing queries to influence sound separation positively or negatively, facilitating the retention or removal of specific sounds as desired. Finally, OmniSep employs a retrieval-augmented approach known as **Query-Aug**, which enables open-vocabulary sound separation. Experimental evaluations on MUSIC, VGGSOUND-CLEAN+, and MUSIC-CLEAN+ datasets demonstrate effectiveness of OmniSep, achieving state-of-the-art performance in text-, image-, and audio-queried sound separation tasks. For samples and further information, please visit the demo page at \url{https://omnisep.github.io/}. Xize Cheng, Zehan Wang 0001, Minghui Fang 0002, Rongjie Huang 0001, Shengpeng Ji, Jialong Zuo, Tao Jin 0004, Zhou Zhao 0001 |
ICLR | 9 |
| 2025 | Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal NoisesabstractTest-Time Adaptation (TTA) aims to tackle distribution shifts using unlabeled test data without access to the source data. In the context of multimodal data, there are more complex noise patterns than unimodal data such as simultaneous corruptions for multiple modalities and missing modalities. Besides, in real-world applications, corruptions from different distribution shifts are always mixed. Existing TTA methods always fail in such multimodal scenario because the abrupt distribution shifts will destroy the prior knowledge from the source model, thus leading to performance degradation. To this end, we reveal a new challenge named *multimodal wild TTA*. To address this challenging problem, we propose two novel strategies: sample identification with interquartile range **S**moothing and **u**nimodal assistance, and **M**utual **i**nformation sharing (SuMi). SuMi smooths the adaptation process by interquartile range which avoids the abrupt distribution shifts. Then, SuMi fully utilizes the unimodal features to select low-entropy samples with rich multimodal information for optimization. Furthermore, mutual information sharing is introduced to align the information, reduce the discrepancies and enhance the information utilization across different modalities. Extensive experiments show the effectiveness and superiority over existing methods under the complex noise patterns in multimodal data. Code is available at https://github.com/zrguo/SuMi. Zirun Guo, Tao Jin 0004 |
ICLR | 2 |
| 2025 | Diff-Prompt: Diffusion-Driven Prompt Generator with Mask SupervisionabstractPrompt learning has demonstrated promising results in fine-tuning pre-trained multimodal models. However, the performance improvement is limited when applied to more complex and fine-grained tasks. The reason is that most existing methods directly optimize the parameters involved in the prompt generation process through loss backpropagation, which constrains the richness and specificity of the prompt representations. In this paper, we propose Diffusion-Driven Prompt Generator (Diff-Prompt), aiming to use the diffusion model to generate rich and fine-grained prompt information for complex downstream tasks. Specifically, our approach consists of three stages. In the first stage, we train a Mask-VAE to compress the masks into latent space. In the second stage, we leverage an improved Diffusion Transformer (DiT) to train a prompt generator in the latent space, using the masks for supervision. In the third stage, we align the denoising process of the prompt generator with the pre-trained model in the semantic space, and use the generated prompts to fine-tune the model. We conduct experiments on a complex pixel-level downstream task, referring expression comprehension, and compare our method with various parameter-efficient fine-tuning approaches. Diff-Prompt achieves a maximum improvement of 8.87 in R@1 and 14.05 in R@5 compared to the foundation model and also outperforms other state-of-the-art methods across multiple metrics. The experimental results validate the effectiveness of our approach and highlight the potential of using generative models for prompt generation. Code is available at https://github.com/Kelvin-ywc/diff-prompt. Weicai Yan, Zirun Guo, Ye Wang 0018, Fangming Feng, Xiaoda Yang, Zehan Wang 0001, Tao Jin 0004 |
ICLR | 8 |
| 2025 | IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion ModelsabstractBridge models in image restoration construct a diffusion process from degraded to clear images. However, existing methods typically require training a bridge model from scratch for each specific type of degradation, resulting in high computational costs and limited performance. This work aims to efficiently leverage pretrained generative priors within existing image restoration bridges to eliminate this requirement. The main challenge is that standard generative models are typically designed for a diffusion process that starts from pure noise, while restoration tasks begin with a low-quality image, resulting in a mismatch in the state distributions between the two processes. To address this challenge, we propose a transition equation that bridges two diffusion processes with the same endpoint distribution. Based on this, we introduce the IRBridge framework, which enables the direct utilization of generative models within image restoration bridges, offering a more flexible and adaptable approach to image restoration. Extensive experiments on six image restoration tasks demonstrate that IRBridge efficiently integrates generative priors, resulting in improved robustness and generalization performance. Code will be available at GitHub. Hanting Wang, Tao Jin 0004, Shulei Wang, Hai Huang 0013, Shengpeng Ji, Zhou Zhao 0001 |
ICML | 2 |
| 2025 | Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
Ruofan Hu 0002, Yan Xia 0006, Minjie Hong, Jieming Zhu, Bo Chen 0023, Xiaoda Yang, Minghui Fang 0002, Tao Jin 0004 |
INTERSPEECH | 8 |
| 2025 | Multimodal Conditional Retrieval with High ControllabilityabstractSearching for images using text has limitations because language has difficulties in expressing certain abstract intentions, e.g. artistic styles are difficult to describe for non-experts. As for the image search image model, images can convey abstract intentions, but cannot express the specific purpose, so many of the current graph search works only have a single function, such as content search and style search. Our work aims to combine the strengths of both, merging the ability of text to express specific ideas with the ability of images to convey abstract concepts, thus achieving a better capture of the user's intentions. To this end, we propose CCSR, a multimodal conditional content-style joint retrieval model. Our model is the first to apply contrastive learning to conditional retrieval and introduces a novel Mixture-of-Expert models (MOE) system to enable collaboration between multiple expert systems. We adopt a novel prompt learning strategy that allows the model to adaptively select specific prompts, thereby enhancing its focus on the current task. In addition, to evaluate the joint content-style retrieval capability of our model, we present a new dataset, StyleCoco, containing rich content categories and style categories. The experimental results indicate that CCSR has achieved state-of-the-art performance in conditional style retrieval, content retrieval, and style-content retrieval. The dataset and code will be publicly available on https://mccsr.github.io/. Xiaoda Yang, Xize Cheng, Minghui Fang 0002, Hongshun Qiu, Jiaqi Duan, Sihang Cai, Zehan Wang 0001, Ruofan Hu 0002, Zhou Zhao 0001, Tao Jin 0004 |
KDD (2) | 13 |
| 2025 | Speech Token Prediction via Compressed-to-fine Language Modeling for Speech GenerationabstractNeural audio codecs, used as speech tokenizers, have demonstrated remarkable potential in the field of speech generation. However, to ensure high-fidelity audio reconstruction, neural audio codecs typically encode audio into long sequences of speech tokens, posing a significant challenge for downstream language models in long-context modeling. We observe that speech token sequences exhibit short-range dependency: due to the monotonic alignment between text and speech in text-to-speech (TTS) tasks, the prediction of the current token primarily relies on its local context, while long-range tokens contribute less to the current token prediction and often contain redundant information. Inspired by this observation, we propose a compressed-to-fine language modeling approach to address the challenge of long sequence speech tokens within neural codec language models: (1) Fine-grained Initial and Short-range Information: Our approach retains the prompt and local tokens during prediction to ensure text alignment and the integrity of paralinguistic information; (2) Compressed Long-range Context: Our approach compresses long-range token spans into compact representations to reduce redundant information while preserving essential semantics. Extensive experiments on various neural audio codecs and downstream language models validate the effectiveness and generalizability of the proposed approach, highlighting the importance of token compression in improving speech generation within neural codec language models. The demo of audio samples will be available at https://anonymous.4open.science/r/SpeechTokenPredictionViaCompressedToFinedLM. Wenrui Liu 0003, Qian Chen 0003, Wen Wang 0019, Guanrou Yang, Minghui Fang 0002, Jialong Zuo, Xiaoda Yang, Tao Jin 0004, Jin Xu 0010, Yafeng Chen, Jionghao Bai, Zhifang Guo |
ACM Multimedia | 9 |
| 2025 | ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
Yu Zhang 0126, Wenxiang Guo, Changhao Pan, Tao Jin 0004, Zhou Zhao 0001 |
ACM Multimedia | 5 |
| 2025 | TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather RemovalabstractImage restoration under adverse weather conditions has been extensively explored, leading to numerous high-performance methods. In particular, recent advances in All-in-One approaches have shown impressive results by training on multi-task image restoration datasets. However, most of these methods rely on dedicated network modules or parameters for each specific degradation type, resulting in a significant parameter overhead. Moreover, the relatedness across different restoration tasks is often overlooked. In light of these issues, we propose a parameter-efficient All-in-One image restoration framework that leverages task-aware enhanced prompts to tackle various adverse weather degradations. Specifically, we adopt a two-stage training paradigm consisting of a pretraining phase and a prompt-tuning phase to mitigate parameter conflicts across tasks. We first employ supervised learning to acquire general restoration knowledge, and then adapt the model to handle specific degradation via trainable soft prompts. Crucially, we enhance these task-specific prompts in a task-aware manner. We apply low-rank decomposition to these prompts to capture both task-general and task-specific characteristics, and impose contrastive constraints to better align them with the actual inter-task relatedness. These enhanced prompts not only improve the parameter efficiency of the restoration model but also enable more accurate task modeling, as evidenced by t-SNE analysis. Experimental results on different restoration tasks demonstrate that the proposed method achieves superior performance with only 2.75M parameters. Hanting Wang, Shengpeng Ji, Shulei Wang, Hai Huang 0013, Qifei Zhang 0001, Tao Jin 0004 |
ACM Multimedia | 7 |
| 2025 | Efficient Prompting for Continual Adaptation to Missing ModalitiesabstractZirun Guo, Shulei Wang, Wang Lin, Weicai Yan, Yangyang Wu, Tao Jin. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zirun Guo, Shulei Wang, Weicai Yan, Tao Jin 0004 |
NAACL (Long Papers) | 6 |
| 2025 | AHa-Bench: Benchmarking Audio Hallucinations in Large Audio-Language ModelsabstractHallucinations present a significant challenge in the development and evaluation of large language models (LLMs), directly affecting their reliability and accuracy. While notable advancements have been made in research on textual and visual hallucinations, there is still a lack of a comprehensive benchmark for evaluating auditory hallucinations in large audio language models (LALMs). To fill this gap, we introduce AHa-Bench, a systematic and comprehensive benchmark for audio hallucinations. Audio data, in particular, uniquely combines the multi-attribute complexity of visual data with the semantic richness of textual data, leading to auditory hallucinations that share characteristics with both visual and textual hallucinations. Based on the source of these hallucinations, AHa-Bench categorizes them into semantic hallucinations, acoustic hallucinations, and semantic-acoustic confusion hallucinations. In addition, we systematically evaluate seven open-source local perception language models (LALMs), demonstrating the challenges these models face in audio understanding, especially when it comes to jointly understanding semantic and acoustic information. Through the development of a comprehensive evaluation framework, AHa-Bench aims to enhance the robustness and stability of LALMs, fostering more reliable and nuanced audio understanding in LALMs. The benchmark dataset is available at \url{https://huggingface.co/datasets/ahabench/AHa-Bench}. Xize Cheng, Chenyuhao Wen, Shannon Yu, Zehan Wang 0001, Shengpeng Ji, Siddhant Arora, Tao Jin 0004, Shinji Watanabe 0001, Zhou Zhao 0001 |
NeurIPS | 8 |
| 2025 | New Concentration Bounds and Their Applications in Online Resource Allocation
Jinshan Zhang 0001, Biaoshuai Tao, Meng Xi 0002, Tao Jin 0004, Jianwei Yin |
WINE | 5 |
| 2025 | Recognize-and-tell: Generating video captions with textual cue in scene
Tao Jin 0004, Hao Jiang 0062, Jian Wang 0119, Jingyuan Chen 0003, Zhou Zhao 0001, Zhongfei Zhang |
Expert Syst. Appl. | 1 |
| 2025 | Multi-talker audio-visual speech recognition towards diverse scenariosabstractRecently, audio–visual speech recognition (AVSR) has attracted increasing attention. However, most existing works simplify the complex challenges in real-world applications and only focus on scenarios with two speakers and perfectly aligned audio-video clips. In this work, we study the effect of speaker number and modal misalignment in the AVSR task, and propose an end-to-end AVSR framework under a more realistic condition. Specifically, we propose a speaker-number-aware mixture-of-experts (SA-MoE) mechanism to explicitly model the characteristic difference in scenarios with different speaker numbers, and a cross-modal realignment (CMR) module for robust handling of asynchronous inputs. We also use the underlying difficulty difference and introduce a new training strategy named challenge-based curriculum learning (CBCL), which forces the model to focus on difficult, challenging data instead of simple data to improve efficiency. Yuxiao Lin, Tao Jin 0004, Xize Cheng, Zhou Zhao 0001, Fei Wu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2024 | Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion RecognitionabstractThe development of multimodal models has significantly advanced multimodal sentiment analysis and emotion recognition.However, in real-world applications, the presence of various missing modality cases often leads to a degradation in the model's performance.In this work, we propose a novel multimodal Transformer framework using prompt learning to address the issue of missing modalities.Our method introduces three types of prompts: generative prompts, missing-signal prompts, and missingtype prompts.These prompts enable the generation of missing modality features and facilitate the learning of intra-and inter-modality information.Through prompt learning, we achieve a substantial reduction in the number of trainable parameters.Our proposed method outperforms other methods significantly across all evaluation metrics.Extensive experiments and ablation studies are conducted to demonstrate the effectiveness and robustness of our method, showcasing its ability to effectively handle missing modalities.Codes are available at https://github.com/zrguo/MPLMM. Zirun Guo, Tao Jin 0004, Zhou Zhao 0001 |
ACL (1) | 2 |
| 2024 | Rethinking the Multimodal Correlation of Multimodal Sequential Learning via Generalizable Attentional Results AlignmentabstractTransformer-based methods have gone mainstream in multimodal sequential learning.The intra and inter modality interactions are captured by the query-key associations of multihead attention.In this way, the calculated multimodal contexts (attentional results) are expected to be relevant to the query modality.However, in existing literature, the alignment degree between different calculated attentional results of the same query are under-explored.Based on this concern, we propose a new constrained scheme called Multimodal Contextual Contrast (MCC), which could align the multiple attentional results from both local and global perspectives, making the information capture more efficient.Concretely, the calculated attentional results of different modalities are mapped into a common feature space, those attentional vectors with the same query are considered as a positive group and the remaining sets are negative.From local perspective, we sample the negative groups for a positive group by randomly changing the sequential step of one specific context and keeping the other stay the same.From coarse global perspective, we divide all the contextual groups into two sets (i.e., aligned and unaligned), making the total score of aligned group relatively large.We extend the vectorial inner product operation for more input and calculate the aligned score for each multimodal group.Considering that the computational complexity scales exponentially to the number of modalities, we adopt stochastic expectation approximation (SEA) for the real process.The extensive experimental results on several tasks reveal the effectiveness of our contributions. Tao Jin 0004, Ye Wang 0018, Linjun Li, Xize Cheng, Zhou Zhao 0001 |
ACL (1) | 1 |
| 2024 | Uni-Dubbing: Zero-Shot Speech Synthesis from Visual ArticulationabstractSongju Lei, Xize Cheng, Mengjiao Lyu, Jianqiao Hu, Jintao Tan, Runlin Liu, Lingyu Xiong, Tao Jin, Xiandong Li, Zhou Zhao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Songju Lei, Xize Cheng, Mengjiao Lyu, Jianqiao Hu, Jintao Tan, Runlin Liu, Lingyu Xiong, Tao Jin 0004, Xiandong Li, Zhou Zhao 0001 |
ACL (1) | 8 |
| 2024 | MPOD123: One Image to 3D Content Generation Using Mask-Enhanced Progressive Outline-to-Detail OptimizationabstractRecent advancements in single image driven 3D content generation have been propelled by leveraging prior knowledge from pretrained 2D diffusion models. However, the 3D content generated by existing methods often exhibits distorted outline shapes and inadequate details. To solve this problem, we propose a novel framework called Mask-enhanced Progressive Outline-to-Detail optimization (aka. MPOD123), which consists of two stages. Specifically, in the first stage, MPOD123 utilizes the pretrained view-conditioned diffusion model to guide the outline shape optimization of the 3D content. Given certain viewpoint, we estimate outline shape priors in the form of 2D mask from the 3D content by leveraging opacity calculation. In the second stage, MPOD123 incorporates Detail Appearance Inpainting (DAI) to guide the refinement on local geometry and texture with the shape priors. The essence of DAI lies in the Mask Rectified Cross-Attention (MRCA), which can be conveniently plugged in the stable diffusion model. The MRCA module utilizes the mask to rectify the attention map from each cross-attention layer. Accompanied with this new module, DAI is capable of guiding the detail refinement of the 3D content, while better preserves the outline shape. To assess the applicability in practical scenarios, we contribute a new dataset modeled on real-world e-commerce environments. Extensive quantitative and qualitative experiments on this dataset and open benchmarks demonstrate the effectiveness of MPOD123 over the state-of-the-arts. Jimin Xu, Tianbao Wang, Tao Jin 0004, Shengyu Zhang 0001, Jiangjing Lyu, Chengfei Lv, Chaoyue Niu, Zhou Yu 0001, Zhou Zhao 0001, Fei Wu 0001 |
CVPR | 3 |
| 2024 | AudioVSR: Enhancing Video Speech Recognition with Audio DataabstractXiaoda Yang, Xize Cheng, Jiaqi Duan, Hongshun Qiu, Minjie Hong, Minghui Fang, Shengpeng Ji, Jialong Zuo, Zhiqing Hong, Zhimeng Zhang, Tao Jin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Xiaoda Yang, Xize Cheng, Jiaqi Duan, Hongshun Qiu, Minjie Hong, Minghui Fang 0002, Shengpeng Ji, Jialong Zuo, Zhiqing Hong, Tao Jin 0004 |
EMNLP | 11 |
| 2024 | FreeBind: Free Lunch in Unified Multimodal Space via Knowledge FusionabstractUnified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting problems make it challenging to further enhance pre-trained unified spaces. In this work, we propose FreeBind, an idea that treats multimodal representation spaces as basic units, and freely augments pre-trained unified space by integrating knowledge from extra expert spaces via “space bonds". Specifically, we introduce two kinds of basic space bonds: 1) Space Displacement Bond and 2) Space Combination Bond. Based on these basic bonds, we design Complex Sequential & Parallel Bonds to effectively integrate multiple spaces simultaneously. Benefiting from the modularization concept, we further propose a coarse-to-fine customized inference strategy to flexibly adjust the enhanced unified space for different purposes. Experimentally, we bind ImageBind with extra image-text and audio-text expert spaces, resulting in three main variants: ImageBind++, InternVL_IB, and InternVL_IB++. These resulting spaces outperform ImageBind on 5 audio-image-text downstream tasks across 9 datasets. Moreover, via customized inference, it even surpasses the advanced audio-text and image-text expert spaces. Our code and checkpoints are released at https://github.com/zehanwang01/FreeBind Zehan Wang 0001, Xize Cheng, Rongjie Huang 0001, Luping Liu, Zhenhui Ye, Haifeng Huang 0001, Yang Zhao 0022, Tao Jin 0004, Peng Gao 0007, Zhou Zhao 0001 |
ICML | 9 |
| 2024 | Non-confusing Generation of Customized Concepts in Diffusion ModelsabstractWe tackle the common challenge of inter-concept visual confusion in compositional concept generation using text-guided diffusion models (TGDMs). It becomes even more pronounced in the generation of customized concepts, due to the scarcity of user-provided concept visual examples. By revisiting the two major stages leading to the success of TGDMs---1) contrastive image-language pre-training (CLIP) for text encoder that encodes visual semantics, and 2) training TGDM that decodes the textual embeddings into pixels---we point that existing customized generation methods only focus on fine-tuning the second stage while overlooking the first one. To this end, we propose a simple yet effective solution called CLIF: contrastive image-language fine-tuning. Specifically, given a few samples of customized concepts, we obtain non-confusing textual embeddings of a concept by fine-tuning CLIP via contrasting a concept and the over-segmented visual regions of other concepts. Experimental results demonstrate the effectiveness of CLIF in preventing the confusion of multi-customized concept generation. Project page: https://clif-official.github.io/clif. Jingyuan Chen 0003, Jiaxin Shi, Junzhong Miao, Tao Jin 0004, Zhou Zhao 0001, Fei Wu 0001, Shuicheng Yan, Hanwang Zhang |
ICML | 7 |
| 2024 | EAGER: Two-Stream Generative Recommender with Behavior-Semantic CollaborationabstractGenerative retrieval has recently emerged as a promising approach to sequential recommendation, framing candidate item retrieval as an autoregressive sequence generation problem. However, existing generative methods typically focus solely on either behavioral or semantic aspects of item information, neglecting their complementary nature and thus resulting in limited effectiveness. To address this limitation, we introduce EAGER, a novel generative recommendation framework that seamlessly integrates both behavioral and semantic information. Specifically, we identify three key challenges in combining these two types of information: a unified generative architecture capable of handling two feature types, ensuring sufficient and independent learning for each type, and fostering subtle interactions that enhance collaborative information utilization. To achieve these goals, we propose (1) a two-stream generation architecture leveraging a shared encoder and two separate decoders to decode behavior tokens and semantic tokens with a confidence-based ranking strategy; (2) a global contrastive task with summary tokens to achieve discriminative decoding for each type of information; and (3) a semantic-guided transfer task designed to implicitly promote cross-interactions through reconstruction and estimation objectives. We validate the effectiveness of EAGER on four public benchmarks, demonstrating its superior performance compared to existing methods. Our source code will be publicly available on PapersWithCode.com. Ye Wang 0018, Jiahao Xun, Minjie Hong, Jieming Zhu, Tao Jin 0004, Haoyuan Li 0002, Linjun Li, Yan Xia 0006, Zhou Zhao 0001, Zhenhua Dong |
KDD | 5 |
| 2024 | Calibrating Prompt from History for Continual Vision-Language Retrieval and GroundingabstractIn the field of machine learning, continual learning is a crucial concept that allows models to adapt to non-stationary data distributions. However, most of the existing works focus on uni-modal settings and ignore the multi-modal data. In this paper, to enable neural networks better understand diverse modalities in real-world scenario, we investigate continual learning for two typical vision-language applications, i.e. retrieval and grounding. Instead of conventional exemplar-based methods, we leverage the pre-trained transformer model (e.g. CLIP/GLIP) and the prompt technique to tackle this problem. Under this scheme, we identify two critical limitations in existing methods: (1) Unfamiliarity across tasks, which prevents task-specific prompts from achieving forward propagation; and (2) Heterogeneity between modalities, which makes it difficult to guarantee a consistent optimization direction for prompts of different modalities. To overcome these constraints, we design Historical Prompt Calibration that includes two objectives to calibrate prompts. First, the intra-modal relevance estimation helps encode sufficient task-specific information for prompts, with the help a relevance estimator developed for recognizing task relevance. Second, the inter-modal consistency alignment enhances the agreement of the two modality-specific prompts in the current task by contrasting them with the prompts from previous tasks. We evaluate the superiority of our strategy over state-of-the arts methods by four vision-language applications, including two retrieval tasks (i.e. image- and video-text retrieval) and two grounding tasks (i.e. referring expression comprehension and segmentation). Tao Jin 0004, Weicai Yan, Ye Wang 0018, Sihang Cai, Qifan Shuai, Zhou Zhao 0001 |
ACM Multimedia | 1 |
| 2024 | Boosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented PromptsabstractIn the burgeoning field of Audio-Visual Speech Recognition (AVSR), extant research has predominantly concentrated on the training paradigms tailored for high-quality resources. However, owing to the challenges inherent in real-world data collection, audio-visual data are frequently affected by modality-distortion, which encompasses audio-visual asynchrony, video noise and audio noise. The recognition accuracy of existing AVSR method is significantly compromised when multiple modality-distortion coexist in low-resource data. In light of the above challenges, we propose PCD: cluster-Prompt with Contrastive Decomposition, a robust framework for modality-distortion speech recognition, specifically devised to transpose the pre-trained knowledge from high-resource domain to the targeted domain by leveraging contrast-augmented prompts. In contrast to previous studies, we take into consideration the possibility of various types of distortion in both the audio and visual modalities. Concretely, we design bespoke prompts to delineate each modality-distortion, guiding the model to achieve speech recognition applicable to various distortion scenarios with quite few learnable parameters. To materialize the prompt mechanism, we employ multiple cluster-based strategies that better suits the pre-trained audio-visual model. Additionally, we design a contrastive decomposition mechanism to restrict the explicit relationships among various modality conditions, given their shared task knowledge and disparate modality priors. Extensive results on LRS2 dataset demonstrate that PCD achieves state-of-the-art performance for audio-visual speech recognition under the constraints of distorted resources. Code is available at https://github.com/ballooncatt/PCD. Xize Cheng, Xiaoda Yang, Hanting Wang, Zhou Zhao 0001, Tao Jin 0004 |
ACM Multimedia | 6 |
| 2024 | Low-rank Prompt Interaction for Continual Vision-Language RetrievalabstractResearch on continual learning in multi-modal tasks has been receiving increasing attention. However, most existing work overlooks the explicit cross-modal and cross-task interactions. In this paper, we innovatively propose the Low-rank Prompt Interaction (LPI) to address this general problem of multi-modal understanding, which considers both cross-modal and cross-task interactions. Specifically, as for the former, we employ multi-modal correlation modules for corresponding Transformer layers. Considering that the training parameters scale to the number of layers and tasks, we propose low-rank interaction-augmented decomposition to avoid memory explosion while enhancing the cross-modal association through sharing and separating common-specific low-rank factors. In addition, due to the multi-modal semantic differences carried by the low-rank initialization, we adopt hierarchical low-rank contrastive learning to ensure training robustness. As for the latter, we initially employ a visual analysis and identify that different tasks have clear distinctions in proximity. Therefore, we introduce explicit task contrastive constraints in the prompt learning process based on task semantic distances. Experiments on two retrieval tasks show performance improvements with the introduction of a minimal number of parameters, demonstrating the effectiveness of our method. Code is available at https://github.com/Kelvin-ywc/LPI. Weicai Yan, Ye Wang 0018, Zirun Guo, Zhou Zhao 0001, Tao Jin 0004 |
ACM Multimedia | 6 |
| 2024 | SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task LearningabstractTalking Face Generation (TFG) reconstructs facial motions concerning lips given speech input, which aims to generate highquality, synchronized, and lip-readable videos. Previous efforts have achieved success in generating quality and synchronization, and recently, there has been an increasing focus on the importance of intelligibility. Despite these efforts, there remains a challenge in achieving a balance among quality, synchronization, and intelligibility, often resulting in trade-offs that compromise one aspect in favor of another. In light of this, we propose SyncTalklip, a novel dual-tower framework designed to overcome the challenges of synchronization while improving lip-reading performance. To enhance the performance of SyncTalklip in both synchronization and intelligibility, we design AV-SyncNet, a pre-trained multi-task model, aiming to achieve a dual-focus on synchronization and intelligibility. Moreover, we propose a novel cross-modal contrastive learning bringing audio and video closer to enhance synchronization. Experimental results demonstrate that SyncTalklip achieves state-of-the-art performance in quality, intelligibility, and synchronization. Furthermore, extensive experiments have demonstrated our model's generalizability across domains. The code and demo is available at https://sync-talklip.github.io. Xiaoda Yang, Xize Cheng, Minghui Fang 0002, Jialong Zuo, Shengpeng Ji, Zhou Zhao 0001, Tao Jin 0004 |
ACM Multimedia | 8 |
| 2024 | Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language PromptabstractYongqi Wang, Ruofan Hu, Rongjie Huang, Zhiqing Hong, Ruiqi Li, Wenrui Liu, Fuming You, Tao Jin, Zhou Zhao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Ruofan Hu 0002, Rongjie Huang 0001, Zhiqing Hong, Ruiqi Li 0002, Wenrui Liu 0003, Fuming You, Tao Jin 0004, Zhou Zhao 0001 |
NAACL-HLT | 8 |
| 2024 | Classifier-guided Gradient Modulation for Enhanced Multimodal LearningabstractMultimodal learning has developed very fast in recent years. However, during the multimodal training process, the model tends to rely on only one modality based on which it could learn faster, thus leading to inadequate use of other modalities. Existing methods to balance the training process always have some limitations on the loss functions, optimizers and the number of modalities and only consider modulating the magnitude of the gradients while ignoring the directions of the gradients. To solve these problems, in this paper, we present a novel method to balance multimodal learning with **C**lassifier-**G**uided **G**radient **M**odulation (CGGM), considering both the magnitude and directions of the gradients. We conduct extensive experiments on four multimodal datasets: UPMC-Food 101, CMU-MOSI, IEMOCAP and BraTS 2021, covering classification, regression and segmentation tasks. The results show that CGGM outperforms all the baselines and other state-of-the-art methods consistently, demonstrating its effectiveness and versatility. Our code is available at https://github.com/zrguo/CGGM. Zirun Guo, Tao Jin 0004, Jingyuan Chen 0003, Zhou Zhao 0001 |
NeurIPS | 2 |
| 2024 | Action Imitation in Common Action Space for Customized Action Image SynthesisabstractWe propose a novel method, \textbf{TwinAct}, to tackle the challenge of decoupling actions and actors in order to customize the text-guided diffusion models (TGDMs) for few-shot action image generation. TwinAct addresses the limitations of existing methods that struggle to decouple actions from other semantics (e.g., the actor's appearance) due to the lack of an effective inductive bias with few exemplar images. Our approach introduces a common action space, which is a textual embedding space focused solely on actions, enabling precise customization without actor-related details. Specifically, TwinAct involves three key steps: 1) Building common action space based on a set of representative action phrases; 2) Imitating the customized action within the action space; and 3) Generating highly adaptable customized action images in diverse contexts with action similarity loss. To comprehensively evaluate TwinAct, we construct a novel benchmark, which provides sample images with various forms of actions. Extensive experiments demonstrate TwinAct's superiority in generating accurate, context-independent customized actions while maintaining the identity consistency of different subjects, including animals, humans, and even customized actors. Jingyuan Chen 0003, Jiaxin Shi, Zirun Guo, Zehan Wang 0001, Tao Jin 0004, Zhou Zhao 0001, Fei Wu 0001, Shuicheng Yan, Hanwang Zhang |
NeurIPS | 7 |
| 2024 | E3: Exploring Embodied Emotion Through A Large-Scale Egocentric Video DatasetabstractUnderstanding human emotions is fundamental to enhancing human-computer interaction, especially for embodied agents that mimic human behavior. Traditional emotion analysis often takes a third-person perspective, limiting the ability of agents to interact naturally and empathetically. To address this gap, this paper presents $E^3$ for Exploring Embodied Emotion, the first massive first-person view video dataset. $E^3$ contains more than $50$ hours of video, capturing $8$ different emotion types in diverse scenarios and languages. The dataset features videos recorded by individuals in their daily lives, capturing a wide range of real-world emotions conveyed through visual, acoustic, and textual modalities. By leveraging this dataset, we define $4$ core benchmark tasks - emotion recognition, emotion classification, emotion localization, and emotion reasoning - supported by more than $80$k manually crafted annotations, providing a comprehensive resource for training and evaluating emotion analysis models. We further present Emotion-LlaMa, which complements visual modality with acoustic modality to enhance the understanding of emotion in first-person videos. The results of comparison experiments with a large number of baselines demonstrate the superiority of Emotion-LlaMa and set a new benchmark for embodied emotion analysis. We expect that $E^3$ can promote advances in multimodal understanding, robotics, and augmented reality, and provide a solid foundation for the development of more empathetic and context-aware embodied agents. Yueying Feng, WenKang Han, Tao Jin 0004, Zhou Zhao 0001, Fei Wu 0001, Chang Yao 0001, Jingyuan Chen 0003 |
NeurIPS | 4 |
| 2024 | Extending Multi-modal Contrastive RepresentationsabstractMulti-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high dependence on large-scale, high-quality paired data and the expensive training costs limit their further development. Inspired by recent C-MCR, this paper proposes $\textbf{Ex}$tending $\textbf{M}$ultimodal $\textbf{C}$ontrastive $\textbf{R}$epresentation (Ex-MCR), a training-efficient and paired-data-free method to build unified contrastive representation for many modalities. Since C-MCR is designed to learn a new latent space for the two non-overlapping modalities and projects them onto this space, a significant amount of information from their original spaces is lost in the projection process. To address this issue, Ex-MCR proposes to extend one modality's space into the other's, rather than mapping both modalities onto a completely new space. This method effectively preserves semantic alignment in the original space. Experimentally, we extend pre-trained audio-text and 3D-image representations to the existing vision-text space. Without using paired data, Ex-MCR achieves comparable performance to advanced methods on a series of audio-image-text and 3D-image-text tasks and achieves superior performance when used in parallel with data-driven methods. Moreover, semantic alignment also emerges between the extended modalities (e.g., audio and 3D). Zehan Wang 0001, Luping Liu, Rongjie Huang 0001, Xize Cheng, Zhenhui Ye, Huadai Liu, Haifeng Huang 0001, Yang Zhao 0022, Tao Jin 0004, Zhou Zhao 0001 |
NeurIPS | 11 |
| 2024 | GTADT: Gated tone-sensitive acne grading via augmented domain transfer
Min Tan 0005, Ruirui Wang, Ankur Purwar, Tao Jin 0004, Jun Yu 0002, Alex Chichung Kot |
Multim. Tools Appl. | 4 |
| 2024 | Multi-Granularity Relational Attention Network for Audio-Visual Question AnsweringabstractRecent methods for video question answering (VideoQA), aiming to generate answers based on given questions and video content, have made significant progress in cross-modal interaction. From the perspective of video understating, these existing frameworks concentrate on the various levels of visual content, partially assisted by subtitles. However, audio information is also instrumental in helping get correct answers, especially in videos with real-life scenarios. Indeed, in some cases, both audio and visual contents are required and complement each other to answer questions, which is defined as audio-visual question answering (AVQA). In this paper, we focus on importing raw audio for AVQA and contribute in three ways. Firstly, due to no dataset annotating QA pairs for raw audio, we introduce E-AVQA, a manually annotated and large-scale dataset involving multiple modalities. E-AVQA consists of 34,033 QA pairs on 33,340 clips of 18,786 videos from the e-commerce scenarios. Secondly, we propose a multi-granularity relational attention method with contrastive constraints between audio and visual features after the interaction, named MGN, which captures local sequential representation by leveraging the pairwise potential attention mechanism and obtains global multi-modal representation via designing the novel ternary potential attention mechanism. Thirdly, our proposed MGN outperforms the baseline on dataset E-AVQA, achieving 20.73% on [email protected] and 19.81% on BLEU@1, demonstrating its superiority with at least 1.02 improvement on [email protected] and about 10% on timing complexity over the baseline. Linjun Li, Tao Jin 0004, Hao Jiang 0062, Wenwen Pan 0003, Jian Wang 0119, Shuwen Xiao, Yan Xia 0006, Zhou Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality AlignmentabstractSpeech Recognition builds a bridge between the multimedia streaming (audio-only, visualonly or audio-visual) and the corresponding text transcription.However, when training the specific model of new domain, it often gets stuck in the lack of new-domain utterances, especially the labeled visual utterances.To break through this restriction, we attempt to achieve zero-shot modality transfer by maintaining the multi-modality alignment in phoneme space learned with unlabeled multimedia utterances in the high resource domain during the pretraining (Shi et al., 2022), and propose a training system Open-modality Speech Recognition (OpenSR) that enables the models trained on a single modality (e.g., audio-only) applicable to more modalities (e.g., visual-only and audio-visual).Furthermore, we employ a cluster-based prompt tuning strategy to handle the domain shift for the scenarios with only common words in the new domain utterances.We demonstrate that OpenSR enables modality transfer from one to any in three different settings (zero-, few-and fullshot), and achieves highly competitive zeroshot performance compared to the existing fewshot and full-shot lip-reading methods.To the best of our knowledge, OpenSR achieves the state-of-the-art performance of word error rate in LRS2 on audio-visual speech recognition and lip-reading with 2.7% and 25.0%, respectively.The code and demo are available at https://github.com/Exgc/OpenSR. Xize Cheng, Tao Jin 0004, Linjun Li, Xinyu Duan, Zhou Zhao 0001 |
ACL (1) | 2 |
| 2023 | TAVT: Towards Transferable Audio-Visual Text GenerationabstractAudio-visual text generation aims to understand multi-modality contents and translate them into texts.Although various transfer learning techniques of text generation have been proposed, they focused on uni-modal analysis (e.g., text-to-text, visual-to-text) and lack consideration of multi-modal content and cross-modal relation.Motivated by the fact that humans can recognize the timbre of the same low-level concepts (e.g., footstep, rainfall, and laughing), even in different visual conditions, we aim to mitigate the domain discrepancies by audiovisual correlation.In this paper, we propose a novel Transferable Audio-Visual Text Generation framework, named TAVT, which consists of two key components: Audio-Visual Meta-Mapper (AVMM) and Dual Counterfactual Contrastive Learning (DCCL).(1) AVMM first introduces a universal auditory semantic space and drifts the domain-invariant low-level concepts into visual prefixes.Then the reconstructbased learning encourages the AVMM to learn "which pixels belong to the same sound" and achieve audio-enhanced visual prefix.The welltrained AVMM can be further applied to unimodal setting.(2) Furthermore, DCCL leverages the destructive counterfactual transformations to provide cross-modal constraints for AVMM from the perspective of feature distribution and text generation.(3) The experimental results show that TAVT outperforms the stateof-the-art methods across multiple domains (cross-datasets, cross-categories) and various modal settings (uni-modal, multi-modal). Tao Jin 0004, Wenwen Pan 0003, Linjun Li, Xize Cheng, Ye Wang 0018, Zhou Zhao 0001 |
ACL (1) | 2 |
| 2023 | Weakly-Supervised Spoken Video Grounding via Semantic Interaction LearningabstractThe task of spoken video grounding aims to localize moments in videos that are relevant to descriptive spoken queries.However, extracting semantic information from speech and modeling the cross-modal correlation pose two critical challenges.Previous studies solve them by representing spoken queries based on the matched video frames, which require tremendous effort for frame-level labeling.In this work, we investigate weakly-supervised spoken video grounding, i.e., learning to localize moments without expensive temporal annotations.To effectively represent the cross-modal semantics, we propose Semantic Interaction Learning (SIL), a novel framework consisting of the acoustic-semantic pre-training (ASP) and acoustic-visual contrastive learning (AVCL).In ASP, we pre-train an effective encoder for the grounding task with three comprehensive tasks, where the robustness task enhances stability by explicitly capturing the invariance between time-and frequency-domain features, the conciseness task avoids over-smooth attention by compressing long sequence into segments, and the semantic task improves spoken language understanding by modeling the precise semantics.In AVCL, we mine pseudo labels with discriminative sampling strategies and directly strengthen the interaction between speech and video by maximizing their mutual information.Extensive experiments demonstrate the effectiveness and superiority of our method.1 Ye Wang 0018, Shengyu Zhang 0001, Tao Jin 0004, Linjun Li, Xize Cheng, Zhou Zhao 0001 |
ACL (1) | 4 |
| 2023 | DATE: Domain Adaptive Product Seeker for E-CommerceabstractProduct Retrieval (PR) and Grounding (PG), aiming to seek image and object-level products respectively according to a textual query, have attracted great interest recently for better shopping experience. Owing to the lack of relevant datasets, we collect two large-scale benchmark datasets from Taobao Mall and Live domains with about 474k and 101k image-query pairs for PR, and manually annotate the object bounding boxes in each image for PG. As annotating boxes is expensive and time-consuming, we attempt to transfer knowledge from annotated domain to unannotated for PG to achieve unsupervised Domain Adaptation (PG-DA). We propose a Domain Adaptive Product Seeker (DATE) framework, regarding PR and PG as Product Seeking problem at different levels, to assist the query date the product. Concretely, we first design a semantics-aggregated feature extractor for each modality to obtain concentrated and comprehensive features for following efficient retrieval and fine-grained grounding tasks. Then, we present two cooperative seekers to simultaneously search the image for PR and localize the product for PG. Besides, we devise a domain aligner for PG-DA to alleviate unimodal marginal and multi-modal conditional distribution shift between source and target domains, and design a pseudo box generator to dynamically select reliable instances and generate bounding boxes for further knowledge transfer. Extensive experiments show that our DATE achieves satisfactory performance in fully-supervised PR, PG and unsupervised PG-DA. Our desensitized datasets will be publicly available here11https://github.com/Taobao-live/Product-Seeking. Haoyuan Li 0002, Hao Jiang 0062, Tao Jin 0004, Zhijie Lin 0001, Yang Zhao 0022, Zhou Zhao 0001 |
CVPR | 3 |
| 2023 | Gloss Attention for Gloss-free Sign Language TranslationabstractMost sign language translation (SLT) methods to date require the use of gloss annotations to provide additional supervision information, however, the acquisition of gloss is not easy. To solve this problem, we first perform an analysis of existing models to confirm how gloss annotations make SLT easier. We find that it can provide two aspects of information for the model, 1) it can help the model implicitly learn the location of semantic boundaries in continuous sign language videos, 2) it can help the model understand the sign language video globally. We then propose gloss attention, which enables the model to keep its attention within video segments that have the same semantics locally, just as gloss helps existing models do. Furthermore, we transfer the knowledge of sentence-to-sentence similarity from the natural language model to our gloss attention SLT network (GASLT) to help it understand sign language videos at the sentence level. Experimental results on multiple large-scale sign language datasets show that our proposed GASLT model significantly outperforms existing methods. Our code is provided in https://github.com/YinAoXiong/GASLT. Aoxiong Yin, Tianyun Zhong, Weike Jin, Tao Jin 0004, Zhou Zhao 0001 |
CVPR | 5 |
| 2023 | MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and RecognitionabstractMulti-media communications facilitate global interaction among people. However, despite researchers exploring cross-lingual translation techniques such as machine translation and audio speech translation to overcome language barriers, there is still a shortage of cross-lingual studies on visual speech. This lack of research is mainly due to the absence of datasets containing visual speech and translated text pairs. In this paper, we present AVMuST-TED, the first dataset for Audio-Visual Multilingual Speech Translation, derived from TED talks. Nonetheless, visual speech is not as distinguishable as audio speech, making it difficult to develop a mapping from source speech phonemes to the target language text. To address this issue, we propose MixSpeech, a cross-modality self-learning framework that utilizes audio speech to regularize the training of visual speech tasks. To further minimize the cross-modality gap and its impact on knowledge transfer, we suggest adopting mixed speech, which is created by interpolating audio and visual streams, along with a curriculum learning strategy to adjust the mixing ratio as needed. MixSpeech enhances speech translation in noisy environments, improving BLEU scores for four languages on AVMuST-TED by +1.4 to +4.2. Moreover, it achieves state-of-the-art performance in lip reading on CMLR (11.1%), LRS2 (25.5%), and LRS3 (28.0%). Xize Cheng, Tao Jin 0004, Rongjie Huang 0001, Linjun Li, Zehan Wang 0001, Ye Wang 0018, Huadai Liu, Aoxiong Yin, Zhou Zhao 0001 |
ICCV | 2 |
| 2023 | Exploring Group Video Captioning with Efficient Relational ApproximationabstractCurrent video captioning efforts most focus on describing a single video while the need for captioning videos in groups has increased considerably. In this study, we propose a new task, group video captioning, which aims to infer the desired content among a group of target videos and describe it with another group of related reference videos. This task requires the model to effectively summarize the target videos and accurately describe the distinguishing content compared to the reference videos, and it becomes more difficult as the video length increases. To solve this problem, 1) First, we propose an efficient relational approximation (ERA) to identify the shared content among videos while the complexity is linearly related to the number of videos. 2) Then, we introduce a contextual feature refinery with intra-group self-supervision to capture the contextual information and further refine the common properties. 3) In addition, we construct two group video captioning datasets derived from the YouCook2 and the ActivityNet Captions. The experimental results demonstrate the effectiveness of our method on this new task. Tao Jin 0004, Ye Wang 0018, Wenwen Pan 0003, Linjun Li, Xize Cheng, Zhou Zhao 0001 |
ICCV | 2 |
| 2023 | Rethinking Missing Modality Learning from a Decoding PerspectiveabstractConventional pipeline of multimodal learning consists of three stages, including encoding, fusion, and decoding. Most existing methods under missing modality condition focus on the first stage and aim to learn the modality invariant representation or reconstruct missing features. However, these methods rely on strong assumptions (i.e., all the pre-defined modalities are available for each input sample during training and the number of modalities is fixed). To solve this problem, we propose a simple yet effective method called Interaction Augmented Prototype Decomposition (IPD) for a more general setting, where the number of modalities is arbitrary and there are various incomplete modality conditions happening in both training and inference phases, even there are unseen testing conditions. Different from the previous methods, we improve the decoding stage. Concretely, IPD jointly learns the common and modality-specific task prototypes. Considering that the number of missing modality conditions scales exponentially with the number of modalities O(2n) and different conditions may have implicit interaction, the low-rank partial prototype decomposition with enough theoretical analysis is employed for modality-specific components to reduce the complexity. The decomposition also can promote unseen generalization with the modality factors of existing conditions. To simulate the low-rank setup, we further constrain the explicit interaction of specific modality conditions by employing disentangled contrastive constraints. Extensive results on the newly-created benchmarks of multiple tasks illustrate the effectiveness of our proposed model. Tao Jin 0004, Xize Cheng, Linjun Li, Ye Wang 0018, Zhou Zhao 0001 |
ACM Multimedia | 1 |
| 2023 | Electromagnetic Imaging Boosted Visual Object Recognition Under Difficult Visual ConditionsabstractObject imaging and recognition under difficult visual conditions is extremely challenging due to the captured low-quality images, and traditional optical-based recognition methods always fail in this task. In this paper, we propose to utilize the visual-microwave image pairs captured by both visual cameras and microwave sensors for imaging and recognition. To address the heavy noises in the low-quality optical images, we retrieve the physically quantitative images from associated scattered field data, and enhance visual features by both optical and retrieval images. We develop a cross-modal Enhanced Attentive Visual-Microwave Fusion (EAVMF) object recognition model to jointly learn the cross-modal generator and multimodal recognizer. In addition, an attention module for the visual subnetwork is utilized to highlight the regions of interest. Two multimodal datasets with synthetic visual-microwave image pairs are built to simulate the difficult visual condition. The numerical results on these datasets demonstrate that: 1) both the multimodal fusion, cross-modal enhancement, and visual attention module can enhance the performance; and 2) compared with existing methods, the proposed EAVMF not only performs better in terms of accuracy but also has good scalability and one-shot learning ability. Min Tan 0005, Tao Jin 0004, Danhui Ye, Kuiwen Xu, Xiaoling Gu, Jun Yu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | MC-SLT: Towards Low-Resource Signer-Adaptive Sign Language TranslationabstractOne of the challenging factors in real application of sign language translation (SLT) is inter-signer variation. With the assumption that the pre-trained translation model cannot cover all the signers, the adaptation capability for unseen signers is of great concern. In this paper, we take a completely different perspective for SLT, called signer-adaptive SLT, which mainly considers the transferable ability of SLT systems. To attack this challenging problem, we propose MC-SLT, a novel meta-learning framework that could exploit additional new-signer data via a support set, and output a signer-adaptive model via a few-gradient-step update. Considering the various degrees of style discrepancies of different words performed by multiple signers, we further devise diversity-aware meta-adaptive weights for the token-wise cross-entropy losses. Besides, to improve the training robustness, we adopt the self-guided curriculum learning scheme that first captures the global curricula from each signer to avoid falling into a bad local optimum early, and then learns the curricula of individualities to improve the model adaptability for learning signer-specific knowledge. We re-construct the existing standard datasets of SLT for the signer-adaptive setting and establish a new benchmark for subsequent research. Tao Jin 0004, Zhou Zhao 0001, Meng Zhang 0019, Xingshan Zeng |
ACM Multimedia | 1 |
| 2022 | Interaction augmented transformer with decoupled decoding for video captioning
Tao Jin 0004, Zhou Zhao 0001, Jun Yu 0002, Fei Wu 0001 |
Neurocomputing | 1 |
| 2021 | Contrastive Disentangled Meta-Learning for Signer-Independent Sign Language TranslationabstractSign language translation aims at directly translating a sign language video into a natural sentence. The majority of existing methods take the video-sentence pairs labeled by multiple specific signers as training and testing samples. However, such setting does not fit in with the real-world applications. A practicable sign language translation system is supposed to provide accurate translation results for unseen signers. In this paper, we mainly attack the signer-independent setting and focus on augmenting the generalization ability of translation model. To adapt to the challenging setting, we propose a novel framework called contrastive disentangled meta-learning (CDM), which develops several improvements in both deep architecture and training mode. Specifically, based on the minimax entropy objective, a disentangled module with adaptive gated units is developed to decouple the signer-specific and task-specific representation in the encoder. Besides, we facilitate the frame-word alignments by leveraging contrastive constraints between the obtained task-specific representation and the decoding output. The disentangled and contrastive modules could provide complementary information for each other. As for the training mode, we encourage the model to perform well in the simulated signer-independent scenarios by finding the generalized learning directions in the meta-learning process. Considering that vanilla meta-learning methods utilize the multiple specific signers insufficiently, we adopt a fine-grained learning strategy that simultaneously conducts meta-learning in a variety of domain shift scenarios in each iteration. Extensive experiments on the benchmark dataset RWTH-PHOENIX-Weather-2014T(PHOENIX14T) show that CDM could achieve competitive results compared with the state-of-the-art methods. Tao Jin 0004, Zhou Zhao 0001 |
ACM Multimedia | 1 |
| 2021 | Generalizable Multi-linear Attention NetworkabstractThe majority of existing multimodal sequential learning methods focus on how to obtain effective representations and ignore the importance of multimodal fusion. Bilinear attention network (BAN) is a commonly used fusion method, which leverages tensor operations to associate the features of different modalities. However, BAN has a poor compatibility for more modalities, since the computational complexity of the attention map increases exponentially with the number of modalities. Based on this concern, we propose a new method called generalizable multi-linear attention network (MAN), which can associate as many modalities as possible in linear complexity with hierarchical approximation decomposition (HAD). Besides, considering the fact that softmax attention kernels cannot be decomposed as linear operation directly, we adopt the addition random features (ARF) mechanism to approximate the non-linear softmax functions with enough theoretical analysis. We conduct extensive experiments on four datasets of three tasks (multimodal sentiment analysis, multimodal speaker traits recognition, and video retrieval), the experimental results show that MAN could achieve competitive results compared with the state-of-the-art methods, showcasing the effectiveness of the approximation decomposition and addition random features mechanism. Tao Jin 0004, Zhou Zhao 0001 |
NeurIPS | 1 |
| 2020 | SBAT: Video Captioning with Sparse Boundary-Aware TransformerabstractIn this paper, we focus on the problem of applying the transformer structure to video captioning effectively. The vanilla transformer is proposed for uni-modal language generation task such as machine translation. However, video captioning is a multimodal learning problem, and the video features have much redundancy between different time steps. Based on these concerns, we propose a novel method called sparse boundary-aware transformer (SBAT) to reduce the redundancy in video representation. SBAT employs boundary-aware pooling operation for scores from multihead attention and selects diverse features from different scenarios. Also, SBAT includes a local correlation scheme to compensate for the local information loss brought by sparse operation. Based on SBAT, we further propose an aligned cross-modal encoding scheme to boost the multimodal interaction. Experimental results on two benchmark datasets show that SBAT outperforms the state-of-the-art methods under most of the metrics. Tao Jin 0004, Siyu Huang, Yingming Li, Zhongfei Zhang |
IJCAI | 1 |
| 2019 | Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video CaptioningabstractTao Jin, Siyu Huang, Yingming Li, Zhongfei Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tao Jin 0004, Siyu Huang, Yingming Li, Zhongfei Zhang |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Recurrent convolutional video captioning with global and local attention
Tao Jin 0004, Yingming Li, Zhongfei Zhang |
Neurocomputing | 1 |