VLDB 2026 Research / reviewers in the wild / expert
Yapeng Tian
dblp:176/4020
· DBLP profile ↗
72ranked-venue papers
7as first author
61since 2021 · last 2026
0000-0003-1423-4513ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 7 first-author · 39 since 2021Artificial intelligence and machine learning · 46 · 6 first-author · 40 since 2021Human-computer interaction and ubiquitous computing · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Gaze Target Detection of Young Autistic ChildrenabstractThe automatic detection of gaze targets in autistic children through artificial intelligence can be impactful, especially for those who lack access to a sufficient number of professionals to improve their quality of life. This paper introduces a new, real-world AI application for gaze target detection in autistic children, which predicts a child's point of gaze from an activity image. This task is foundational for building automated systems that can measure joint attention—a core challenge in Autism Spectrum Disorder (ASD). To facilitate the study of this challenging application, we collected the first-ever Autism Gaze Target (AGT) Dataset. We further propose a novel social-aware coarse-to-fine (SACF) gaze detection framework that explicitly leverages the social context of a scene to overcome the class imbalance common in autism datasets—a consequence of autistic children's tendency to show reduced gaze to faces. It utilizes a two-pathway architecture with expert models specialized in social and non-social gaze, guided by a context-awareness gate module. The results of our comprehensive experiments demonstrate that our framework achieves new state-of-the-art performance for gaze target detection in this population, significantly outperforming existing methods, especially on the critical minority class of face-directed gaze. Shijian Deng, Erin E. Kosloski, Siva Sai Nagender Vasireddy, Randi Sierra Sherwood, Feroz Mohamed Hatha, Siddhi Patel, Pamela R. Rollins, Yapeng Tian |
AAAI | 9 |
| 2026 | Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?abstractUnlike traditional visual segmentation, audio-visual segmentation (AVS) requires the model not only to identify and segment objects but also to determine whether they are sound sources. Recent AVS approaches have achieved impressive performance on standard benchmarks. Yet, an important question remains: Do these models genuinely integrate audio-visual cues to segment sounding objects? Our study reveals a fundamental bias in current methods: they tend to generate segmentation masks based predominantly on visual salience, irrespective of the audio context, resulting in unreliable predictions when sounds are absent or irrelevant. To address this challenge, we introduce AVSBench-Robust, a comprehensive benchmark incorporating diverse negative audio scenarios, including silence, noise, and off-screen sounds. We also propose a simple yet effective approach combining balanced training with negative samples and classifier-guided similarity learning. Our extensive experiments show that while state-of-the-art AVS methods consistently fail under negative audio conditions, our approach achieves remarkable improvements in both standard metrics and robustness measures, maintaining near-perfect false positive rates while preserving high-quality segmentation performance. Ziru Huang, Yunhui Guo, Yapeng Tian |
AAAI | 5 |
| 2026 | Touch with Meaning: A Contextual Analysis of Social TouchabstractSocial touch is a rich channel of human communication, conveying emotion, intent, and meaning embedded in context. Yet most HCI studies treat touch in isolation, overlooking the layered subtleties that shape interpretation. We present a contextual analysis of 5,016 social touch events, grounded in a large collection of annotated scenes from films, dramas, and documentaries. Using a computer vision pipeline, we segmented touch events from video and annotated them across dimensions, including who is involved, how the gesture is performed, where on the body it occurs, and the cultural backdrop. Our analysis shows that identical gestures can convey distinct meanings depending on body location, relationship type, and context. Similar intentions—like comfort, encouragement, or dominance—may be expressed through different gestures or locations, shaped by relational dynamics, cultural norms, and public or private settings. These insights inform the design of socially aware touch technologies, including avatars, social agents, and mediated communication systems. Ayush Bhardwaj, Ashish Pratap, Abbas Khawaja, Yapeng Tian, Uison Ju, Dajin Lee, Seungmoon Choi, Jin Ryong Kim |
CHI | 4 |
| 2026 | High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
Chao Huang 0033, Susan Liang, Yapeng Tian, Chenliang Xu |
Int. J. Comput. Vis. | 3 |
| 2025 | CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIPabstractTianyu Yang, Lisen Dai, Xiangqi Wang, Minhao Cheng, Yapeng Tian, Xiangliang Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lisen Dai, Xiangqi Wang, Minhao Cheng, Yapeng Tian, Xiangliang Zhang 0001 |
ACL (1) | 5 |
| 2025 | Vision Token Reduction via Attention-Driven Self-Compression for Efficient Multimodal Large Language Models
Omer Faruk Deniz, Ruiyu Mao, Yapeng Tian, Latifur Khan |
IEEE Big Data | 4 |
| 2025 | Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel LevelabstractWe introduce Motion-Grounded Video Reasoning, a new motion understanding task that requires generating visual answers (video segmentation masks) according to the input question, and hence needs implicit spatiotemporal reasoning and grounding. It extends existing spatiotemporal grounding work focusing on explicit action/motion grounding, to a more general format by enabling implicit reasoning via questions. To facilitate the development of the new task, we collect a large-scale dataset called GroundMoRe, which comprises 1,715 video clips, 249K object masks that are deliberately designed with 4 question types for benchmarking deep and comprehensive motion reasoning abilities. GroundMoRe uniquely requires models to generate visual answers, providing a more concrete and visually interpretable response than plain texts. It evaluates models on both spatiotemporal grounding and reasoning, fostering to address complex challenges in motion-related video reasoning, temporal perception, and pixel-level understanding. Furthermore, we introduce a novel baseline model named MoRA, which achieves respectable performance on GroundMoRe outperforming the best existing visual grounding baseline model by an average of 21.5% relatively. We hope this novel and challenging task will pave the way for future advancements in robust and general motion understanding via video reasoning segmentation. Project available at: https://groundmore.github.io/ Andong Deng, Tongjia Chen, Shoubin Yu, Taojiannan Yang, Lincoln Spencer, Yapeng Tian, Ajmal Mian, Mohit Bansal, Chen Chen 0001 |
CVPR | 6 |
| 2025 | VinTAGe: Joint Video and Text Conditioning for Holistic Audio GenerationabstractRecent advances in audio generation have focused on text-to-audio (T2A) and video-to-audio (V2A) tasks. However, T2A or V2A methods cannot generate holistic sounds (onscreen and off-screen). This is because T2A cannot generate sounds aligning with onscreen objects, while V2A cannot generate semantically complete (offscreen sounds missing). In this work, we address the task of holistic audio generation: given a video and a text prompt, we aim to generate both onscreen and offscreen sounds that are temporally synchronized with the video and semantically aligned with text and video. Previous approaches for joint text and video-to-audio generation often suffer from modality bias, favoring one modality over the other. To overcome this limitation, we introduce VinTAGe, a flow-based transformer model that jointly considers text and video to guide audio generation. Our framework comprises two key components: a Visual-Text Encoder and a Joint VT-SiT model. To reduce modality bias and improve generation quality, we employ pretrained uni-modal text-to-audio and video-to-audio generation models for additional guidance. Due to the lack of appropriate benchmarks, we also introduce VinTAGeBench, a dataset of 636 video-text-audio pairs containing both onscreen and offscreen sounds. Our comprehensive experiments on VinTAGe-Bench demonstrate that joint text and visual interaction is necessary for holistic audio generation. Furthermore, VinTAGe achieves state-of-the-art results on the VGGSound benchmark. We will release our pretrained models and the VinTAGe-Bench dataset to facilitate future research in this exciting field. Saksham Singh Kushwaha, Yapeng Tian |
CVPR | 2 |
| 2025 | SignDiff: Diffusion Model for American Sign Language ProductionabstractIn this paper, we propose a dual-condition diffusion pre-training model named SignDIFF that can generate human sign language speakers from a skeleton pose. SignDiff has a novel Frame Reinforcement Network called FR-Net, similar to dense human pose estimation work, which enhances the correspondence between text lexical symbols and sign language dense pose frames, reduces the occurrence of multiple fingers in the diffusion model. In addition, we propose a new method for American Sign Language Production (ASLP), which can generate ASL skeletal pose videos from text input, integrating two new improved modules and a new loss function to improve the accuracy and quality of sign language skeletal posture and enhance the ability of the model to train on largescale data. We propose a simple baseline for ASL production and report the scores of 17.19 and 12.85 on BLEU-4 on the How2Sign dev/test sets. We evaluated our model on the previous mainstream dataset PHOENIX14T, and our method achieved the SOTA results. In addition, our image quality far exceeds all previous results by 10 percentage points in terms of SSIM. Sen Fang, Chunyu Sui, Yanghao Zhou, Hongbin Zhong, Yapeng Tian, Chen Chen 0001 |
FG | 6 |
| 2025 | Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion ModelsabstractSpatial audio is a crucial component in creating immersive experiences. Traditional simulation-based approaches to generate spatial audio rely on expertise, have limited scalability, and assume independence between semantic and spatial information. To address these issues, we explore end-to-end spatial audio generation. We introduce and formulate a new task of generating first-order Ambisonics (FOA) given a sound category and sound source spatial location. We propose Diff-SAGe, an end-to-end, flow-based diffusion-transformer model for this task. Diff-SAGe utilizes a complex spectrogram representation for FOA, preserving the phase information crucial for accurate spatial cues. Additionally, a multi-conditional encoder integrates the input conditions into a unified representation, guiding the generation of FOA waveforms from noise. Through extensive evaluations on two datasets, we demonstrate that our method consistently outperforms traditional simulation-based baselines across both objective and subjective metrics. Saksham Singh Kushwaha, Mark R. P. Thomas, Yapeng Tian, Avery Bruni |
ICASSP | 4 |
| 2025 | ZFusion: Efficient Deep Compositional Zero-Shot Learning for Blind Image Super-Resolution with Generative Diffusion Prior
Alireza Esmaeilzehi, Hossein Zaredar, Yapeng Tian, Laleh Seyyed-Kalantari |
ICCV | 3 |
| 2025 | PRVQL: Progressive Knowledge-Guided Refinement for Robust Egocentric Visual Query LocalizationabstractEgocentric visual query localization (EgoVQL) focuses on localizing the target of interest in space and time from first-person videos, given a visual query. Despite recent progressive, existing methods often struggle to handle severe object appearance changes and cluttering background in the video due to lacking sufficient target cues, leading to degradation. Addressing this, we introduce PRVQL, a novel Progressive knowledge-guided Refinement framework for EgoVQL. The core is to continuously exploit target-relevant knowledge directly from videos and utilize it as guidance to refine both query and video features for improving target localization. Our PRVQL contains multiple processing stages. The target knowledge from one stage, comprising appearance and spatial knowledge extracted via two specially designed knowledge learning modules, are utilized as guidance to refine the query and videos features for the next stage, which are used to generate more accurate knowledge for further feature refinement. With such a progressive process, target knowledge in PRVQL can be gradually improved, which, in turn, leads to better refined query and video features for localization in the final stage. Compared to previous methods, our PRVQL, besides the given object cues, enjoys additional crucial target information from a video as guidance to refine features, and hence enhances EgoVQL in complicated scenes. In our experiments on challenging Ego4D, PRVQL achieves state-of-the-art result and largely surpasses other methods, showing its efficacy. Our code, model and results will be released at https://github.com/fb-reps/PRVQL. Bing Fan, Yunhe Feng, Yapeng Tian, James Liang, Yuewei Lin, Yan Huang 0002, Heng Fan 0001 |
ICCV | 3 |
| 2025 | AV-DiT: Taming Image Diffusion Transformers for Efficient Joint Audio and Video GenerationabstractRecent Diffusion Transformers (DiTs) have shown impressive capabilities in generating single-modality content, including images, videos, and audio. However, the potential of DiTs to enable superb multimodal content creation remains underexplored. To bridge this gap, we introduce AV-DiT, a novel and efficient audio-visual diffusion transformer designed to generate high-quality, realistic videos with synchronized audio tracks. To minimize model complexity and computational costs, our AV-DiT utilizes a modality-shared DiT backbone pre-trained on image-only data, with only newly inserted adapters being trainable. This shared backbone facilitates the generation of both audio and video. Specifically, the video branch incorporates a trainable temporal attention layer into a pre-trained DiT block for capturing the temporal consistency for video generation. In addition, a small number of trainable parameters adapt the image-based DiT block to learn the acoustic characteristics for audio generation. An extra shared self-attention block reused from the DiT block, equipped with lightweight parameters, facilitates feature interaction between audio and visual modalities for alignment. Extensive experiments on the datasets demonstrate that our AV-DiT achieves state-of-the-art performance in joint audio-visual generation with significantly fewer tunable parameters. Furthermore, our results highlight that a single shared image generative backbone with modality-specific adaptations is sufficient for constructing a joint audio-video generator. Kai Wang 0012, Shijian Deng, Jing Shi 0005, Dimitrios Hatzinakos, Yapeng Tian |
ACM Multimedia | 5 |
| 2025 | <tt>AVROBUSTBENCH</tt>: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time
Sarthak Kumar Maharana, Saksham Singh Kushwaha, Baoming Zhang, Adrian Rodriguez, Songtao Wei, Yapeng Tian, Yunhui Guo |
NeurIPS | 6 |
| 2025 | Language-Guided Adaptive Vision Token Pruning for Efficient Multimodal Large Language Models
Omer Faruk Deniz, Tarik Arici, Fatemeh Sheikholeslami, Burak Gozluklu, Ameni Trabelsi, Suleiman Ali Khan, Yapeng Tian, Latifur Khan |
PAKDD (5) | 7 |
| 2025 | VRSight: An AI-Driven Scene Description System to Improve Virtual Reality Accessibility for Blind People
Daniel Killough, Justin Feng, Zheng Xue "ZX" Ching, Rithvik Dyava, Yapeng Tian, Yuhang Zhao 0001 |
UIST | 6 |
| 2025 | AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multimodal Information Between Reality and VideosabstractVideos offer rich audiovisual information that can support people in performing activities of daily living (ADLs), but they remain largely inaccessible to blind or low-vision (BLV) individuals.In cooking, BLV people often rely on non-visual cues-such as touch, taste, and smell-to navigate their environment, making it difficult to follow UIST '25, September 28-October 01, 2025, Busan, Republic of Korea Ning et al.the predominantly audiovisual instructions found in video recipes.To address this problem, we introduce Aroma, an AI system that provides timely responses to the user based on real-time, contextaware assistance by integrating non-visual cues perceived by the user, a wearable camera feed, and video recipe content.Aroma uses a mixed-initiative approach: it responds to user requests while also proactively monitoring the video stream to offer timely alerts and guidance.This collaborative design leverages the complementary strengths of the user and AI system to align the physical environment with the video recipe, helping the user interpret their current state and make sense of the steps.We evaluated Aroma through a study with eight BLV participants and offered insights for designing interactive AI systems to support BLV individuals in performing ADLs. Zheng Ning, Leyang Li, Daniel Killough, JooYoung Seo, Patrick Carrington, Yapeng Tian, Yuhang Zhao 0001, Franklin Mingzhe Li, Toby Jia-Jun Li |
UIST | 6 |
| 2025 | Joint Co-Speech Gesture and Expressive Talking Face Generation Using Diffusion with AdaptersabstractRecent advances in co-speech gesture and talking head generation have been impressive, yet most methods focus on only one of the two tasks. Those that attempt to generate both often rely on separate models or network modules, increasing training complexity and ignoring the inherent relationship between face and body movements. To address the challenges, in this paper, we propose a novel model architecture that jointly generates face and body motions within a single network. This approach leverages shared weights between modalities, facilitated by adapters that enable adaptation to a common latent space. Our experiments demonstrate that the proposed framework not only maintains state-of-the-art co-speech gesture and talking head generation performance but also significantly reduces the number of parameters required. Steven Hogue, Yapeng Tian, Xiaohu Guo |
WACV | 3 |
| 2025 | MagicTalk: Implicit and Explicit Correlation Learning for Diffusion-Based Emotional Talking Face GenerationabstractGenerating emotional talking faces from a single portrait image remains a significant challenge. The simultaneous achievement of expressive emotional talking and accurate lip-sync is particularly difficult, as expressiveness is often compromised for lip-sync accuracy. Prevailing generative works usually struggle to juggle to generate subtle variations of emotional expression and lip-synchronized talking. To address these challenges, we suggest modeling the implicit and explicit correlations between audio and emotional talking faces with a unified framework. As human emotional expressions usually present subtle and implicit relations with speech audio, we propose incorporating audio and emotional style embeddings into the diffusion-based generation process, for realistic generation while concentrating on emotional expressions. We then propose lip-based explicit correlation learning to construct a strong mapping of audio to lip motions, assuring lip-audio synchronization. Furthermore, we deploy a video-to-video rendering module to transfer expressions and lip motions from a proxy 3D avatar to an arbitrary portrait. Both quantitatively and qualitatively, MagicTalk outperforms state-of-the-art methods in terms of expressiveness, lip-sync, and perceptual quality. Chao Wang 0088, Guoxian Song, You Xie, Linjie Luo, Yapeng Tian, Jiashi Feng, Xiaohu Guo |
Comput. Vis. Media | 8 |
| 2025 | DiffI2I: Efficient Diffusion Model for Image-to-Image TranslationabstractThe Diffusion Model (DM) has emerged as the SOTA approach for image synthesis. However, the existing DM cannot perform well on some image-to-image translation (I2I) tasks. Different from image synthesis, some I2I tasks, such as super-resolution, require generating results in accordance with GT images. Traditional DMs for image synthesis require extensive iterations and large denoising models to estimate entire images, which gives their strong generative ability but also leads to artifacts and inefficiency for I2I. To tackle this challenge, we propose a simple, efficient, and powerful DM framework for I2I, called DiffI2I. Specifically, DiffI2I comprises three key components: a compact I2I prior extraction network (CPEN), a dynamic I2I transformer (DI2Iformer), and a denoising network. We train DiffI2I in two stages: pretraining and DM training. For pretraining, GT and input images are fed into CPEN to capture a compact I2I prior representation (IPR) guiding DI2Iformer. In the second stage, the DM is trained to only use the input images to estimate the same IRP as CPEN. Compared to traditional DMs, the compact IPR enables DiffI2I to obtain more accurate outcomes and employ a lighter denoising network and fewer iterations. Through extensive experiments on various I2I tasks, we demonstrate that DiffI2I achieves SOTA performance while significantly reducing computational burdens. Bin Xia 0014, Yulun Zhang 0001, Shiyin Wang, Yapeng Tian, Wenming Yang, Radu Timofte, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Hear Me, See Me, Understand Me: Audio-Visual Autism Behavior RecognitionabstractIn this article, we introduce a novel problem of audio-visual autism behavior recognition, which includes social behavior recognition, an essential aspect previously omitted in AI-assisted autism screening research. We define the task at hand as one that is audio-visual autism behavior recognition, which uses audio and visual cues, including any speech present in the audio, to recognize autism-related behaviors. To facilitate this new research direction, we collected an audio-visual autism spectrum dataset (AV-ASD), currently the largest video dataset for autism screening using a behavioral approach. It covers an extensive range of autism-associated behaviors, including those related to social communication and interaction. To pave the way for further research on this new problem, we intensively explored leveraging foundation models and multimodal large language models across different modalities. Our experiments on the AV-ASD dataset demonstrate that integrating audio (mainly ambient sound), visual, and speech (predominately spoken language) modalities significantly enhances the performance in autism behavior recognition. Additionally, we explored the use of apost-hoc to ad-hocpipeline in a multimodal large language model to investigate its potential to augment the model's explanatory capability during autism behavior recognition. Shijian Deng, Erin E. Kosloski, Siddhi Patel, Zeke A. Barnett, Yiyang Nan, Alexander Kaplan, Sisira Aarukapalli, William T. Doan, Matthew Wang, Harsh Singh, Pamela R. Rollins, Yapeng Tian |
IEEE Trans. Multim. | 12 |
| 2024 | High-Quality Visually-Guided Sound Separation from Diverse Categories
Chao Huang 0033, Susan Liang, Yapeng Tian, Chenliang Xu |
ACCV (6) | 3 |
| 2024 | Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation
Susan Liang, Chao Huang 0033, Yapeng Tian, Chenliang Xu |
ACCV (6) | 3 |
| 2024 | Benchmarking and Optimizing Federated Learning with Hardware-related Metrics
Kai Pan, Yapeng Tian, Yinhe Han 0001, Yiming Gan |
BMVC | 2 |
| 2024 | MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on VideosabstractSpatial audio offers more immersive video consumption experiences to viewers; however, creating and editing spatial audio often expensive and requires specialized hardware equipment and skills, posing a high barrier for amateur video creators. We present Mimosa, a human-AI co-creation tool that enables amateur users to computationally generate and manipulate spatial audio effects. For a video with only monaural or stereo audio, Mimosa automatically grounds each sound source to the corresponding sounding object in the visual scene and enables users to further validate and fix errors in the location of the sounding objects. Users can also augment the spatial audio effect by flexibly manipulating the sounding source positions and creatively customizing the audio effect. The design of Mimosa exemplifies a human-AI collaboration approach that, instead of utilizing state-of-art end-to-end “black-box” ML models, uses a multistep pipeline that aligns its interpretable intermediate results with the user’s workflow. A lab user study with 15 participants demonstrates Mimosa’s usability, usefulness, expressiveness, and capability in creating immersive spatial audio effects in collaboration with users. Zheng Ning, Zheng Zhang 0043, Jerrick Ban, Ruohong Gan, Yapeng Tian, Toby Jia-Jun Li |
Creativity & Cognition | 6 |
| 2024 | SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersabstractBlind or Low-Vision (BLV) users often rely on audio descriptions (AD) to access video content. However, conventional static ADs can leave out detailed information in videos, impose a high mental load, neglect the diverse needs and preferences of BLV users, and lack immersion. To tackle these challenges, we introduce Spica, an AI-powered system that enables BLV users to interactively explore video content. Informed by prior empirical studies on BLV video consumption, Spica offers interactive mechanisms for supporting temporal navigation of frame captions and spatial exploration of objects within key frames. Leveraging an audio-visual machine learning pipeline, Spica augments existing ADs by adding interactivity, spatial sound effects, and individual object descriptions without requiring additional human annotation. Through a user study with 14 BLV participants, we evaluated the usability and usefulness of Spica and explored user behaviors, preferences, and mental models when interacting with augmented ADs. Zheng Ning, Brianna L. Wimer, Keyi Chen 0008, Jerrick Ban, Yapeng Tian, Yuhang Zhao 0001, Toby Jia-Jun Li |
CHI | 6 |
| 2024 | T-VSL: Text-Guided Visual Sound Source Localization in MixturesabstractVisual sound source localization poses a significant chal-lenge in identifying the semantic region of each sounding source within a video. Existing self-supervised and weakly supervised source localization methods struggle to accu-rately distinguish the semantic regions of each sounding object, particularly in multi-source mixtures. These methods often rely on audio-visual correspondence as guidance, which can lead to substantial performance drops in com-plex multi-source localization scenarios. The lack of access to individual source sounds in multi-source mixtures during training exacerbates the difficulty of learning effective audio-visual correspondence for localization. To ad-dress this limitation, in this paper, we propose incorpo-rating the text modality as an intermediate feature guide using tri-modal joint embedding models (e.g., Audio Clip) to disentangle the semantic audio-visual source correspon-dence in multi-source mixtures. Our framework, dubbed T-VSL, begins by predicting the class of sounding enti-ties in mixtures. Subsequently, the textual representation of each sounding source is employed as guidance to dis-entangle fine-grained audio-visual source correspondence from multi-source mixtures, leveraging the tri-modal Audio-CLIP embedding. This approach enables our framework to handle a flexible number of sources and exhibits promising zero-shot transferability to unseen classes during test time. Extensive experiments conducted on the MUSIC, VG-GSound, and VGGSound-Instruments datasets demonstrate significant performance improvements over state-of-the-art methods. Code is released at https://github.com/enyac-group/T-VSL/tree/main. Tanvir Mahmud, Yapeng Tian, Diana Marculescu |
CVPR | 2 |
| 2024 | Continual Audio-Visual Sound SeparationabstractIn this paper, we introduce a novel continual audio-visual sound separation task, aiming to continuously separate sound sources for new classes while preserving performance on previously learned classes, with the aid of visual guidance. This problem is crucial for practical visually guided auditory perception as it can significantly enhance the adaptability and robustness of audio-visual sound separation models, making them more applicable for real-world scenarios where encountering new sound sources is commonplace. The task is inherently challenging as our models must not only effectively utilize information from both modalities in current tasks but also preserve their cross-modal association in old tasks to mitigate catastrophic forgetting during audio-visual continual learning. To address these challenges, we propose a novel approach named ContAV-Sep ($\textbf{Cont}$inual $\textbf{A}$udio-$\textbf{V}$isual Sound $\textbf{Sep}$aration). ContAV-Sep presents a novel Cross-modal Similarity Distillation Constraint (CrossSDC) to uphold the cross-modal semantic similarity through incremental tasks and retain previously acquired knowledge of semantic similarity in old models, mitigating the risk of catastrophic forgetting. The CrossSDC can seamlessly integrate into the training process of different audio-visual sound separation frameworks. Experiments demonstrate that ContAV-Sep can effectively mitigate catastrophic forgetting and achieve significantly better performance compared to other continual learning baselines for audio-visual sound separation. Code is available at: https://github.com/weiguoPian/ContAV-Sep_NeurIPS2024. Weiguo Pian, Yiyang Nan, Shijian Deng, Shentong Mo, Yunhui Guo, Yapeng Tian |
NeurIPS | 6 |
| 2024 | CookAR: Affordance Augmentations in Wearable AR to Support Kitchen Tool Interactions for People with Low VisionabstractCooking is a central activity of daily living, supporting independence as well as mental and physical health. However, prior work has highlighted key barriers for people with low vision (LV) to cook, particularly around safely interacting with tools, such as sharp knives or hot pans. Drawing on recent advancements in computer vision (CV), we present CookAR, a head-mounted AR system with real-time object affordance augmentations to support safe and efficient interactions with kitchen tools. To design and implement CookAR, we collected and annotated the first egocentric dataset of kitchen tool affordances, fine-tuned an affordance segmentation model, and developed an AR system with a stereo camera to generate visual augmentations. To validate CookAR, we conducted a technical evaluation of our fine-tuned model as well as a qualitative lab study with 10 LV participants for suitable augmentation design. Our technical evaluation demonstrates that our model outperforms the baseline on our tool affordance dataset, while our user study indicates a preference for affordance augmentations over the traditional whole object augmentations. Jaewook Lee 0005, Andrew D. Tjahjadi, Junpu Yu, Minji Park, Jon Froehlich, Yapeng Tian, Yuhang Zhao 0001 |
UIST | 8 |
| 2024 | LAVSS: Location-Guided Audio-Visual Spatial Audio SeparationabstractExisting machine learning research has achieved promising results in monaural audio-visual separation (MAVS). However, most MAVS methods purely consider what the sound source is, not where it is located. This can be a problem in VR/AR scenarios, where listeners need to be able to distinguish between similar audio sources located in different directions. To address this limitation, we have generalized MAVS to spatial audio separation and proposed LAVSS: a location-guided audio-visual spatial audio separator. LAVSS is inspired by the correlation between spatial audio and visual location. We introduce the phase difference carried by binaural audio as spatial cues, and we utilize positional representations of sounding objects as additional modality guidance. We also leverage multi-level cross-modal attention to perform visual-positional collaboration with audio features. In addition, we adopt a pre-trained monaural separator to transfer knowledge from rich mono sounds to boost spatial audio separation. This exploits the correlation between monaural and binaural channels. Experiments on the FAIR-Play dataset demonstrate the superiority of the proposed LAVSS over existing benchmarks of audio-visual separation. Our project page: https://yyx666660.github.io/LAVSS/. Wenming Yang, Yapeng Tian |
WACV | 3 |
| 2024 | STADNet: Spatial-Temporal Attention-Guided Dual-Path Network for cardiac cine MRI super-resolution
Shuo Wang 0011, Yapeng Tian, Shunjie Dong, Chengyan Wang, Angelica I. Avilés-Rivero, Harry Qin |
Medical Image Anal. | 3 |
| 2024 | EgoVSR: Toward High-Quality Egocentric Video Super-ResolutionabstractDue to the limitations of capture devices and scenarios, egocentric videos frequently have low visual quality, mainly caused by high compression and severe motion blur. With the increasing application of egocentric videos, there is an urgent need to enhance the quality of these videos through super-resolution. However, existing Video Super-Resolution (VSR) works, focusing on third-person view videos, are actually unsuitable for handling blurring artifacts caused by rapid ego-motion and object motion in egocentric videos. To this end, we propose EgoVSR, a VSR framework specifically designed for egocentric videos. We explicitly tackle motion blurs in egocentric videos using a Dual Branch Deblur Network (DB2Net) in the VSR framework. Meanwhile, a blurring mask is introduced to guide the DB2Net learning, and can be used to localize blurred areas in video frames. We also design a MaskNet to predict the mask, as well as a mask loss to optimize the mask estimation. Additionally, an online motion blur synthesis model for common VSR training data is proposed to simulate motion blurs as in egocentric videos. In order to validate the effectiveness of our proposed method, we introduce an EgoVSR dataset containing a large amount of fast-motion egocentric video sequences. Extensive experiments demonstrate that our EgoVSR model can efficiently super-resolve low-quality egocentric videos and outperform strong comparison baselines. Our code, pre-trained models and data can be found athttps://github.com/chiyich/EGOVSR/. Yichen Chi, Junhao Gu, Jiamiao Zhang, Wenming Yang, Yapeng Tian |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Cross Modality Bias in Visual Question Answering: A Causal View With Possible Worlds VQAabstractTo increase the generalization capability of VQA systems, many recent studies have tried to de-bias spurious language or vision associations that shortcut the question or image to the answer. Despite these efforts, the literature fails to address the confounding effect of vision and language simultaneously. As a result, when they reduce bias learned from one modality, they usually increase bias from the other. In this paper, we first model a confounding effect that causes language and vision bias simultaneously, then propose a counterfactual inference to remove the influence of this effect. The model trained in this strategy can concurrently and efficiently reduce vision and language bias. To the best of our knowledge, this is the first work to reduce biases resulting from confounding effects of vision and language in VQA, leveraging causal explain-away relations. We accompany our method with an explain-away strategy, pushing the accuracy of the questions with numerical answers results compared to existing methods that have been an open problem. The proposed method outperforms the state-of-the-art methods in VQA-CP v2 datasets. R2: Providing brief insights into the experimental setup and results would add valuable context for readers. In response to R2, we released the code and documentation for the implementation as follows. Our codes are available at https://github.com/ali-vosoughi/PW-VQA. Ali Vosoughi, Shijian Deng, Songyang Zhang 0004, Yapeng Tian, Chenliang Xu, Jiebo Luo 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | STDAN: Deformable Attention Network for Space-Time Video Super-ResolutionabstractThe target of space-time video super-resolution (STVSR) is to increase the spatial-temporal resolution of low-resolution (LR) and low-frame-rate (LFR) videos. Recent approaches based on deep learning have made significant improvements, but most of them only use two adjacent frames, that is, short-term features, to synthesize the missing frame embedding, which cannot fully explore the information flow of consecutive input LR frames. In addition, existing STVSR models hardly exploit the temporal contexts explicitly to assist high-resolution (HR) frame reconstruction. To address these issues, in this article, we propose a deformable attention network called STDAN for STVSR. First, we devise a long short-term feature interpolation (LSTFI) module that is capable of excavating abundant content from more neighboring input frames for the interpolation process through a bidirectional recurrent neural network (RNN) structure. Second, we put forward a spatial-temporal deformable feature aggregation (STDFA) module, in which spatial and temporal contexts in dynamic video frames are adaptively captured and aggregated to enhance SR reconstruction. Experimental results on several datasets demonstrate that our approach outperforms state-of-the-art STVSR methods. The code is available at https://github.com/littlewhitesea/STDAN. Hai Wang 0020, Xiaoyu Xiang, Yapeng Tian, Wenming Yang, Qingmin Liao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Towards Unified, Explainable, and Robust Multisensory PerceptionabstractHumans perceive surrounding scenes through multiple senses with multisensory integration. For example, hearing helps capture the spatial location of a racing car behind us; seeing peoples' talking faces can strengthen our perception of their speech. However, today's state-of-the-art scene understanding systems are usually designed to rely on a single audio or visual modality. Ignoring multisensory cooperation has become one of the key bottlenecks in creating intelligent systems with human-level perception capability, which impedes the real-world applications of existing scene understanding models. To address this limitation, my research has pioneered marrying computer vision with computer audition to create multimodal systems that can learn to understand audio and visual data. In particular, my current research focuses on asking and solving fundamental problems in a fresh research area: audio-visual scene understanding and strives to develop unified, explainable, and robust multisensory perception machines. The three themes are distinct yet interconnected, and all of them are essential for designing powerful and trustworthy perception systems. In my talk, I will give a brief overview about this new research area and then introduce my works in the three research thrusts. Yapeng Tian |
AAAI | 1 |
| 2023 | Egocentric Audio-Visual Object LocalizationabstractHumans naturally perceive surrounding scenes by unifying sound and sight from a first-person view. Likewise, machines are advanced to approach human intelligence by learning with multisensory inputs from an egocentric perspective. In this paper, we explore the challenging egocentric audio-visual object localization task and observe that 1) egomotion commonly exists in first-person recordings, even within a short duration; 2) The out-of-view sound components can be created when wearers shift their attention. To address the first problem, we propose a geometryaware temporal aggregation module that handles the egomotion explicitly. The effect of egomotion is mitigated by estimating the temporal geometry transformation and exploiting it to update visual representations. Moreover, we propose a cascaded feature enhancement module to over-come the second issue. It improves cross-modal localization robustness by disentangling visually-indicated audio representation. During training, we take advantage of the naturally occurring audio-visual temporal synchronization as the “free” self-supervision to avoid costly labeling. We also annotate and create the Epic Sounding Object dataset for evaluation purposes. Extensive experiments show that our method achieves state-of-the-art localization performance in egocentric videos and can be generalized to diverse audio-visual scenes. Code is available at https://github.com/WikiChao/Ego-AV-Loc. Chao Huang 0033, Yapeng Tian, Chenliang Xu |
CVPR | 2 |
| 2023 | Audio-Visual Grouping Network for Sound Localization from MixturesabstractSound source localization is a typical and challenging task that predicts the location of sound sources in a video. Previous single-source methods mainly used the audio-visual association as clues to localize sounding objects in each image. Due to the mixed property of multiple sound sources in the original space, there exist rare multi-source approaches to localizing multiple sources simultaneously, except for one recent work using a contrastive random walk in the graph with images and separated sound as nodes. Despite their promising performance, they can only handle a fixed number of sources, and they cannot learn compact class-aware representations for individual sources. To alleviate this shortcoming, in this paper, we propose a novel audio-visual grouping network, namely AVGN, that can directly learn category-wise semantic features for each source from the input audio mixture and image to localize multiple sources simultaneously. Specifically, our AVGN leverages learnable audio-visual class tokens to aggregate class-aware source features. Then, the aggregated semantic features for each source can be used as guidance to localize the corresponding visual regions. Compared to existing multi-source methods, our new framework can localize a flexible number of sources and disentangle category-aware audio-visual representations for individual sound sources. We conduct extensive experiments on MUSIC, VGGSound-Instruments, and VGG-Sound Sources benchmarks. The results demonstrate that the proposed AVGN can achieve state-of-the-art sounding object localization performance on both single-source and multi-source scenarios. Code is available at https://github.com/stoneMo/AVGN. Shentong Mo, Yapeng Tian |
CVPR | 2 |
| 2023 | Structured Sparsity Learning for Efficient Video Super-ResolutionabstractThe high computational costs of video super-resolution (VSR) models hinder their deployment on resource-limited devices, e.g., smartphones and drones. Existing VSR models contain considerable redundant filters, which drag down the inference efficiency. To prune these unimportant filters, we develop a structured pruning scheme called Structured Sparsity Learning (SSL) according to the properties of VSR. In SSL, we design pruning schemes for several key components in VSR models, including residual blocks, recurrent networks, and upsampling networks. Specifically, we develop a Residual Sparsity Connection (RSC) scheme for residual blocks of recurrent networks to liberate pruning restrictions and preserve the restoration information. For upsampling networks, we design a pixel-shuffle pruning scheme to guarantee the accuracy of feature channel-space conversion. In addition, we observe that pruning error would be amplified as the hidden states propagate along with recurrent networks. To alleviate the issue, we design Temporal Finetuning (TF). Extensive experiments show that SSL can significantly outperform recent methods quantitatively and qualitatively. The code is available at https://github.com/Zj-BinXia/SSL. Bin Xia 0014, Jingwen He, Yulun Zhang 0001, Yapeng Tian, Wenming Yang, Luc Van Gool |
CVPR | 5 |
| 2023 | Class-Incremental Grouping Network for Continual Audio-Visual LearningabstractContinual learning is a challenging problem in which models need to be trained on non-stationary data across sequential tasks for class-incremental learning. While previous methods have focused on using either regularization or rehearsal-based frameworks to alleviate catastrophic forgetting in image classification, they are limited to a single modality and cannot learn compact class-aware cross-modal representations for continual audio-visual learning. To address this gap, we propose a novel class-incremental grouping network (CIGN) that can learn category-wise semantic features to achieve continual audio-visual learning. Our CIGN leverages learnable audio-visual class tokens and audio-visual grouping to continually aggregate class-aware features. Additionally, it utilizes class tokens distillation and continual grouping to prevent forgetting parameters learned from previous tasks, thereby improving the model’s ability to capture discriminative audiovisual categories. We conduct extensive experiments on VGG-Sound-Instruments, VGGSound-100, and VGG-Sound Sources benchmarks. Our experimental results demonstrate that the CIGN achieves state-of-the-art audio-visual class-incremental learning performance. Code is available at https://github.com/stoneMo/CIGN. Shentong Mo, Weiguo Pian, Yapeng Tian |
ICCV | 3 |
| 2023 | Audio-Visual Class-Incremental LearningabstractIn this paper, we introduce audio-visual class-incremental learning, a class-incremental learning scenario for audio-visual video recognition. We demonstrate that joint audio-visual modeling can improve class-incremental learning, but current methods fail to preserve semantic similarity between audio and visual features as incremental step grows. Furthermore, we observe that audio-visual correlations learned in previous tasks can be forgotten as incremental steps progress, leading to poor performance. To overcome these challenges, we propose AV-CIL, which incorporates Dual-Audio-Visual Similarity Constraint (D-AVSC) to maintain both instance-aware and class-aware semantic similarity between audio-visual modalities and Visual Attention Distillation (VAD) to retain previously learned audio-guided visual attentive ability. We create three audio-visual class-incremental datasets, AVE-Class-Incremental (AVE-CI), Kinetics-Sounds-Class-Incremental (K-S-CI), and VGGSound100-Class-Incremental (VS100-CI) based on the AVE, Kinetics-Sounds, and VGGSound datasets, respectively. Our experiments on AVE-CI, K-SCI, and VS100-CI demonstrate that AV-CIL significantly outperforms existing class-incremental learning methods in audio-visual class-incremental learning. Code and data are available at: https://github.com/weiguoPian/AV-CIL_ICCV2023. Weiguo Pian, Shentong Mo, Yunhui Guo, Yapeng Tian |
ICCV | 4 |
| 2023 | DiffIR: Efficient Diffusion Model for Image RestorationabstractDiffusion model (DM) has achieved SOTA performance by modeling the image synthesis process into a sequential application of a denoising network. However, different from image synthesis, image restoration (IR) has a strong constraint to generate results in accordance with ground-truth. Thus, for IR, traditional DMs running massive iterations on a large model to estimate whole images or feature maps is inefficient. To address this issue, we propose an efficient DM for IR (DiffIR), which consists of a compact IR prior extraction network (CPEN), dynamic IR transformer (DIRformer), and denoising network. Specifically, DiffIR has two training stages: pretraining and training DM. In pretraining, we input ground-truth images into CPENS1to capture a compact IR prior representation (IPR) to guide DIRformer. In the second stage, we train the DM to directly estimate the same IRP as pretrained CPENS1only using LQ images. We observe that since the IPR is only a compact vector, DiffIR can use fewer iterations than traditional DM to obtain accurate estimations and generate more stable and realistic results. Since the iterations are few, our DiffIR can adopt a joint optimization of CPENS2, DIRformer, and denoising network, which can further reduce the estimation error influence. We conduct extensive experiments on several IR tasks and achieve SOTA performance while consuming less computational costs. Code is available at https://github.com/Zj-BinXia/DiffIR. Bin Xia 0014, Yulun Zhang 0001, Shiyin Wang, Yapeng Tian, Wenming Yang, Luc Van Gool |
ICCV | 6 |
| 2023 | Basic Binary Convolution Unit for Binarized Image Restoration Network
Bin Xia 0014, Yulun Zhang 0001, Yapeng Tian, Wenming Yang, Radu Timofte, Luc Van Gool |
ICLR | 4 |
| 2023 | Knowledge Distillation based Degradation Estimation for Blind Super-Resolution
Bin Xia 0014, Yulun Zhang 0001, Yapeng Tian, Wenming Yang, Radu Timofte, Luc Van Gool |
ICLR | 4 |
| 2023 | Dual Arbitrary Scale Super-Resolution for Multi-contrast MRI
Jiamiao Zhang, Yichen Chi, Wenming Yang, Yapeng Tian |
MICCAI (10) | 5 |
| 2023 | AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene SynthesisabstractCan machines recording an audio-visual scene produce realistic, matching audio-visual experiences at novel positions and novel view directions? We answer it by studying a new task---real-world audio-visual scene synthesis---and a first-of-its-kind NeRF-based approach for multimodal learning. Concretely, given a video recording of an audio-visual scene, the task is to synthesize new videos with spatial audios along arbitrary novel camera trajectories in that scene. We propose an acoustic-aware audio generation module that integrates prior knowledge of audio propagation into NeRF, in which we implicitly associate audio generation with the 3D geometry and material properties of a visual environment. Furthermore, we present a coordinate transformation module that expresses a view direction relative to the sound source, enabling the model to learn sound source-centric acoustic fields. To facilitate the study of this new task, we collect a high-quality Real-World Audio-Visual Scene (RWAVS) dataset. We demonstrate the advantages of our method on this real-world dataset and the simulation-based SoundSpaces dataset. Notably, we refer readers to view our demo videos for convincing comparisons. Susan Liang, Chao Huang 0033, Yapeng Tian, Chenliang Xu |
NeurIPS | 3 |
| 2023 | Disentangled Counterfactual Learning for Physical Audiovisual Commonsense ReasoningabstractIn this paper, we propose a Disentangled Counterfactual Learning (DCL) approach for physical audiovisual commonsense reasoning. The task aims to infer objects’ physics commonsense based on both video and audio input, with the main challenge is how to imitate the reasoning ability of humans. Most of the current methods fail to take full advantage of different characteristics in multi-modal data, and lacking causal reasoning ability in models impedes the progress of implicit physical knowledge inferring. To address these issues, our proposed DCL method decouples videos into static (time-invariant) and dynamic (time-varying) factors in the latent space by the disentangled sequential encoder, which adopts a variational autoencoder (VAE) to maximize the mutual information with a contrastive loss function. Furthermore, we introduce a counterfactual learning module to augment the model’s reasoning ability by modeling physical knowledge relationships among different objects under counterfactual intervention. Our proposed method is a plug-and-play module that can be incorporated into any baseline. In experiments, we show that our proposed method improves baseline methods and achieves state-of-the-art performance. Our source code is available at https://github.com/Andy20178/DCL. Changsheng Lv, Yapeng Tian, Mengshi Qi, Huadong Ma |
NeurIPS | 3 |
| 2023 | PEANUT: A Human-AI Collaborative Tool for Annotating Audio-Visual DataabstractAudio-visual learning seeks to enhance the computer’s multi-modal perception leveraging the correlation between the auditory and visual modalities. Despite their many useful downstream tasks, such as video retrieval, AR/VR, and accessibility, the performance and adoption of existing audio-visual models have been impeded by the availability of high-quality datasets. Annotating audio-visual datasets is laborious, expensive, and time-consuming. To address this challenge, we designed and developed an efficient audio-visual annotation tool called Peanut. Peanut’s human-AI collaborative pipeline separates the multi-modal task into two single-modal tasks, and utilizes state-of-the-art object detection and sound-tagging models to reduce the annotators’ effort to process each frame and the number of manually-annotated frames needed. A within-subject user study with 20 participants found that Peanut can significantly accelerate the audio-visual data annotation process while maintaining high annotation accuracy. Zheng Zhang 0043, Zheng Ning, Chenliang Xu, Yapeng Tian, Toby Jia-Jun Li |
UIST | 4 |
| 2023 | GDSSR: Toward Real-World Ultra-High-Resolution Image Super-ResolutionabstractAlthough single image super-resolution (SR) has achieved great success, super-resolving the real-world Ultra-High-Resolution (UHR) image remains a challenging issue. Confronted with UHR images, most existing SR methods resort to patch-splitting so that the interconnections among the cropped patches are not attracted reasonable attention during the training and inference procedure. Rather than considering global image degradation levels and types, previous methods only focus on local degradation and unavoidably lead to inter-patch inconsistency, like blocking artifacts in the UHR image. To address this issue, we propose a real-world super-resolution framework to integrate the restoration of different patches through a Global Degradation Supervision Super-Resolution (GDSSR) method. Specifically, a lightweight Global Degradation Extractor is used for extracting global degradation features, which can facilitate restoring better local patches independently and enforce inter-patch consistency. Additionally, a joint training method of local and global patches is proposed to exercise global supervision during the training process, which enhances the degradation estimation and restores more natural results. Experiments show that our GDSSR method achieves superior restoration performance on real-world and UHR image SR datasets. Yichen Chi, Wenming Yang, Yapeng Tian |
IEEE Signal Process. Lett. | 3 |
| 2023 | Meta-Learning-Based Degradation Representation for Blind Super-ResolutionabstractBlind image super-resolution (blind SR) aims to generate high-resolution (HR) images from low-resolution (LR) input images with unknown degradations. To enhance the performance of SR, the majority of blind SR methods introduce an explicit degradation estimator, which helps the SR model adjust to unknown degradation scenarios. Unfortunately, it is impractical to provide concrete labels for the multiple combinations of degradations (e.g., blurring, noise, or JPEG compression) to guide the training of the degradation estimator. Moreover, the special designs for certain degradations hinder the models from being generalized for dealing with other degradations. Thus, it is imperative to devise an implicit degradation estimator that can extract discriminative degradation representations for all types of degradations without requiring the supervision of degradation ground-truth. To this end, we propose a Meta-Learning based Region Degradation Aware SR Network (MRDA), including Meta-Learning Network (MLN), Degradation Extraction Network (DEN), and Region Degradation Aware SR Network (RDAN). To handle the lack of ground-truth degradation, we use the MLN to rapidly adapt to the specific complex degradation after several iterations and extract implicit degradation information. Subsequently, a teacher network MRDAT is designed to further utilize the degradation information extracted by MLN for SR. However, MLN requires iterating on paired LR and HR images, which is unavailable in the inference phase. Therefore, we adopt knowledge distillation (KD) to make the student network learn to directly extract the same implicit degradation representation (IDR) as the teacher from LR images. Furthermore, we introduce an RDAN module that is capable of discerning regional degradations, allowing IDR to adaptively influence various texture patterns. Extensive experiments under classic and real-world degradation settings show that MRDA achieves SOTA performance and can generalize to various degradation processes. Bin Xia 0014, Yapeng Tian, Yulun Zhang 0001, Yucheng Hang, Wenming Yang, Qingmin Liao |
IEEE Trans. Image Process. | 2 |
| 2022 | Efficient Non-local Contrastive Attention for Image Super-resolutionabstractNon-Local Attention (NLA) brings significant improvement for Single Image Super-Resolution (SISR) by leveraging intrinsic feature correlation in natural images. However, NLA gives noisy information large weights and consumes quadratic computation resources with respect to the input size, limiting its performance and application. In this paper, we propose a novel Efficient Non-Local Contrastive Attention (ENLCA) to perform long-range visual modeling and leverage more relevant non-local features. Specifically, ENLCA consists of two parts, Efficient Non-Local Attention (ENLA) and Sparse Aggregation. ENLA adopts the kernel method to approximate exponential function and obtains linear computation complexity. For Sparse Aggregation, we multiply inputs by an amplification factor to focus on informative features, yet the variance of approximation increases exponentially. Therefore, contrastive learning is applied to further separate relevant and irrelevant features. To demonstrate the effectiveness of ENLCA, we build an architecture called Efficient Non-Local Contrastive Network (ENLCN) by adding a few of our modules in a simple backbone. Extensive experimental results show that ENLCN reaches superior performance over state-of-the-art approaches on both quantitative and qualitative evaluations. Bin Xia 0014, Yucheng Hang, Yapeng Tian, Wenming Yang, Qingmin Liao, Jie Zhou 0001 |
AAAI | 3 |
| 2022 | Coarse-to-Fine Embedded PatchMatch and Multi-Scale Dynamic Aggregation for Reference-Based Super-resolutionabstractReference-based super-resolution (RefSR) has made significant progress in producing realistic textures using an external reference (Ref) image. However, existing RefSR methods obtain high-quality correspondence matchings consuming quadratic computation resources with respect to the input size, limiting its application. Moreover, these approaches usually suffer from scale misalignments between the low-resolution (LR) image and Ref image. In this paper, we propose an Accelerated Multi-Scale Aggregation network (AMSA) for Reference-based Super-Resolution, including Coarse-to-Fine Embedded PatchMatch (CFE-PatchMatch) and Multi-Scale Dynamic Aggregation (MSDA) module. To improve matching efficiency, we design a novel Embedded PatchMacth scheme with random samples propagation, which involves end-to-end training with asymptotic linear computational cost to the input size. To further reduce computational cost and speed up convergence, we apply the coarse-to-fine strategy on Embedded PatchMacth constituting CFE-PatchMatch. To fully leverage reference information across multiple scales and enhance robustness to scale misalignment, we develop the MSDA module consisting of Dynamic Aggregation and Multi-Scale Aggregation. The Dynamic Aggregation corrects minor scale misalignment by dynamically aggregating features, and the Multi-Scale Aggregation brings robustness to large scale misalignment by fusing multi-scale information. Experimental results show that the proposed AMSA achieves superior performance over state-of-the-art approaches on both quantitative and qualitative evaluations. Bin Xia 0014, Yapeng Tian, Yucheng Hang, Wenming Yang, Qingmin Liao, Jie Zhou 0001 |
AAAI | 2 |
| 2022 | Transformer-empowered Multi-scale Contextual Matching and Aggregation for Multi-contrast MRI Super-resolutionabstractMagnetic resonance imaging (MRI) can present multicontrast images of the same anatomical structures, enabling multi-contrast super-resolution (SR) techniques. Compared with SR reconstruction using a single-contrast, multicontrast SR reconstruction is promising to yield SR images with higher quality by leveraging diverse yet complementary information embedded in different imaging modalities. However, existing methods still have two shortcomings: (1) they neglect that the multi-contrast features at different scales contain different anatomical details and hence lack effective mechanisms to match and fuse these features for better reconstruction; and (2) they are still deficient in capturing long-range dependencies, which are essential for the regions with complicated anatomical structures. We propose a novel network to comprehensively address these problems by developing a set of innovative Transformer-empowered multi-scale contextual matching and aggregation techniques; we call it McMRSR. Firstly, we tame transformers to model long-range dependencies in both reference and target images. Then, a new multi-scale contextual matching method is proposed to capture corresponding contexts from reference features at different scales. Furthermore, we introduce a multi-scale aggregation mechanism to gradually and interactively aggregate multi-scale matched features for reconstructing the target SR MR image. Extensive experiments demonstrate that our network outperforms state-of-the-art approaches and has great potential to be applied in clinical practice. Codes are available at https://github.com/XAIMI-Lab/McMRSR. Yapeng Tian, Qi Dou 0001, Chengyan Wang, Chenliang Xu, Harry Qin |
CVPR | 3 |
| 2022 | Learning to Answer Questions in Dynamic Audio-Visual ScenariosabstractIn this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive multimodal understanding and spatio-temporal reasoning over audio-visual scenes. To benchmark this task and facilitate our study, we introduce a large-scale MUSIC-AVQA dataset, which contains more than 45K question-answer pairs covering 33 different question templates spanning over different modalities and question types. We develop several baselines and introduce a spatio-temporal grounded audio-visual network for the AVQA problem. Our results demonstrate that AVQA benefits from multisensory perception and our model outperforms recent A-, V-, and AVQA approaches. We believe that our built dataset has the potential to serve as testbed for evaluating and promoting progress in audio-visual scene understanding and spatio-temporal reasoning. Code and dataset: http://gewu-lab.github.io/MUSIC-AVQA/ Guangyao Li 0001, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, Di Hu 0001 |
CVPR | 3 |
| 2022 | Learning Spatio-Temporal Downsampling for Effective Video Upscaling
Xiaoyu Xiang, Yapeng Tian, Vijay Rengarajan, Lucas D. Young, Bo Zhu 0011 |
ECCV (18) | 2 |
| 2022 | DuDoCAF: Dual-Domain Cross-Attention Fusion with Recurrent Transformer for Fast Multi-contrast MR Imaging
Bin Sui, Chengyan Wang, Yapeng Tian, Qi Dou 0001, Harry Qin |
MICCAI (6) | 4 |
| 2022 | Multi-modal Grouping Network for Weakly-Supervised Audio-Visual Video ParsingabstractThe audio-visual video parsing task aims to parse a video into modality- and category-aware temporal segments. Previous work mainly focuses on weakly-supervised approaches, which learn from video-level event labels. During training, they do not know which modality perceives and meanwhile which temporal segment contains the video event. Since there is no explicit grouping in the existing frameworks, the modality and temporal uncertainties make these methods suffer from false predictions. For instance, segments in the same category could be predicted in different event classes. Learning compact and discriminative multi-modal subspaces is essential for mitigating the issue. To this end, in this paper, we propose a novel Multi-modal Grouping Network, namely MGN, for explicitly semantic-aware grouping. Specifically, MGN aggregates event-aware unimodal features through unimodal grouping in terms of learnable categorical embedding tokens. Furthermore, it leverages the cross-modal grouping for modality-aware prediction to match the video-level target. Our simple framework achieves improving results against previous baselines on weakly-supervised audio-visual video parsing. In addition, our MGN is much more lightweight, using only 47.2% of the parameters of baselines (17 MB vs. 36 MB). Code is available at https://github.com/stoneMo/MGN. Shentong Mo, Yapeng Tian |
NeurIPS | 2 |
| 2021 | Space-Time Memory Network for Sounding Object Localization in Videos
Yapeng Tian, Chenliang Xu |
BMVC | 2 |
| 2021 | Cyclic Co-Learning of Sounding Object Visual Grounding and Sound SeparationabstractThere are rich synchronized audio and visual events in our daily life. Inside the events, audio scenes are associated with the corresponding visual objects; meanwhile, sounding objects can indicate and help to separate their individual sounds in the audio track. Based on this observation, in this paper, we propose a cyclic co-learning (CCoL) paradigm that can jointly learn sounding object visual grounding and audio-visual sound separation in a unified framework. Concretely, we can leverage grounded object-sound relations to improve the results of sound separation. Meanwhile, benefiting from discriminative information from separated sounds, we improve training example sampling for sounding object grounding, which builds a co-learning cycle for the two tasks and makes them mutually beneficial. Extensive experiments show that the proposed framework outperforms the compared recent approaches on both tasks, and they can benefit from each other with our cyclic co-learning. The source code and pre-trained models are released in https://github.com/YapengTian/CCOL-CVPR21. Yapeng Tian, Di Hu 0001, Chenliang Xu |
CVPR | 1 |
| 2021 | Can Audio-Visual Integration Strengthen Robustness Under Multimodal Attacks?abstractIn this paper, we propose to make a systematic study on machines’ multisensory perception under attacks. We use the audio-visual event recognition task against multi-modal adversarial attacks as a proxy to investigate the robustness of audio-visual learning. We attack audio, visual, and both modalities to explore whether audio-visual integration still strengthens perception and how different fusion mechanisms affect the robustness of audio-visual models. For interpreting the multimodal interactions under attacks, we learn a weakly-supervised sound source visual localization model to localize sounding regions in videos. To mitigate multimodal attacks, we propose an audio-visual defense approach based on an audio-visual dissimilarity constraint and external feature memory banks. Extensive experiments demonstrate that audio-visual models are susceptible to multimodal adversarial attacks; audio-visual integration could decrease the model robustness rather than strengthen under multimodal attacks; even a weakly-supervised sound source visual localization model can be successfully fooled; our defense method can improve the invulnerability of audio-visual networks without significantly sacrificing clean model performance. The source code and pre-trained models are released in https://github.com/YapengTian/AV-Robustness-CVPR21. Yapeng Tian, Chenliang Xu |
CVPR | 1 |
| 2021 | Video Matting via Consistency-Regularized Graph Neural NetworksabstractLearning temporally consistent foreground opacity from videos, i.e., video matting, has drawn great attention due to the blossoming of video conferencing. Previous approaches are built on top of image matting models, which fail in maintaining the temporal coherence when being adapted to videos. They either utilize the optical flow to smooth frame-wise prediction, where the performance is dependent on the selected optical flow model; or naively combine feature maps from multiple frames, which does not model well the correspondence of pixels in adjacent frames. In this paper, we propose to enhance the temporal coherence by Consistency-Regularized Graph Neural Networks (CRGNN) with the aid of a synthesized video matting dataset. CRGNN utilizes Graph Neural Networks (GNN) to relate adjacent frames such that pixels or regions that are incorrectly predicted in one frame can be corrected by leveraging information from its neighboring frames. To generalize our model from synthesized videos to real-world videos, we propose a consistency regularization technique to enforce the consistency on the alpha and foreground when blending them with different backgrounds. To evaluate the efficacy of CRGNN, we further collect a real-world dataset with annotated alpha mattes. Compared with state-of-the-art methods that require hand-crafted trimaps or backgrounds for modeling training, CRGNN generates favorably results with the help of unlabeled real training dataset. The source code and datasets are available at https://github.com/TiantianWang/VideoMattingCRGNN.git. Sifei Liu, Yapeng Tian, Ming-Hsuan Yang 0001 |
ICCV | 3 |
| 2021 | Residual Dense Network for Image RestorationabstractRecently, deep convolutional neural network (CNN) has achieved great success for image restoration (IR) and provided hierarchical features at the same time. However, most deep CNN based IR models do not make full use of the hierarchical features from the original low-quality images; thereby, resulting in relatively-low performance. In this work, we propose a novel and efficient residual dense network (RDN) to address this problem in IR, by making a better tradeoff between efficiency and effectiveness in exploiting the hierarchical features from all the convolutional layers. Specifically, we propose residual dense block (RDB) to extract abundant local features via densely connected convolutional layers. RDB further allows direct connections from the state of preceding RDB to all the layers of current RDB, leading to a contiguous memory mechanism. To adaptively learn more effective features from preceding and current local features and stabilize the training of wider network, we proposed local feature fusion in RDB. After fully obtaining dense local features, we use global feature fusion to jointly and adaptively learn global hierarchical features in a holistic way. We demonstrate the effectiveness of RDN with several representative IR applications, single image super-resolution, Gaussian image denoising, image compression artifact reduction, and image deblurring. Experiments on benchmark and real-world datasets show that our RDN achieves favorable performance against state-of-the-art methods for each IR task quantitatively and visually. Yulun Zhang 0001, Yapeng Tian, Yu Kong 0001, Bineng Zhong 0001, Yun Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | TDAN: Temporally-Deformable Alignment Network for Video Super-ResolutionabstractVideo super-resolution (VSR) aims to restore a photo-realistic high-resolution (HR) video frame from both its corresponding low-resolution (LR) frame (reference frame) and multiple neighboring frames (supporting frames). Due to varying motion of cameras or objects, the reference frame and each support frame are not aligned. Therefore, temporal alignment is a challenging yet important problem for VSR. Previous VSR methods usually utilize optical flow between the reference frame and each supporting frame to warp the supporting frame for temporal alignment. However, both inaccurate flow and the image-level warping strategy will lead to artifacts in the warped supporting frames. To overcome the limitation, we propose a temporally-deformable alignment network (TDAN) to adaptively align the reference frame and each supporting frame at the feature level without computing optical flow. The TDAN uses features from both the reference frame and each supporting frame to dynamically predict offsets of sampling convolution kernels. By using the corresponding kernels, TDAN transforms supporting frames to align with the reference frame. To predict the HR video frame, a reconstruction network taking aligned frames and the reference frame is utilized. Experimental results demonstrate that the TDAN is capable of alleviating occlusions and artifacts for temporal alignment and the TDAN-based VSR model outperforms several recent state-of-the-art VSR networks with a comparable or even much smaller model size. The source code and pre-trained models are released in https://github.com/YapengTian/TDAN-VSR. Yapeng Tian, Yulun Zhang 0001, Yun Fu 0001, Chenliang Xu |
CVPR | 1 |
| 2020 | Zooming Slow-Mo: Fast and Accurate One-Stage Space-Time Video Super-ResolutionabstractIn this paper, we explore the space-time video super-resolution task, which aims to generate a high-resolution (HR) slow-motion video from a low frame rate (LFR), low-resolution (LR) video. A simple solution is to split it into two sub-tasks: video frame interpolation (VFI) and video super-resolution (VSR). However, temporal interpolation and spatial super-resolution are intra-related in this task. Two-stage methods cannot fully take advantage of the natural property. In addition, state-of-the-art VFI or VSR networks require a large frame-synthesis or reconstruction module for predicting high-quality video frames, which makes the two-stage methods have large model sizes and thus be time-consuming. To overcome the problems, we propose a one-stage space-time video super-resolution framework, which directly synthesizes an HR slow-motion video from an LFR, LR video. Rather than synthesizing missing LR video frames as VFI networks do, we firstly temporally interpolate LR frame features in missing LR video frames capturing local temporal contexts by the proposed feature temporal interpolation network. Then, we propose a deformable ConvLSTM to align and aggregate temporal information simultaneously for better leveraging global temporal contexts. Finally, a deep reconstruction network is adopted to predict HR slow-motion video frames. Extensive experiments on benchmark datasets demonstrate that the proposed method not only achieves better quantitative and qualitative performance but also is more than three times faster than recent two-stage state-of-the-art methods, e.g., DAIN+EDVR and DAIN+RBPN. Xiaoyu Xiang, Yapeng Tian, Yulun Zhang 0001, Yun Fu 0001, Jan P. Allebach, Chenliang Xu |
CVPR | 2 |
| 2020 | Unified Multisensory Perception: Weakly-Supervised Audio-Visual Video Parsing
Yapeng Tian, Dingzeyu Li, Chenliang Xu |
ECCV (3) | 1 |
| 2020 | LCSCNet: Linear Compressing-Based Skip-Connecting Network for Image Super-ResolutionabstractIn this paper, we develop a concise but efficient network architecture called linear compressing based skipconnecting network (LCSCNet) for image super-resolution. Compared with two representative network architectures with skip connections, ResNet and DenseNet, a linear compressing layer is designed in LCSCNet for skip connection, which connects former feature maps and distinguishes them from newly-explored feature maps. In this way, the proposed LCSCNet enjoys the merits of the distinguish feature treatment of DenseNet and the parametereconomic form of ResNet. Moreover, to better exploit hierarchical information from both low and high levels of various receptive fields in deep models, inspired by gate units in LSTM, we also propose an adaptive element-wise fusion strategy with multisupervised training. Experimental results in comparison with state-of-the-art algorithms validate the effectiveness of LCSCNet. Wenming Yang, Xuechen Zhang 0003, Yapeng Tian, Wei Wang 0194, Jing-Hao Xue, Qingmin Liao |
IEEE Trans. Image Process. | 3 |
| 2019 | CFSNet: Toward a Controllable Feature Space for Image RestorationabstractDeep learning methods have witnessed the great progress in image restoration with specific metrics (e.g., PSNR, SSIM). However, the perceptual quality of the restored image is relatively subjective, and it is necessary for users to control the reconstruction result according to personal preferences or image characteristics, which cannot be done using existing deterministic networks. This motivates us to exquisitely design a unified interactive framework for general image restoration tasks. Under this framework, users can control continuous transition of different objectives, e.g., the perception-distortion trade-off of image super-resolution, the trade-off between noise reduction and detail preservation. We achieve this goal by controlling the latent features of the designed network. To be specific, our proposed framework, named Controllable Feature Space Network (CFSNet), is entangled by two branches based on different objectives. Our framework can adaptively learn the coupling coefficients of different layers and channels, which provides finer control of the restored image quality. Experiments on several typical image restoration tasks fully validate the effective benefits of the proposed method. Code is available at https://github.com/qibao77/CFSNet. Wei Wang 0194, Ruiming Guo, Yapeng Tian, Wenming Yang |
ICCV | 3 |
| 2019 | Deep Learning for Single Image Super-Resolution: A Brief ReviewabstractSingle image super-resolution (SISR) is a notoriously challenging ill-posed problem that aims to obtain a high-resolution output from one of its low-resolution versions. Recently, powerful deep learning algorithms have been applied to SISR and have achieved state-of-the-art performance. In this survey, we review representative deep learning-based SISR methods and group them into two categories according to their contributions to two essential aspects of SISR: The exploration of efficient neural network architectures for SISR and the development of effective optimization objectives for deep SISR learning. For each category, a baseline is first established, and several critical limitations of the baseline are summarized. Then, representative works on overcoming these limitations are presented based on their original content, as well as our critical exposition and analyses, and relevant comparisons are conducted from a variety of perspectives. Finally, we conclude this review with some current challenges and future trends in SISR that leverage deep learning algorithms. Wenming Yang, Xuechen Zhang 0003, Yapeng Tian, Wei Wang 0194, Jing-Hao Xue, Qingmin Liao |
IEEE Trans. Multim. | 3 |
| 2018 | Residual Dense Network for Image Super-ResolutionabstractA very deep convolutional neural network (CNN) has recently achieved great success for image super-resolution (SR) and offered hierarchical features as well. However, most deep CNN based SR models do not make full use of the hierarchical features from the original low-resolution (LR) images, thereby achieving relatively-low performance. In this paper, we propose a novel residual dense network (RDN) to address this problem in image SR. We fully exploit the hierarchical features from all the convolutional layers. Specifically, we propose residual dense block (RDB) to extract abundant local features via dense connected convolutional layers. RDB further allows direct connections from the state of preceding RDB to all the layers of current RDB, leading to a contiguous memory (CM) mechanism. Local feature fusion in RDB is then used to adaptively learn more effective features from preceding and current local features and stabilizes the training of wider network. After fully obtaining dense local features, we use global feature fusion to jointly and adaptively learn global hierarchical features in a holistic way. Experiments on benchmark datasets with different degradation models show that our RDN achieves favorable performance against state-of-the-art methods. Yulun Zhang 0001, Yapeng Tian, Yu Kong 0001, Bineng Zhong 0001, Yun Fu 0001 |
CVPR | 2 |
| 2018 | Audio-Visual Event Localization in Unconstrained Videos
Yapeng Tian, Jing Shi 0005, Bochen Li, Zhiyao Duan, Chenliang Xu |
ECCV (2) | 1 |
| 2017 | Adaptive anchor-point selection for single image super-resolutionabstractThis paper presents an adaptive anchor-point selection method for single image super-resolution (SR), which is based upon internal example-based SR model via locality constrained anchored neighborhood regression. The anchor points are fixed in anchored SR methods and are not flexible and customized for different input low-resolution (LR) images. To overcome this defect, we adaptively select anchor points via constructing customized training set for different input LR images, which can be realized by an internal example-based SR method. We introduce a locality-constrained anchored neighborhood regression to learn the relationship between LR space and high-resolution (HR) space. Extensive experimental results demonstrate that the performance of proposed method is competitive with several state-of-the-art SR methods. Xuesen Shang, Wenming Yang, Shuifa Sun, Yapeng Tian, Hai Chen, Kaiquan Chen |
VCIP | 4 |
| 2016 | Anchored neighborhood regression based single image super-resolution from self-examplesabstractIn this paper, we present a novel self-learning single image super-resolution (SR) method, which restores a high-resolution (HR) image from self-examples extracted from the low-resolution (LR) input image itself without relying on extra external training images. In the proposed method, we directly use sampled image patches as the anchor points, and then learn multiple linear mapping functions based on anchored neighborhood regression to transform LR space into HR space. Moreover, we utilize the flipped and rotated versions of the self-examples to expand the internal patch space. Experimental comparison on standard benchmarks with state-of-the-art methods validates the effectiveness of the proposed approach. Yapeng Tian, Fei Zhou 0001, Wenming Yang, Xuesen Shang, Qingmin Liao |
ICIP | 1 |
| 2016 | Consistent Coding Scheme for Single-Image Super-Resolution Via Independent DictionariesabstractIn this paper, we present a unified frame based on collaborative representation (CR) for single-image super-resolution (SR), which learns low-resolution (LR) and high-resolution (HR) dictionaries independently in the training stage and adopts a consistent coding scheme (CCS) to guarantee the prediction accuracy of HR coding coefficients during SR reconstruction. The independent LR and HR dictionaries are learned based on CR with l2-norm regularization, which can well describe the corresponding LR and HR patch space, respectively. Furthermore, a mapping function is learned to map LR coding coefficients onto the corresponding HR coding coefficients. Propagation filtering can achieve smoothing over an image while preserving image context like edges or textural regions. Moreover, to preserve the edge structures of a super-resolved image and suppress artifacts, a propagation filtering-based constraint and image nonlocal self-similarity regularization are introduced into the SR reconstruction framework. Experimental comparison with state-of-the-art single image SR algorithms validates the effectiveness of proposed approach. Wenming Yang, Yapeng Tian, Fei Zhou 0001, Qingmin Liao, Hai Chen, Chenglin Zheng |
IEEE Trans. Multim. | 2 |