Yan-Bo Lin

dblp:06/11431 · DBLP profile ↗
← Back
13ranked-venue papers
10as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 7 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising
abstract
In this paper, we introduce zero-shot audio-video editing, a novel task that requires transforming original audio-visual content to align with a specified textual prompt without additional model training. To evaluate this task, we curate a benchmark dataset, AVED-Bench, designed explicitly for zero-shot audio-video editing. AVED-Bench includes 110 videos, each with a 10-second duration, spanning 11 categories from VGGSound. It offers diverse prompts and scenarios that require precise alignment between auditory and visual elements, enabling robust evaluation. We identify limitations in existing zero-shot audio and video editing methods, particularly in synchronization and coherence between modalities, which often result in inconsistent outcomes. To address these challenges, we propose AVED, a zero-shot cross-modal delta denoising framework that leverages audio-video interactions to achieve synchronized and coherent edits. AVED demonstrates superior results on both AVED-Bench and the recent OAVE dataset to validate its generalization capabilities. Our codes, data, and results are available at https://genjib.github.io/project_page/AVED/index.html
Yan-Bo Lin, Zhengyuan Yang, Chung-Ching Lin, Gedas Bertasius
WACV1
2025 Dam: Dynamic Adapter Merging for Continual Video QA Learning
abstract
We present a parameter-efficient method for continual video question-answering (VidQA) learning. Our method, named Dam, uses the proposed Dynamic Adapter Merging to (i) mitigate catastrophic forgetting, (ii) enable efficient adaptation to continually arriving datasets, (iii) handle inputs from unknown datasets during inference, and (iv) enable knowledge sharing across similar dataset domains. Given a set of continually streaming VidQA datasets, we sequentially train dataset-specific adapters for each dataset while freezing the parameters of a large pretrained video-language backbone. During inference, given a video-question sample from an unknown domain, our method first uses the proposed non-parametric router function to compute a probability for each adapter, reflecting how relevant that adapter is to the current video-question input instance. Subsequently, the proposed dynamic adapter merging scheme aggregates all the adapter weights into a new adapter instance tailored for that particular test sample to compute the final VidQA prediction, mitigating the impact of inaccurate router predictions and facilitating knowledge sharing across domains. Our Dam model outperforms prior state-of-the-art continual learning approaches by 9.1% while exhibiting 1.9% less forgetting on 6 VidQA datasets spanning various domains. We further extend Dam to continual image classification and image QA and outperform prior methods by a large margin. The code will be publicly released.
Yi-Lin Sung, Yan-Bo Lin, Mohit Bansal, Gedas Bertasius
WACV4
2025 VMAs: Video-to-Music Generation via Semantic Alignment in Web Music Videos
abstract
We present a framework for learning to generate background music from video inputs. Unlike existing works that rely on symbolic musical annotations, which are limited in quantity and diversity, our method leverages large-scale web videos accompanied by background music. This enables our model to learn to generate realistic and diverse music. To accomplish this goal, we develop a generative video-music Transformer with a novel semantic video-music alignment scheme. Our model uses a joint autoregressive and contrastive learning objective, which encourages the generation of music aligned with high-level video content. We also introduce a novel video-beat alignment scheme to match the generated music beats with the low-level motions in the video. Lastly, to capture fine-grained visual cues in a video needed for realistic background music generation, we introduce a new temporal video encoder architecture, allowing us to efficiently process videos consisting of many densely sampled frames. We train our frame-work on our newly curated DISCO-MV dataset, consisting of 2.2M video-music samples, which is orders of magnitude larger than any prior datasets used for video music generation. Our method outperforms existing approaches on the DISCO-MV and MusicCaps datasets according to various music generation evaluation metrics, including human evaluation. Results are available at https://genjib.github.io/project_page/VMAs/index.html
Yan-Bo Lin, Gedas Bertasius
WACV1
2024 Siamese Vision Transformers are Scalable Audio-Visual Learners
Yan-Bo Lin, Gedas Bertasius
ECCV (14)1
2023 Vision Transformers are Parameter-Efficient Audio-Visual Learners
abstract
Vision transformers (ViTs) have achieved impressive results on various computer vision tasks in the last several years. In this work, we study the capability of frozen ViTs, pretrained only on visual data, to generalize to audio-visual data without finetuning any of its original parameters. To do so, we propose a latent audio-visual hybrid (LAVISH) adapter that adapts pretrained ViTs to audio-visual tasks by injecting a small number of trainable parameters into every layer of a frozen ViT. To efficiently fuse visual and audio cues, our LAVISH adapter uses a small set of latent tokens, which form an attention bottleneck, thus, eliminating the quadratic cost of standard cross-attention. Compared to the existing modality-specific audio-visual methods, our approach achieves competitive or even better performance on various audio-visual tasks while using fewer tunable parameters and without relying on costly audio pretraining or external audio encoders. Our code is available at https://genjib.github.io/project_page/LAVISH/
Yan-Bo Lin, Yi-Lin Sung, Jie Lei 0003, Mohit Bansal, Gedas Bertasius
CVPR1
2023 Unsupervised sound localization via iterative contrastive learning
Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee 0001, Yen-Yu Lin, Ming-Hsuan Yang 0001
Comput. Vis. Image Underst.1
2022 EclipSE: Efficient Long-Range Video Retrieval Using Sight and Sound
Yan-Bo Lin, Jie Lei 0003, Mohit Bansal, Gedas Bertasius
ECCV (34)1
2021 Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation
abstract
Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which would degrade the user experience due to the lack of ambient information. To address this issue, we propose an audio spatialization framework to convert a monaural video into a binaural one exploiting the relationship across audio and visual components. By preserving the left-right consistency in both audio and visual modalities, our learning strategy can be viewed as a self-supervised learning technique, and alleviates the dependency on a large amount of video data with ground truth binaural audio data during training. Experiments on benchmark datasets confirm the effectiveness of our proposed framework in both semi-supervised and fully supervised scenarios, with ablation studies and visualization further support the use of our model for audio spatialization.
Yan-Bo Lin, Yu-Chiang Frank Wang
AAAI1
2021 Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video Parsing
abstract
The audio-visual video parsing task aims to temporally parse a video into audio or visual event categories. However, it is labor intensive to temporally annotate audio and visual events and thus hampers the learning of a parsing model. To this end, we propose to explore additional cross-video and cross-modality supervisory signals to facilitate weakly-supervised audio-visual video parsing. The proposed method exploits both the common and diverse event semantics across videos to identify audio or visual events. In addition, our method explores event co-occurrence across audio, visual, and audio-visual streams. We leverage the explored cross-modality co-occurrence to localize segments of target events while excluding irrelevant ones. The discovered supervisory signals across different videos and modalities can greatly facilitate the training with only video-level annotations. Quantitative and qualitative results demonstrate that the proposed method performs favorably against existing methods on weakly-supervised audio-visual video parsing.
Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee 0001, Yen-Yu Lin, Ming-Hsuan Yang 0001
NeurIPS1
2020 Audiovisual Transformer with Instance Attention for Audio-Visual Event Localization
Yan-Bo Lin, Yu-Chiang Frank Wang
ACCV (6)1
2019 Dual-modality Seq2Seq Network for Audio-visual Event Localization
abstract
Audio-visual event localization requires one to identify the event which is both visible and audible in a video (either at a frame or video level). To address this task, we propose a deep neural network named Audio-Visual sequence-to-sequence dual network (AVSDN). By jointly taking both audio and visual features at each time segment as inputs, our proposed model learns global and local event information in a sequence to sequence manner, which can be realized in either fully supervised or weakly supervised settings. Empirical results confirm that our proposed method performs favorably against recent deep learning approaches in both settings.
Yan-Bo Lin, Yu-Jhe Li, Yu-Chiang Frank Wang
ICASSP1
2019 Cross-Dataset Person Re-Identification via Unsupervised Pose Disentanglement and Adaptation
abstract
Person re-identification (re-ID) aims at recognizing the same person from images taken across different cameras. To address this challenging task, existing re-ID models typically rely on a large amount of labeled training data, which is not practical for real-world applications. To alleviate this limitation, researchers now targets at cross-dataset re-ID which focuses on generalizing the discriminative ability to the unlabeled target domain when given a labeled source domain dataset. To achieve this goal, our proposed Pose Disentanglement and Adaptation Network (PDA-Net) aims at learning deep image representation with pose and domain information properly disentangled. With the learned cross-domain pose invariant feature space, our proposed PDA-Net is able to perform pose disentanglement across domains without supervision in identities, and the resulting features can be applied to cross-dataset re-ID. Both of our qualitative and quantitative results on two benchmark datasets confirm the effectiveness of our approach and its superiority over the state-of-the-art cross-dataset Re-ID approaches.
Yu-Jhe Li, Ci-Siang Lin, Yan-Bo Lin, Yu-Chiang Frank Wang
ICCV3
2017 Single channel source separation using graph sparse NMF and adaptive dictionary learning
abstract
The aim of single channel source separation is to accurately recover signals from mixtures. Non-negative matrix factorization (NMF) is a popular method to separate mixed signals using learned dictionaries. These dictionaries can be produced efficiently by sparse NMF to approximate the input signal as closely as possible. However, the literature does not consider the structure of the data in terms of the similarity among vertices of the input signal. Furthermore, state-of-art variants of NMF that are more efficient than conventional ones have not been utilized, and the learned dictionary is typically fixed in the separating phase. This strategy is not favorable because the training data and the testing data totally differ. To deal with these issues, our work proposes a method that incorporates the graph regularization into group sparsity β-NMF to improve the performance of source separation. The proposed algorithms differ from those in the literature by using an adaptive dictionary in which particular characteristics of the testing data are updated to produce newer dictionaries. Experimental results demonstrate that our proposed method is outstandingly effective in speech separation in various scenarios, relative to the baseline.
Yuan-Shan Lee, Yan-Bo Lin, Yung-Hui Li, Tzu-Chiang Tai, Jia-Ching Wang
Intell. Data Anal.3