VLDB 2026 Research / reviewers in the wild / expert
Taiga Yamane
dblp:348/8898
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
0009-0004-6254-8810ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Difference Vector Equalization for Robust Fine-tuning of Vision-Language ModelsabstractContrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image and text encoders. This paper aims to robustly fine-tune these vision-language models on in-distribution (ID) data without compromising their generalization abilities in out-of-distribution (OOD) and zero-shot settings. Current robust fine-tuning methods tackle this challenge by reusing contrastive learning, which was used in pre-training, for fine-tuning. However, we found that these methods distort the geometric structure of the embeddings, which plays a crucial role in the generalization of vision-language models, resulting in limited OOD and zero-shot performance. To address this, we propose Difference Vector Equalization (DiVE), which preserves the geometric structure during fine-tuning. The idea behind DiVE is to constrain difference vectors, each of which is obtained by subtracting the embeddings extracted from the pre-trained and fine-tuning models for the same data sample. By constraining the difference vectors to be equal across various data samples, we effectively preserve the geometric structure. Therefore, we introduce two losses: average vector loss (AVL) and pairwise vector loss (PVL). AVL preserves the geometric structure globally by constraining difference vectors to be equal to their weighted average. PVL preserves the geometric structure locally by ensuring a consistent multimodal alignment. Our experiments demonstrate that DiVE effectively preserves the geometric structure, achieving strong results across ID, OOD, and zero-shot metrics. Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane, Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
AAAI | 4 |
| 2025 | Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language ModelabstractThis paper proposes a personalization method for speech emotion recognition (SER) through in-context learning (ICL). Since the expression of emotions varies from person to person, speaker-specific adaptation is crucial for improving the SER performance. Conventional SER methods have been personalized using emotional utterances of a target speaker, but it is often difficult to prepare utterances corresponding to all emotion labels in advance. Our idea to overcome this difficulty is to obtain speaker characteristics by conditioning a few emotional utterances of the target speaker in ICL-based inference. ICL is a method to perform unseen tasks by conditioning a few inputoutput examples through inference in large language models (LLMs). We meta-train a speech-language model extended from the LLM to learn how to perform personalized SER via ICL. Experimental results using our newly collected SER dataset demonstrate that the proposed method outperforms conventional methods. Mana Ihori, Taiga Yamane, Naotaka Kawata, Naoki Makishima, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
ASRU | 2 |
| 2025 | Phoneme Overlapping-Aware Pre-Training with External Text Resources for Multi-Talker ASRabstractThis paper proposes a new pre-training method utilizing external text resources to improve the robustness of single-channel multi-talker automatic speech recognition (MTASR) across various linguistic domains. In the development of single-talker ASR systems, various pre-training methods have been studied to acquire knowledge about word order and correspondence between phonetic information and text from external text resources. However, methods focusing on improving MT-ASR performance by leveraging external text resources remain underexplored. To bridge this gap, we aim to acquire the ability to discover multiple texts contained within overlapping phonetic information. The key idea of the proposed method is to induce overlapping phenomena in phoneme sequences in order to reproduce a task similar to MT-ASR using external text resources. Our experiments demonstrate that the proposed method significantly improves the MT-ASR performance on both in-domain and out-of-domain linguistic tasks. Ryo Masumura, Tomohiro Tanaka, Naoki Makishima, Mana Ihori, Shota Orihashi, Naotaka Kawata, Taiga Yamane, Takafumi Moriya |
ASRU | 7 |
| 2025 | MVTrajecter: Multi-View Pedestrian Tracking With Trajectory Motion Cost and Trajectory Appearance CostabstractMulti-View Pedestrian Tracking (MVPT) aims to track pedestrians in the form of a bird's eye view occupancy map from multi-view videos. End-to-end methods that detect and associate pedestrians within one model have shown great progress in MVPT. The motion and appearance information of pedestrians is important for the association, but previous end-to-end MVPT methods rely only on the current and its single adjacent past timestamp, discarding the past trajectories before that. This paper proposes a novel end-to-end MVPT method called Multi-View Trajectory Tracker (MVTrajecter) that utilizes information from multiple timestamps in past trajectories for robust association. MVTrajecter introduces trajectory motion cost and trajectory appearance cost to effectively incorporate motion and appearance information, respectively. These costs calculate which pedestrians at the current and each past timestamp are likely identical based on the information between those timestamps. Even if a current pedestrian could be associated with a false pedestrian at some past timestamp, these costs enable the model to associate that current pedestrian with the correct past trajectory based on other past timestamps. In addition, MVTrajecter effectively captures the relationships between multiple timestamps leveraging the attention mechanism. Extensive experiments demonstrate the effectiveness of each component in MVTrajecter and show that it outperforms the previous state-of-the-art methods. Taiga Yamane, Ryo Masumura, Shota Orihashi |
ICCV | 1 |
| 2025 | Unified Audio-Visual Modeling for Recognizing Which Face Spoke When and What in Multi-Talker Overlapped Speech and Video
Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
INTERSPEECH | 3 |
| 2025 | SOMSRED-SVC: Sequential Output Modeling with Speaker Vector Constraints for Joint Multi-Talker Overlapped ASR and Speaker Diarization
Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
INTERSPEECH | 3 |
| 2024 | MVAFormer: RGB-Based Multi-View Spatio-Temporal Action Recognition with TransformerabstractMulti-view action recognition aims to recognize human actions using multiple camera views and deals with occlusion caused by obstacles or crowds. In this task, cooperation among views, which generates a joint representation by combining multiple views, is vital. Previous studies have explored promising cooperation methods for improving performance. However, since their methods focus only on the task setting of recognizing a single action from an entire video, they are not applicable to the recently popular spatio-temporal action recognition (STAR) setting, in which each person’s action is recognized sequentially. To address this problem, this paper proposes a multi-view action recognition method for the STAR setting, called MVAFormer. In MVAFormer, we introduce a novel transformer-based cooperation module among views. In contrast to previous studies, which utilize embedding vectors with lost spatial information, our module utilizes the feature map for effective cooperation in the STAR setting, which preserves the spatial information. Furthermore, in our module, we divide the self-attention for the same and different views to model the relationship between multiple views effectively. The results of experiments using a newly collected dataset demonstrate that MVAFormer outperforms the comparison baselines by approximately 4.4 points on the F-measure. Taiga Yamane, Ryo Masumura, Shotaro Tora |
ICIP | 1 |
| 2024 | Unified Multi-Talker ASR with and without Target-speaker Enrollment
Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Mana Ihori, Naotaka Kawata, Shota Orihashi, Kazutoshi Shinoda, Taiga Yamane, Saki Mizuno, Keita Suzuki, Nobukatsu Hojo, Takafumi Moriya, Atsushi Ando |
INTERSPEECH | 8 |
| 2023 | Leveraging Language Embeddings for Cross-Lingual Self-Supervised Speech Representation LearningabstractIn this paper, we propose novel cross-lingual self-supervised speech representation learning methods that explicitly consider language information. Cross-lingual self-supervised speech representation learning has been studied to make effective use of diverse data in various languages. Previous methods train models from multilingual datasets without taking language into account. However, it is difficult to train speech representations from multilingual datasets in the same space without language specification since there are clear differences in the acoustic context between languages. To solve this problem, we propose leveraging language IDs to build self-supervised speech representation learning models that explicitly consider language information. Our proposed models utilize fixed-dimensional language embeddings converted from language IDs for the model learning the relationship between related speech representations in different languages. We investigate two strategies to introduce language embeddings into the models: adding the embeddings to all of the inputs and concatenating to the inputs of the Transformer. We experimentally investigated how the difference between the two strategies affects the downstream tasks. Experimental results on the English and Japanese datasets show that the proposed methods improve the accuracies of downstream automatic speech recognition tasks. Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Hiroshi Sato 0002, Taiga Yamane, Takanori Ashihara, Kohei Matsuura, Takafumi Moriya |
ICASSP | 5 |
| 2023 | OnDA-DETR: Online Domain Adaptation for Detection Transformers with Self-Training FrameworkabstractThis paper presents a novel method for online domain adaptation (OnDA) for DEtection TRansformer (DETR)-based object detection models called OnDA-DETR. OnDA is a domain adaptation paradigm that adapts a model trained on the source domain data to perform well on the target domain in an online manner during testing, using only the unlabeled test data from the target domain. Due to challenging and realistic problem settings, OnDA has garnered significant attention. However, OnDA methods for DETR-based models, which have demonstrated excellent performance in object detection research fields, had not been developed. OnDA-DETR is the first OnDA method specifically designed for DETR-based models. OnDA-DETR incorporates a self-training framework that generates pseudo-labels for the unlabeled target domain data. To effectively incorporate the self-training framework into DETR-based models, we leverage recall-aware pseudo-labeling and quality-aware training in OnDA-DETR. Experimental results indicate that OnDA-DETR improves the performance of the source-trained model by about 3.0 % points through OnDA. Taiga Yamane, Naoki Makishima, Keita Suzuki, Atsushi Ando, Ryo Masumura |
ICIP | 2 |
| 2023 | End-to-End Joint Target and Non-Target Speakers ASR
Ryo Masumura, Naoki Makishima, Taiga Yamane, Yoshihiko Yamazaki, Saki Mizuno, Mana Ihori, Mihiro Uchida, Keita Suzuki, Hiroshi Sato 0002, Tomohiro Tanaka, Akihiko Takashima, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando |
INTERSPEECH | 3 |