VLDB 2026 Research / reviewers in the wild / expert
Naoki Makishima
dblp:246/8004
· DBLP profile ↗
33ranked-venue papers
9as first author
30since 2021 · last 2026
0000-0002-7065-315XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 8 first-author · 30 since 2021Artificial intelligence and machine learning · 23 · 7 first-author · 20 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Difference Vector Equalization for Robust Fine-tuning of Vision-Language ModelsabstractContrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image and text encoders. This paper aims to robustly fine-tune these vision-language models on in-distribution (ID) data without compromising their generalization abilities in out-of-distribution (OOD) and zero-shot settings. Current robust fine-tuning methods tackle this challenge by reusing contrastive learning, which was used in pre-training, for fine-tuning. However, we found that these methods distort the geometric structure of the embeddings, which plays a crucial role in the generalization of vision-language models, resulting in limited OOD and zero-shot performance. To address this, we propose Difference Vector Equalization (DiVE), which preserves the geometric structure during fine-tuning. The idea behind DiVE is to constrain difference vectors, each of which is obtained by subtracting the embeddings extracted from the pre-trained and fine-tuning models for the same data sample. By constraining the difference vectors to be equal across various data samples, we effectively preserve the geometric structure. Therefore, we introduce two losses: average vector loss (AVL) and pairwise vector loss (PVL). AVL preserves the geometric structure globally by constraining difference vectors to be equal to their weighted average. PVL preserves the geometric structure locally by ensuring a consistent multimodal alignment. Our experiments demonstrate that DiVE effectively preserves the geometric structure, achieving strong results across ID, OOD, and zero-shot metrics. Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane, Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
AAAI | 5 |
| 2025 | Multimodal Fine-Grained Apparent Personality Trait Recognition: Joint Modeling of Big Five and Questionnaire Item-level ScoresabstractThis paper presents a novel method for automatically recognizing people's apparent personality traits as perceived by others. In previous studies, apparent personality trait recognition from multimodal human behavior is often modeled to directly estimate personality trait scores, i.e., the ``Big Five'' scores. In the model training phase, ground-truth personality trait scores were often determined from personality test results scored by many other people using fine-grained questionnaires, however, rich information in the personality test results have not been leveraged for anything other than determining the ground-truth Big Five scores. The scores assigned to each questionnaire item are thought to include more meta-level differences in personality characteristics. Therefore, we propose joint modeling methods that can estimate not only the Big Five scores but also questionnaire item-level scores. This enables us to improve awareness of multimodal human behavior. In addition, we present a newly created self-introduction video dataset with 50-item Big Five questionnaire results since previous apparent personality trait recognition datasets do not provide such personality test results. Experiments using the created dataset demonstrate that our proposed joint modeling methods with a multimodal transformer backbone can improve to estimate Big Five scores and effectively estimate questionnaire item-level scores. We also verify that the estimation performance reached human evaluation performance. Ryo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Naoki Makishima, Saki Mizuno, Nobukatsu Hojo |
AAAI | 5 |
| 2025 | Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language ModelabstractThis paper proposes a personalization method for speech emotion recognition (SER) through in-context learning (ICL). Since the expression of emotions varies from person to person, speaker-specific adaptation is crucial for improving the SER performance. Conventional SER methods have been personalized using emotional utterances of a target speaker, but it is often difficult to prepare utterances corresponding to all emotion labels in advance. Our idea to overcome this difficulty is to obtain speaker characteristics by conditioning a few emotional utterances of the target speaker in ICL-based inference. ICL is a method to perform unseen tasks by conditioning a few inputoutput examples through inference in large language models (LLMs). We meta-train a speech-language model extended from the LLM to learn how to perform personalized SER via ICL. Experimental results using our newly collected SER dataset demonstrate that the proposed method outperforms conventional methods. Mana Ihori, Taiga Yamane, Naotaka Kawata, Naoki Makishima, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
ASRU | 4 |
| 2025 | Phoneme Overlapping-Aware Pre-Training with External Text Resources for Multi-Talker ASRabstractThis paper proposes a new pre-training method utilizing external text resources to improve the robustness of single-channel multi-talker automatic speech recognition (MTASR) across various linguistic domains. In the development of single-talker ASR systems, various pre-training methods have been studied to acquire knowledge about word order and correspondence between phonetic information and text from external text resources. However, methods focusing on improving MT-ASR performance by leveraging external text resources remain underexplored. To bridge this gap, we aim to acquire the ability to discover multiple texts contained within overlapping phonetic information. The key idea of the proposed method is to induce overlapping phenomena in phoneme sequences in order to reproduce a task similar to MT-ASR using external text resources. Our experiments demonstrate that the proposed method significantly improves the MT-ASR performance on both in-domain and out-of-domain linguistic tasks. Ryo Masumura, Tomohiro Tanaka, Naoki Makishima, Mana Ihori, Shota Orihashi, Naotaka Kawata, Taiga Yamane, Takafumi Moriya |
ASRU | 3 |
| 2025 | Unified Audio-Visual Modeling for Recognizing Which Face Spoke When and What in Multi-Talker Overlapped Speech and Video
Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
INTERSPEECH | 1 |
| 2025 | SOMSRED-SVC: Sequential Output Modeling with Speaker Vector Constraints for Joint Multi-Talker Overlapped ASR and Speaker Diarization
Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
INTERSPEECH | 1 |
| 2024 | SOMSRED: Sequential Output Modeling for Joint Multi-talker Overlapped Speech Recognition and Speaker Diarization
Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Atsushi Ando, Ryo Masumura |
INTERSPEECH | 1 |
| 2024 | Unified Multi-Talker ASR with and without Target-speaker Enrollment
Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Mana Ihori, Naotaka Kawata, Shota Orihashi, Kazutoshi Shinoda, Taiga Yamane, Saki Mizuno, Keita Suzuki, Nobukatsu Hojo, Takafumi Moriya, Atsushi Ando |
INTERSPEECH | 2 |
| 2023 | Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuracy-Robustness TradeoffabstractThis paper addresses the tradeoff between standard accuracy on clean examples and robustness against adversarial examples in deep neural networks (DNNs). Although adversarial training (AT) improves robustness, it degrades the standard accuracy, thus yielding the tradeoff. To mitigate this tradeoff, we propose a novel AT method called ARREST, which comprises three components: (i) adversarial finetuning (AFT), (ii) representation-guided knowledge distillation (RGKD), and (iii) noisy replay (NR). AFT trains a DNN on adversarial examples by initializing its parameters with a DNN that is standardly pretrained on clean examples. RGKD and NR respectively entail a regularization term and an algorithm to preserve latent representations of clean examples during AFT. RGKD penalizes the distance between the representations of the standardly pretrained and AFT DNNs. NR switches input adversarial examples to nonadversarial ones when the representation changes significantly during AFT. By combining these components, ARREST achieves both high standard accuracy and robustness. Experimental results demonstrate that ARREST mitigates the tradeoff more effectively than previous AT-based methods do. Shin'ya Yamaguchi, Shoichiro Takeda, Sekitoshi Kanai, Naoki Makishima, Atsushi Ando, Ryo Masumura |
ICCV | 5 |
| 2023 | OnDA-DETR: Online Domain Adaptation for Detection Transformers with Self-Training FrameworkabstractThis paper presents a novel method for online domain adaptation (OnDA) for DEtection TRansformer (DETR)-based object detection models called OnDA-DETR. OnDA is a domain adaptation paradigm that adapts a model trained on the source domain data to perform well on the target domain in an online manner during testing, using only the unlabeled test data from the target domain. Due to challenging and realistic problem settings, OnDA has garnered significant attention. However, OnDA methods for DETR-based models, which have demonstrated excellent performance in object detection research fields, had not been developed. OnDA-DETR is the first OnDA method specifically designed for DETR-based models. OnDA-DETR incorporates a self-training framework that generates pseudo-labels for the unlabeled target domain data. To effectively incorporate the self-training framework into DETR-based models, we leverage recall-aware pseudo-labeling and quality-aware training in OnDA-DETR. Experimental results indicate that OnDA-DETR improves the performance of the source-trained model by about 3.0 % points through OnDA. Taiga Yamane, Naoki Makishima, Keita Suzuki, Atsushi Ando, Ryo Masumura |
ICIP | 3 |
| 2023 | Joint Autoregressive Modeling of End-to-End Multi-Talker Overlapped Speech Recognition and Utterance-level Timestamp Prediction
Naoki Makishima, Keita Suzuki, Atsushi Ando, Ryo Masumura |
INTERSPEECH | 1 |
| 2023 | End-to-End Joint Target and Non-Target Speakers ASR
Ryo Masumura, Naoki Makishima, Taiga Yamane, Yoshihiko Yamazaki, Saki Mizuno, Mana Ihori, Mihiro Uchida, Keita Suzuki, Hiroshi Sato 0002, Tomohiro Tanaka, Akihiko Takashima, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando |
INTERSPEECH | 2 |
| 2023 | Multi-region CNN-Transformer for Micro-gesture Recognition in Face and Upper BodyabstractThis paper presents a novel task that recognizes from a video unintentional micro-gestures (UMGs), which are movements made by people unconsciously and unintentionally. Recognizing UMGs is crucial because they reveal a person’s underlying psychological state. Since a UMG is composed of subtle sequential movements, the recognition model must be able to capture accurate information in both the spatial and temporal directions. Therefore, we utilize a convolutional neural network (CNN) to capture information in the spatial direction and a Transformer to merge the features extracted by the CNN in the temporal direction. However, this model often misrecognizes UMGs because it is not possible to capture slight differences in movements, such as in the face and mouth regions. To address this issue, we propose a novel model for UMG recognition, the Multi-Region CNN-Transformer model, that inputs cropped videos from multiple upper body regions simultaneously. The key advance of our method is to capture subtle changes in regions such as the upper body, face, head, and mouth for recognizing UMGs. We demonstrate the effectiveness of the proposed method through experiments using our newly created UMG dataset for this task. Keita Suzuki, Ryo Masumura, Atsushi Ando, Naoki Makishima |
MMAsia | 5 |
| 2022 | Customer Satisfaction Estimation Using Unsupervised Representation Learning with Multi-Format Prediction LossabstractWe propose a new Customer Satisfaction Estimation (CSE) method that utilizes unsupervised representation learning. Though conventional methods have improved both the heuristic features and the estimation models, their performance is still insufficient as only small amounts of labeled training data can be expected. To mitigate this problem, the proposed method leverages a large amount of unlabeled data by unsupervised representation learning based on self-training. The key advance of the proposed method is to introduce a Multi-Format Prediction (MFP) loss to improve the performance of self-training for the inputs that contain both continuous and biased discrete features such as the number of occurrences of a particular word. MFP loss uses two loss functions based on regression and weighted binary classification to reconstruct both types of features with high accuracy. Experiments on real English contact center calls reveal the improved CSE performance attained by the proposed method. Atsushi Ando, Yumiko Murata, Ryo Masumura, Naoki Makishima, Takafumi Moriya, Takanori Ashihara, Hiroshi Sato 0002 |
ICASSP | 5 |
| 2022 | Speaker consistency loss and step-wise optimization for semi-supervised joint training of TTS and ASR using unpaired text dataabstractIn this paper, we investigate the semi-supervised joint training of text to speech (TTS) and automatic speech recognition (ASR), where a small amount of paired data and a large amount of unpaired text data are available.Conventional studies form a cycle called the TTS-ASR pipeline, where the multispeaker TTS model synthesizes speech from text with a reference speech and the ASR model reconstructs the text from the synthesized speech, after which both models are trained with a cycle-consistency loss.However, the synthesized speech does not reflect the speaker characteristics of the reference speech and the synthesized speech becomes overly easy for the ASR model to recognize after training.This not only decreases the TTS model quality but also limits the ASR model improvement.To solve this problem, we propose improving the cycleconsistency-based training with a speaker consistency loss and step-wise optimization.The speaker consistency loss brings the speaker characteristics of the synthesized speech closer to that of the reference speech.In the step-wise optimization, we first freeze the parameter of the TTS model before both models are trained to avoid over-adaptation of the TTS model to the ASR model.Experimental results demonstrate the efficacy of the proposed method. Naoki Makishima, Atsushi Ando, Ryo Masumura |
INTERSPEECH | 1 |
| 2022 | End-to-End Joint Modeling of Conversation History-Dependent and Independent ASR Systems with Multi-History Training
Ryo Masumura, Yoshihiro Yamazaki, Saki Mizuno, Naoki Makishima, Mana Ihori, Mihiro Uchida, Hiroshi Sato 0002, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando |
INTERSPEECH | 4 |
| 2022 | Strategies to Improve Robustness of Target Speech Extraction to Enrollment VariationsabstractTarget speech extraction is a technique to extract the target speaker's voice from mixture signals using a pre-recorded enrollment utterance that characterize the voice characteristics of the target speaker.One major difficulty of target speech extraction lies in handling variability in "intra-speaker" characteristics, i.e., characteristics mismatch between target speech and an enrollment utterance.While most conventional approaches focus on improving average performance given a set of enrollment utterances, here we propose to guarantee the worst performance, which we believe is of great practical importance.In this work, we propose an evaluation metric called worstenrollment source-to-distortion ratio (SDR) to quantitatively measure the robustness towards enrollment variations.We also introduce a novel training scheme that aims at directly optimizing the worst-case performance by focusing on training with difficult enrollment cases where extraction does not perform well.In addition, we investigate the effectiveness of auxiliary speaker identification loss (SI-loss) as another way to improve robustness over enrollments.Experimental validation reveals the effectiveness of both worst-enrollment target training and SI-loss training to improve robustness against enrollment variations, by increasing speaker discriminability. Hiroshi Sato 0002, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Takafumi Moriya, Naoki Makishima, Mana Ihori, Tomohiro Tanaka, Ryo Masumura |
INTERSPEECH | 6 |
| 2022 | On the Use of Modality-Specific Large-Scale Pre-Trained Encoders for Multimodal Sentiment AnalysisabstractThis paper investigates the effectiveness and implementation of modality-specific large-scale pre-trained encoders for multimodal sentiment analysis (MSA). Although the effectiveness of pre-trained encoders in various fields has been reported, conventional MSA methods employ them for only linguistic modality, and their application has not been investigated. This paper compares the features yielded by large-scale pre-trained encoders with conventional heuristic features. One each of the largest pre-trained encoders publicly available for each modality are used; CLIP-ViT, WavLM, and BERT for visual, acoustic, and linguistic modalities, respectively. Experiments on two datasets reveal that methods with domain-specific pre-trained encoders attain better performance than those with conventional features in both unimodal and multimodal scenarios. We also find it better to use the outputs of the intermediate layers of the encoders than those of the output layer. The codes are available at https://github.com/ando-hub/MSA_Pretrain. Atsushi Ando, Ryo Masumura, Akihiko Takashima, Naoki Makishima, Keita Suzuki, Takafumi Moriya, Takanori Ashihara, Hiroshi Sato 0002 |
SLT | 5 |
| 2021 | Hierarchical Knowledge Distillation for Dialogue Sequence LabelingabstractThis paper presents a novel knowledge distillation method for dialogue sequence labeling. Dialogue sequence labeling is a supervised learning task that estimates labels for each utterance in the target dialogue document, and is useful for many applications such as dialogue act estimation. Accurate labeling is often realized by a hierarchically-structured large model consisting of utterance-level and dialogue-level networks that capture the contexts within an utterance and between utterances, respectively. However, due to its large model size, such a model cannot be deployed on resource-constrained devices. To overcome this difficulty, we focus on knowledge distillation which trains a small model by distilling the knowledge of a large and high performance teacher model. Our key idea is to distill the knowledge while keeping the complex contexts captured by the teacher model. To this end, the proposed method, hierarchical knowledge distillation, trains the small model by distilling not only the probability distribution of the label classification, but also the knowledge of utterance-level and dialogue-level contexts trained in the teacher model by training the model to mimic the teacher model's output in each level. Experiments on dialogue act estimation and call scene segmentation demonstrate the effectiveness of the proposed method. Shota Orihashi, Yoshihiro Yamazaki, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Ryo Masumura |
ASRU | 3 |
| 2021 | MAPGN: Masked Pointer-Generator Network for Sequence-to-Sequence Pre-TrainingabstractThis paper presents a self-supervised learning method for pointer-generator networks to improve spoken-text normalization. Spoken-text normalization that converts spoken-style text into style normalized text is becoming an important technology for improving subsequent processing such as machine translation and summarization. The most successful spoken-text normalization method to date is sequence-to-sequence (seq2seq) mapping using pointer-generator networks that possess a copy mechanism from an input sequence. However, these models require a large amount of paired data of spoken-style text and style normalized text, and it is difficult to prepare such a volume of data. In order to construct spoken-text normalization model from the limited paired data, we focus on self-supervised learning which can utilize unpaired text data to improve seq2seq models. Unfortunately, conventional self-supervised learning methods do not assume that pointer-generator networks are utilized. Therefore, we propose a novel self-supervised learning method, MAsked Pointer-Generator Network (MAPGN). The proposed method can effectively pre-train the pointer-generator net-work by learning to fill masked tokens using the copy mechanism. Our experiments demonstrate that MAPGN is more effective for pointer-generator networks than the conventional self-supervised learning methods in two spoken-text normalization tasks. Mana Ihori, Naoki Makishima, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Ryo Masumura |
ICASSP | 2 |
| 2021 | Audio-Visual Speech Separation Using Cross-Modal Correspondence LossabstractWe present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual speech separation is a technique to estimate the individual speech signals from a mixture using the visual signals of the speakers. Conventional studies on audio-visual speech separation mainly train the separation model on the audio-only loss, which reflects the distance between the source signals and the separated signals. However, conventional losses do not reflect the characteristics of the speech signals, including the speaker’s characteristics and phonetic information, which leads to distortion or remaining noise. To address this problem, we propose the cross-modal correspondence (CMC) loss, which is based on the cooccurrence of the speech signal and the visual signal. Since the visual signal is not affected by background noise and contains speaker and phonetic information, using the CMC loss enables the audio-visual speech separation model to remove noise while preserving the speech characteristics. Experimental results demonstrate that the proposed method learns the cooccurrence on the basis of CMC loss, which improves separation performance. Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
ICASSP | 1 |
| 2021 | Hierarchical Transformer-Based Large-Context End-To-End ASR with Large-Context Knowledge DistillationabstractWe present a novel large-context end-to-end automatic speech recognition (E2E-ASR) model and its effective training method based on knowledge distillation. Common E2E-ASR models have mainly focused on utterance-level processing in which each utterance is independently transcribed. On the other hand, large-context E2E-ASR models, which take into account long-range sequential contexts beyond utterance boundaries, well handle a sequence of utterances such as discourses and conversations. However, the transformer architecture, which has recently achieved state-of-the-art ASR performance among utterance-level ASR systems, has not yet been introduced into the large-context ASR systems. We can expect that the transformer architecture can be leveraged for effectively capturing not only input speech contexts but also long-range sequential contexts beyond utterance boundaries. Therefore, this paper proposes a hierarchical transformer-based large-context E2E-ASR model that combines the transformer architecture with hierarchical encoder-decoder based large-context modeling. In addition, in order to enable the proposed model to use long-range sequential contexts, we also propose a large-context knowledge distillation that distills the knowledge from a pre-trained large-context language model in the training phase. We evaluate the effectiveness of the proposed model and proposed training method on Japanese discourse ASR tasks. Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
ICASSP | 2 |
| 2021 | Zero-Shot Joint Modeling of Multiple Spoken-Text-Style Conversion Tasks Using Switching TokensabstractIn this paper, we propose a novel spoken-text-style conversion method that can simultaneously execute multiple style conversion modules such as punctuation restoration and disfluency deletion without preparing matched datasets.In practice, transcriptions generated by automatic speech recognition systems are not highly readable because they often include many disfluencies and do not include punctuation marks.To improve their readability, multiple spoken-text-style conversion modules that individually model a single conversion task are cascaded because matched datasets that simultaneously handle multiple conversion tasks are often unavailable.However, the cascading is unstable against the order of tasks because of the chain of conversion errors.Besides, the computation cost of the cascading must be higher than the single conversion.To execute multiple conversion tasks simultaneously without preparing matched datasets, our key idea is to distinguish individual conversion tasks using the on-off switch.In our proposed zero-shot joint modeling, we switch the individual tasks using multiple switching tokens, enabling us to utilize a zero-shot learning approach to executing simultaneous conversions.Our experiments on joint modeling of disfluency deletion and punctuation restoration demonstrate the effectiveness of our method. Mana Ihori, Naoki Makishima, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Ryo Masumura |
Interspeech | 2 |
| 2021 | Enrollment-Less Training for Personalized Voice Activity DetectionabstractWe present a novel personalized voice activity detection (PVAD) learning method that does not require enrollment data during training.PVAD is a task to detect the speech segments of a specific target speaker at the frame level using enrollment speech of the target speaker.Since PVAD must learn speakers' speech variations to clarify the boundary between speakers, studies on PVAD used large-scale datasets that contain many utterances for each speaker.However, the datasets to train a PVAD model are often limited because substantial cost is needed to prepare such a dataset.In addition, we cannot utilize the datasets used to train the standard VAD because they often lack speaker labels.To solve these problems, our key idea is to use one utterance as both a kind of enrollment speech and an input to the PVAD during training, which enables PVAD training without enrollment speech.In our proposed method, called enrollment-less training, we augment one utterance so as to create variability between the input and the enrollment speech while keeping the speaker identity, which avoids the mismatch between training and inference.Our experimental results demonstrate the efficacy of the method. Naoki Makishima, Mana Ihori, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Ryo Masumura |
Interspeech | 1 |
| 2021 | Unified Autoregressive Modeling for Joint End-to-End Multi-Talker Overlapped Speech Recognition and Speaker Attribute EstimationabstractIn this paper, we present a novel modeling method for singlechannel multi-talker overlapped automatic speech recognition (ASR) systems.Fully neural network based end-to-end models have dramatically improved the performance of multi-taker overlapped ASR tasks.One promising approach for end-toend modeling is autoregressive modeling with serialized output training in which transcriptions of multiple speakers are recursively generated one after another.This enables us to naturally capture relationships between speakers.However, the conventional modeling method cannot explicitly take into account the speaker attributes of individual utterances such as gender and age information.In fact, the performance deteriorates when each speaker is the same gender or is close in age.To address this problem, we propose unified autoregressive modeling for joint end-to-end multi-talker overlapped ASR and speaker attribute estimation.Our key idea is to handle gender and age estimation tasks within the unified autoregressive modeling.In the proposed method, transformer-based autoregressive model recursively generates not only textual tokens but also attribute tokens of each speaker.This enables us to effectively utilize speaker attributes for improving multi-talker overlapped ASR.Experiments on Japanese multi-talker overlapped ASR tasks demonstrate the effectiveness of the proposed method. Ryo Masumura, Daiki Okamura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
Interspeech | 3 |
| 2021 | Cross-Modal Transformer-Based Neural Correction Models for Automatic Speech RecognitionabstractWe propose a cross-modal transformer-based neural correction models that refines the output of an automatic speech recognition (ASR) system so as to exclude ASR errors.Generally, neural correction models are composed of encoder-decoder networks, which can directly model sequence-to-sequence mapping problems.The most successful method is to use both input speech and its ASR output text as the input contexts for the encoder-decoder networks.However, the conventional method cannot take into account the relationships between these two different modal inputs because the input contexts are separately encoded for each modal.To effectively leverage the correlated information between the two different modal inputs, our proposed models encode two different contexts jointly on the basis of cross-modal self-attention using a transformer.We expect that cross-modal self-attention can effectively capture the relationships between two different modals for refining ASR hypotheses.We also introduce a shallow fusion technique to efficiently integrate the first-pass ASR model and our proposed neural correction model.Experiments on Japanese natural language ASR tasks demonstrated that our proposed models achieve better ASR performance than conventional neural correction models. Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Akihiko Takashima, Takafumi Moriya, Takanori Ashihara, Shota Orihashi, Naoki Makishima |
Interspeech | 8 |
| 2021 | End-to-End Rich Transcription-Style Automatic Speech Recognition with Semi-Supervised LearningabstractWe propose a semi-supervised learning method for building end-to-end rich transcription-style automatic speech recognition (RT-ASR) systems from small-scale rich transcriptionstyle and large-scale common transcription-style datasets.In spontaneous speech tasks, various speech phenomena such as fillers, word fragments, laughter and coughs, etc. are often included.While common transcriptions do not give special awareness to these phenomena, rich transcriptions explicitly convert them into special phenomenon tokens as well as textual tokens.In previous studies, the textual and phenomenon tokens were simultaneously estimated in an end-to-end manner.However, it is difficult to build accurate RT-ASR systems because large-scale rich transcription-style datasets are often unavailable.To solve this problem, our training method uses a limited rich transcription-style dataset and common transcriptionstyle dataset simultaneously.The Key process in our semisupervised learning is to convert the common transcription-style dataset into a pseudo-rich transcription-style dataset.To this end, we introduce style tokens which control phenomenon tokens are generated or not into transformer-based autoregressive modeling.We use this modeling for generating the pseudorich transcription-style datasets and for building RT-ASR system from the pseudo and original datasets.Our experiments on spontaneous ASR tasks showed the effectiveness of the proposed method. Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Akihiko Takashima, Shota Orihashi, Naoki Makishima |
Interspeech | 6 |
| 2021 | Utilizing Resource-Rich Language Datasets for End-to-End Scene Text Recognition in Resource-Poor LanguagesabstractThis paper presents a novel training method for end-to-end scene text recognition. End-to-end scene text recognition offers high recognition accuracy, especially when using the encoder-decoder model based on Transformer. To train a highly accurate end-to-end model, we need to prepare a large image-to-text paired dataset for the target language. However, it is difficult to collect this data, especially for resource-poor languages. To overcome this difficulty, our proposed method utilizes well-prepared large datasets in resource-rich languages such as English, to train the resource-poor encoder-decoder model. Our key idea is to build a model in which the encoder reflects knowledge of multiple languages while the decoder specializes in knowledge of just the resource-poor language. To this end, the proposed method pre-trains the encoder by using a multilingual dataset that combines the resource-poor language’s dataset and the resource-rich language’s dataset to learn language-invariant knowledge for scene text recognition. The proposed method also pre-trains the decoder by using the resource-poor language’s dataset to make the decoder better suited to the resource-poor language. Experiments on Japanese scene text recognition using a small, publicly available dataset demonstrate the effectiveness of the proposed method. Shota Orihashi, Yoshihiro Yamazaki, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Ryo Masumura |
MMAsia | 3 |
| 2021 | Large-Context Conversational Representation Learning: Self-Supervised Learning For Conversational DocumentsabstractThis paper presents a novel self-supervised learning method for handling conversational documents consisting of transcribed text of human-to-human conversations. One of the key technologies for understanding conversational documents is utterance-level sequential labeling, where labels are estimated from the documents in an utterance-by-utterance manner. The main issue with utterance-level sequential labeling is the difficulty of collecting labeled conversational documents, as manual annotations are very costly. To deal with this issue, we propose large-context conversational representation learning (LC-CRL), a self-supervised learning method specialized for conversational documents. A self-supervised learning task in LC-CRL involves the estimation of an utterance using all the surrounding utterances based on large-context language modeling. In this way, LC-CRL enables us to effectively utilize unlabeled conversational documents and thereby enhances the utterance-level sequential labeling. The results of experiments on scene segmentation tasks using contact center conversational datasets demonstrate the effectiveness of the proposed method. Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
SLT | 2 |
| 2021 | Independent deeply learned matrix analysis with automatic selection of stable microphone-wise update and fast sourcewise update of demixing matrixabstractIndependent deeply learned matrix analysis (IDLMA) is a fast and high-performance method for multichannel audio source separation. IDLMA utilizes the deep neural network inference of source models and the blind estimation of demixing filters based on source independence. In conventional IDLMA, iterative projection (IP) is exploited to estimate the demixing filters. Although IP is a fast algorithm, it sometimes fails to estimate an appropriate solution. This is because IP updates the demixing filters in a sourcewise manner, where only one source model is used for each update, and the update sometimes becomes unstable owing to the specific low-quality source models. In this paper, we first derive a new numerically stable microphone-wise update algorithm that exploits all source model information simultaneously. The microphone-wise update problem cannot be solved by IP; instead, a new type of vectorwise coordinate descent algorithm is introduced. Next, comparison analysis of the proposed microphone-wise update and IP reveals the tradeoff w.r.t. convergence speed and numerical stability. To resolve this tradeoff problem, we propose the automatic selection of update rules on the basis of the likelihood function of observed signals. Finally, experimental results show the efficacy of the proposed IDLMA with the automatic selection of update rules. Naoki Makishima, Yoshiki Mitsui, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo |
Signal Process. | 1 |
| 2020 | Memory Attentive Fusion: External Language Model Integration for Transformer-based Sequence-to-Sequence ModelabstractThis paper presents a novel fusion method for integrating an external language model (LM) into the Transformer based sequenceto-sequence (seq2seq) model.While paired data are basically required to train the seq2seq model, the external LM can be trained with only unpaired data.Thus, it is important to leverage memorized knowledge in the external LM for building the seq2seq model, since it is hard to prepare a large amount of paired data.However, the existing fusion methods assume that the LM is integrated with recurrent neural network-based seq2seq models instead of the Transformer.Therefore, this paper proposes a fusion method that can explicitly utilize network structures in the Transformer.The proposed method, called memory attentive fusion, leverages the Transformer-style attention mechanism that repeats source-target attention in a multi-hop manner for reading the memorized knowledge in the LM.Our experiments on two text-style conversion tasks demonstrate that the proposed method performs better than conventional fusion methods. Mana Ihori, Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi |
INLG | 3 |
| 2020 | Phoneme-to-Grapheme Conversion Based Large-Scale Pre-Training for End-to-End Automatic Speech Recognition
Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
INTERSPEECH | 2 |
| 2019 | Independent Deeply Learned Matrix Analysis for Determined Audio Source SeparationabstractIn this paper, we propose a new framework called independent deeply learned matrix analysis (IDLMA), which unifies a deep neural network (DNN) and independence-based multichannel audio source separation. IDLMA utilizes both pretrained DNN source models and statistical independence between sources for the separation, where the time-frequency structures of each source are iteratively optimized by a DNN while enhancing the estimation accuracy of the spatial demixing filters. As the source generative model, we introduce a complex heavy-tailed distribution to improve the separation performance. In addition, we address a semi-supervised situation; namely, a solo-recorded audio dataset can be prepared for only one source in the mixture signal. To solve the limited-data problem, we propose an appropriate data augmentation method to adapt the DNN source models to the observed signal, which enables IDLMA to work even in the semi-supervised situation. Experiments are conducted using music signals with a training dataset in both supervised and semi-supervised situations. The results show the validity of the proposed method in terms of the separation accuracy. Naoki Makishima, Shinichi Mogami, Norihiro Takamune, Daichi Kitamura, Hayato Sumino, Shinnosuke Takamichi, Hiroshi Saruwatari, Nobutaka Ono |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |