VLDB 2026 Research / reviewers in the wild / expert
Ryo Masumura
dblp:08/10650
· DBLP profile ↗
116ranked-venue papers
33as first author
66since 2021 · last 2026
0000-0002-2415-4149ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 102 · 29 first-author · 63 since 2021Artificial intelligence and machine learning · 82 · 25 first-author · 42 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Difference Vector Equalization for Robust Fine-tuning of Vision-Language ModelsabstractContrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image and text encoders. This paper aims to robustly fine-tune these vision-language models on in-distribution (ID) data without compromising their generalization abilities in out-of-distribution (OOD) and zero-shot settings. Current robust fine-tuning methods tackle this challenge by reusing contrastive learning, which was used in pre-training, for fine-tuning. However, we found that these methods distort the geometric structure of the embeddings, which plays a crucial role in the generalization of vision-language models, resulting in limited OOD and zero-shot performance. To address this, we propose Difference Vector Equalization (DiVE), which preserves the geometric structure during fine-tuning. The idea behind DiVE is to constrain difference vectors, each of which is obtained by subtracting the embeddings extracted from the pre-trained and fine-tuning models for the same data sample. By constraining the difference vectors to be equal across various data samples, we effectively preserve the geometric structure. Therefore, we introduce two losses: average vector loss (AVL) and pairwise vector loss (PVL). AVL preserves the geometric structure globally by constraining difference vectors to be equal to their weighted average. PVL preserves the geometric structure locally by ensuring a consistent multimodal alignment. Our experiments demonstrate that DiVE effectively preserves the geometric structure, achieving strong results across ID, OOD, and zero-shot metrics. Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane, Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
AAAI | 10 |
| 2026 | Distribution Highlighted Reference-based Label Distribution Learning for Facial Age EstimationabstractEstimating age from facial images is a fundamental task. In this task, age labels have ambiguity because faces of the same individual across similar ages are often difficult to distinguish. To model this ambiguity, label distribution learning (LDL) trains a deep neural network (DNN) using a label distribution, i.e., the probability that an image belongs to each age, instead of a single age label. However, the heuristic constraints utilized for LDL often fail to accurately model the label ambiguity. Therefore, we propose a novel LDL method called distribution highlighted reference-based LDL (DHRL), which introduces an input-dependent constraint by utilizing a reference DNN pre-trained with any LDL method and minimizing the gap between the reference and target DNNs’ outputs. DHRL incorporates two techniques to highlight the label ambiguity hidden in the pre-trained reference DNN’s output: noisy augmentation-based ensembling (NAE) and different scale multi-temperature (DSM). NAE inputs noisy images to the reference DNN and provides an ensemble effect by averaging all the outputs. DSM sets multiple temperatures simultaneously in the gap minimization between the two DNNs’ outputs. Experimental results indicate that our method achieves state-of-the-art performance across various datasets and conditions. Shin'ya Yamaguchi, Shoichiro Takeda, Takuhiro Kaneko, Shota Orihashi, Ryo Masumura |
WACV | 6 |
| 2025 | Multimodal Fine-Grained Apparent Personality Trait Recognition: Joint Modeling of Big Five and Questionnaire Item-level ScoresabstractThis paper presents a novel method for automatically recognizing people's apparent personality traits as perceived by others. In previous studies, apparent personality trait recognition from multimodal human behavior is often modeled to directly estimate personality trait scores, i.e., the ``Big Five'' scores. In the model training phase, ground-truth personality trait scores were often determined from personality test results scored by many other people using fine-grained questionnaires, however, rich information in the personality test results have not been leveraged for anything other than determining the ground-truth Big Five scores. The scores assigned to each questionnaire item are thought to include more meta-level differences in personality characteristics. Therefore, we propose joint modeling methods that can estimate not only the Big Five scores but also questionnaire item-level scores. This enables us to improve awareness of multimodal human behavior. In addition, we present a newly created self-introduction video dataset with 50-item Big Five questionnaire results since previous apparent personality trait recognition datasets do not provide such personality test results. Experiments using the created dataset demonstrate that our proposed joint modeling methods with a multimodal transformer backbone can improve to estimate Big Five scores and effectively estimate questionnaire item-level scores. We also verify that the estimation performance reached human evaluation performance. Ryo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Naoki Makishima, Saki Mizuno, Nobukatsu Hojo |
AAAI | 1 |
| 2025 | ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of MindabstractExisting Theory of Mind (ToM) benchmarks diverge from real-world scenarios in three aspects: 1) they assess a limited range of mental states such as beliefs, 2) false beliefs are not comprehensively explored, and 3) the diverse personality traits of characters are overlooked. To address these challenges, we introduce ToMATO, a new ToM benchmark formulated as multiple-choice QA over conversations. ToMATO is generated via LLM-LLM conversations featuring information asymmetry. By employing a prompting method that requires role-playing LLMs to verbalize their thoughts before each utterance, we capture both first- and second-order mental states across five categories: belief, intention, desire, emotion, and knowledge. These verbalized thoughts serve as answers to questions designed to assess the mental states of characters within conversations. Furthermore, the information asymmetry introduced by hiding thoughts from others induces the generation of false beliefs about various mental states. Assigning distinct personality traits to LLMs further diversifies both utterances and thoughts. ToMATO consists of 5.4k questions, 753 conversations, and 15 personality trait patterns. Our analysis shows that this dataset construction approach frequently generates false beliefs due to the information asymmetry between role-playing LLMs, and effectively reflects diverse personalities. We evaluate nine LLMs on ToMATO and find that even GPT-4o mini lags behind human performance, especially in understanding false beliefs, and lacks robustness to various personality traits. Kazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno, Keita Suzuki, Ryo Masumura, Hiroaki Sugiyama, Kuniko Saito |
AAAI | 6 |
| 2025 | Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language ModelabstractThis paper proposes a personalization method for speech emotion recognition (SER) through in-context learning (ICL). Since the expression of emotions varies from person to person, speaker-specific adaptation is crucial for improving the SER performance. Conventional SER methods have been personalized using emotional utterances of a target speaker, but it is often difficult to prepare utterances corresponding to all emotion labels in advance. Our idea to overcome this difficulty is to obtain speaker characteristics by conditioning a few emotional utterances of the target speaker in ICL-based inference. ICL is a method to perform unseen tasks by conditioning a few inputoutput examples through inference in large language models (LLMs). We meta-train a speech-language model extended from the LLM to learn how to perform personalized SER via ICL. Experimental results using our newly collected SER dataset demonstrate that the proposed method outperforms conventional methods. Mana Ihori, Taiga Yamane, Naotaka Kawata, Naoki Makishima, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
ASRU | 8 |
| 2025 | Phoneme Overlapping-Aware Pre-Training with External Text Resources for Multi-Talker ASRabstractThis paper proposes a new pre-training method utilizing external text resources to improve the robustness of single-channel multi-talker automatic speech recognition (MTASR) across various linguistic domains. In the development of single-talker ASR systems, various pre-training methods have been studied to acquire knowledge about word order and correspondence between phonetic information and text from external text resources. However, methods focusing on improving MT-ASR performance by leveraging external text resources remain underexplored. To bridge this gap, we aim to acquire the ability to discover multiple texts contained within overlapping phonetic information. The key idea of the proposed method is to induce overlapping phenomena in phoneme sequences in order to reproduce a task similar to MT-ASR using external text resources. Our experiments demonstrate that the proposed method significantly improves the MT-ASR performance on both in-domain and out-of-domain linguistic tasks. Ryo Masumura, Tomohiro Tanaka, Naoki Makishima, Mana Ihori, Shota Orihashi, Naotaka Kawata, Taiga Yamane, Takafumi Moriya |
ASRU | 1 |
| 2025 | All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASRabstractThis paper proposes a unified framework, All-in-One ASR, that allows a single model to support multiple automatic speech recognition (ASR) paradigms, including connectionist temporal classification (CTC), attention-based encoder-decoder (AED), and Transducer, in both offline and streaming modes. While each ASR architecture offers distinct advantages and trade-offs depending on the application, maintaining separate models for each scenario incurs substantial development and deployment costs. To address this issue, we introduce a multi-mode joiner that enables seamless integration of various ASR modes within a single unified model. Experiments show that All-in-One ASR significantly reduces the total model footprint while matching or even surpassing the recognition performance of individually optimized ASR models. Furthermore, joint decoding leverages the complementary strengths of different ASR modes, yielding additional improvements in recognition accuracy. Takafumi Moriya, Masato Mimura, Tomohiro Tanaka, Hiroshi Sato 0002, Ryo Masumura, Atsunori Ogawa |
ASRU | 5 |
| 2025 | Alignment-Free Training for Transducer-based Multi-Talker ASRabstractExtending the RNN Transducer (RNNT) to recognize multi-talker speech is essential for wider automatic speech recognition (ASR) applications. Multi-talker RNNT (MT-RNNT) aims to achieve recognition without relying on costly front-end source separation. MT-RNNT is conventionally implemented using architectures with multiple encoders or decoders, or by serializing all speakers’ transcriptions into a single output stream. The first approach is computationally expensive, particularly due to the need for multiple encoder processing. In contrast, the second approach involves a complex label generation process, requiring accurate timestamps of all words spoken by all speakers in the mixture, obtained from an external ASR system. In this paper, we propose a novel alignment-free training scheme for the MT-RNNT (MT-RNNT-AFT) that adopts the standard RNNT architecture. The target labels are created by appending a prompt token corresponding to each speaker at the beginning of the transcription, reflecting the order of each speaker’s appearance in the mixtures. Thus, MT-RNNT-AFT can be trained without relying on accurate alignments, and it can recognize all speakers’ speech with just one round of encoder processing. Experiments show that MT-RNNT-AFT achieves performance comparable to that of the state-of-the-art alternatives, while greatly simplifying the training process. Takafumi Moriya, Shota Horiguchi, Marc Delcroix, Ryo Masumura, Takanori Ashihara, Hiroshi Sato 0002, Kohei Matsuura, Masato Mimura |
ICASSP | 4 |
| 2025 | MVTrajecter: Multi-View Pedestrian Tracking With Trajectory Motion Cost and Trajectory Appearance CostabstractMulti-View Pedestrian Tracking (MVPT) aims to track pedestrians in the form of a bird's eye view occupancy map from multi-view videos. End-to-end methods that detect and associate pedestrians within one model have shown great progress in MVPT. The motion and appearance information of pedestrians is important for the association, but previous end-to-end MVPT methods rely only on the current and its single adjacent past timestamp, discarding the past trajectories before that. This paper proposes a novel end-to-end MVPT method called Multi-View Trajectory Tracker (MVTrajecter) that utilizes information from multiple timestamps in past trajectories for robust association. MVTrajecter introduces trajectory motion cost and trajectory appearance cost to effectively incorporate motion and appearance information, respectively. These costs calculate which pedestrians at the current and each past timestamp are likely identical based on the information between those timestamps. Even if a current pedestrian could be associated with a false pedestrian at some past timestamp, these costs enable the model to associate that current pedestrian with the correct past trajectory based on other past timestamps. In addition, MVTrajecter effectively captures the relationships between multiple timestamps leveraging the attention mechanism. Extensive experiments demonstrate the effectiveness of each component in MVTrajecter and show that it outperforms the previous state-of-the-art methods. Taiga Yamane, Ryo Masumura, Shota Orihashi |
ICCV | 2 |
| 2025 | Leveraging LLMs for Written to Spoken Style Data Transformation to Enhance Spoken Dialog State Tracking
Haris Gulzar, Monikka Roslianna Busto, Akiko Masaki, Takeharu Eda, Ryo Masumura |
INTERSPEECH | 5 |
| 2025 | Unified Audio-Visual Modeling for Recognizing Which Face Spoke When and What in Multi-Talker Overlapped Speech and Video
Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
INTERSPEECH | 8 |
| 2025 | SOMSRED-SVC: Sequential Output Modeling with Speaker Vector Constraints for Joint Multi-Talker Overlapped ASR and Speaker Diarization
Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
INTERSPEECH | 8 |
| 2024 | Talking Face Generation for Impression Conversion Considering Speech SemanticsabstractThis study investigates the talking face generation method to convert a speaker’s video to give a target impression, such as “favorable” or “considerate”. Such an impression conversion method needs to consider the input speech semantics because they affect the impression of a speaker’s video along with the facial expression. Conventional emotional talking face generation methods utilize speech information to synchronize the lip and speech of the output video. However, they cannot consider speech semantics because the speech representations contain only phonetic information. To solve this problem, we propose a facial expression conversion model that uses a semantic vector obtained from BERT embeddings of speech recognition results of input speech. We first constructed an audio-visual dataset with impression labels assigned to each utterance. The evaluation results based on the dataset showed that the proposed method could improve the estimation accuracy of the facial expressions of the target video. Saki Mizuno, Nobukatsu Hojo, Kazutoshi Shinoda, Keita Suzuki, Mana Ihori, Tomohiro Tanaka, Naotaka Kawata, Satoshi Kobashikawa, Ryo Masumura |
ICASSP | 10 |
| 2024 | Scene Generalized Multi-View Pedestrian Detection with Rotation-Based Augmentation and RegularizationabstractMulti-view pedestrian detection aims to predict a bird’s eye view (BEV) occupancy map using multiple camera views. While existing deep learning-based methods have shown progress, they are typically trained and tested on the same scene, which has the same camera layout and number of cameras. As a result, they are difficult to generalize to new scenes. A dataset containing multiple scenes has recently been proposed to overcome this limitation, but even with this dataset, the performance on new scenes remains poor due to overfitting to the limited training scenes. To address this problem, we propose a novel data augmentation and regularization method for multi-view pedestrian detection. The unique point of our method is that it rotates the BEV features to generate the features originating from a new scene. This enables our method to effectively train the detection model with the knowledge of new scenes and prevent overfitting to specific scenes through regularization. Experimental results show that our method outperforms existing methods in terms of generalization to new scenes. Shotaro Tora, Ryo Masumura |
ICIP | 3 |
| 2024 | MVAFormer: RGB-Based Multi-View Spatio-Temporal Action Recognition with TransformerabstractMulti-view action recognition aims to recognize human actions using multiple camera views and deals with occlusion caused by obstacles or crowds. In this task, cooperation among views, which generates a joint representation by combining multiple views, is vital. Previous studies have explored promising cooperation methods for improving performance. However, since their methods focus only on the task setting of recognizing a single action from an entire video, they are not applicable to the recently popular spatio-temporal action recognition (STAR) setting, in which each person’s action is recognized sequentially. To address this problem, this paper proposes a multi-view action recognition method for the STAR setting, called MVAFormer. In MVAFormer, we introduce a novel transformer-based cooperation module among views. In contrast to previous studies, which utilize embedding vectors with lost spatial information, our module utilizes the feature map for effective cooperation in the STAR setting, which preserves the spatial information. Furthermore, in our module, we divide the self-attention for the same and different views to model the relationship between multiple views effectively. The results of experiments using a newly collected dataset demonstrate that MVAFormer outperforms the comparison baselines by approximately 4.4 points on the F-measure. Taiga Yamane, Ryo Masumura, Shotaro Tora |
ICIP | 3 |
| 2024 | Born-Again Multi-task Self-training for Multi-task Facial Emotion Recognition
Ryo Masumura, Akihiko Takashima, Shota Orihashi |
ICPR (16) | 1 |
| 2024 | Factor-Conditioned Speaking-Style Captioning
Atsushi Ando, Takafumi Moriya, Shota Horiguchi, Ryo Masumura |
INTERSPEECH | 4 |
| 2024 | SOMSRED: Sequential Output Modeling for Joint Multi-talker Overlapped Speech Recognition and Speaker Diarization
Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Atsushi Ando, Ryo Masumura |
INTERSPEECH | 7 |
| 2024 | Unified Multi-Talker ASR with and without Target-speaker Enrollment
Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Mana Ihori, Naotaka Kawata, Shota Orihashi, Kazutoshi Shinoda, Taiga Yamane, Saki Mizuno, Keita Suzuki, Nobukatsu Hojo, Takafumi Moriya, Atsushi Ando |
INTERSPEECH | 1 |
| 2024 | Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
Takafumi Moriya, Takanori Ashihara, Masato Mimura, Hiroshi Sato 0002, Kohei Matsuura, Ryo Masumura, Taichi Asami |
INTERSPEECH | 6 |
| 2024 | Learning from Multiple Annotator Biased Labels in Multimodal Conversation
Kazutoshi Shinoda, Nobukatsu Hojo, Saki Mizuno, Keita Suzuki, Satoshi Kobashikawa, Ryo Masumura |
INTERSPEECH | 6 |
| 2024 | Participant-Pair-Wise Bottleneck Transformer for Engagement Estimation from Video Conversation
Keita Suzuki, Nobukatsu Hojo, Kazutoshi Shinoda, Saki Mizuno, Ryo Masumura |
INTERSPEECH | 5 |
| 2023 | Leveraging Large Text Corpora For End-To-End Speech SummarizationabstractEnd-to-end speech summarization (E2E SSum) is a technique to directly generate summary sentences from speech. Compared with the cascade approach, which combines automatic speech recognition (ASR) and text summarization models, the E2E approach is more promising because it mitigates ASR errors, incorporates nonverbal information, and simplifies the overall system. However, since collecting a large amount of paired data (i.e., speech and summary) is difficult, the training data is usually insufficient to train a robust E2E SSum system. In this paper, we present two novel methods that leverage a large amount of external text summarization data for E2E SSum training. The first technique is to utilize a text-to-speech (TTS) system to generate synthesized speech, which is used for E2E SSum training with the text summary. The second is a TTS-free method that directly inputs phoneme sequence instead of synthesized speech to the E2E SSum model. Experiments show that our proposed TTS- and phoneme-based methods improve several metrics on the How2 dataset. In particular, our best system outperforms a previous state-of-the-art one by a large margin (i.e., METEOR score improvements of more than 6 points). To the best of our knowledge, this is the first work to use external language resources for E2E SSum. Moreover, we report a detailed analysis of the How2 dataset to confirm the validity of our proposed E2E SSum system. Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Tomohiro Tanaka, Atsunori Ogawa, Marc Delcroix, Ryo Masumura |
ICASSP | 7 |
| 2023 | Next-Speaker Prediction Based on Non-Verbal Information in Multi-Party Video ConversationabstractWe propose a method for next-speaker prediction, a task to predict who speaks in the next turn among multiple current listeners, in multi-party video conversation. Previous studies used non-verbal features, such as head movements and gaze behavior, for next-speaker prediction in face-to-face conversation. However, in video conversation, these non-verbal features are vague and ineffective because they look at the screen displaying other participants. Since non-verbal features include participant characteristics, it is necessary to use training data with rich combinations of participants to robustly predict the next speaker. Previous studies used training data with a limited number of combinations of participants because the data consist only of recorded data. Therefore, the proposed method uses 1) novel non-verbal features for next-speaker prediction in video conversation, specifically facial expressions, hand movements and speech segments, and 2) data augmentation of participant combinations in the training data. We conducted experiments to evaluate the proposed method, and the results using video-conversation data indicate its effectiveness. Saki Mizuno, Nobukatsu Hojo, Satoshi Kobashikawa, Ryo Masumura |
ICASSP | 4 |
| 2023 | Improving Scheduled Sampling for Neural Transducer-Based ASRabstractThe recurrent neural network-transducer (RNNT) is a promising approach for automatic speech recognition (ASR) with the introduction of a prediction network that autoregressively considers linguistic aspects. To train the autoregressive part, the ground-truth tokens are used as substitutions for the previous output token, which leads to insufficient robustness to incorrect past tokens; a recognition error in the decoding leads to further errors. Scheduled sampling (SS) is a technique to train autoregressive model robustly to past errors by randomly replacing some ground-truth tokens with actual outputs generated from a model. SS mitigates the gaps between training and decoding steps, known as exposure bias, and it is often used for attentional encoder-decoder training. However SS has not been fully examined for RNNT because of the difficulty in applying SS to RNNT due to the complicated RNNT output form. In this paper we propose SS approaches suited for RNNT. Our SS approaches sample the tokens generated from the distiribution of RNNT itself, i.e. internal language model or RNNT outputs. Experiments in three datasets confirm that RNNT trained with our SS approach achieves the best ASR performance. In particular, on a Japanese ASR task, our best system outperforms the previous state-of-the-art alternative. Takafumi Moriya, Takanori Ashihara, Hiroshi Sato 0002, Kohei Matsuura, Tomohiro Tanaka, Ryo Masumura |
ICASSP | 6 |
| 2023 | Leveraging Language Embeddings for Cross-Lingual Self-Supervised Speech Representation LearningabstractIn this paper, we propose novel cross-lingual self-supervised speech representation learning methods that explicitly consider language information. Cross-lingual self-supervised speech representation learning has been studied to make effective use of diverse data in various languages. Previous methods train models from multilingual datasets without taking language into account. However, it is difficult to train speech representations from multilingual datasets in the same space without language specification since there are clear differences in the acoustic context between languages. To solve this problem, we propose leveraging language IDs to build self-supervised speech representation learning models that explicitly consider language information. Our proposed models utilize fixed-dimensional language embeddings converted from language IDs for the model learning the relationship between related speech representations in different languages. We investigate two strategies to introduce language embeddings into the models: adding the embeddings to all of the inputs and concatenating to the inputs of the Transformer. We experimentally investigated how the difference between the two strategies affects the downstream tasks. Experimental results on the English and Japanese datasets show that the proposed methods improve the accuracies of downstream automatic speech recognition tasks. Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Hiroshi Sato 0002, Taiga Yamane, Takanori Ashihara, Kohei Matsuura, Takafumi Moriya |
ICASSP | 2 |
| 2023 | Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuracy-Robustness TradeoffabstractThis paper addresses the tradeoff between standard accuracy on clean examples and robustness against adversarial examples in deep neural networks (DNNs). Although adversarial training (AT) improves robustness, it degrades the standard accuracy, thus yielding the tradeoff. To mitigate this tradeoff, we propose a novel AT method called ARREST, which comprises three components: (i) adversarial finetuning (AFT), (ii) representation-guided knowledge distillation (RGKD), and (iii) noisy replay (NR). AFT trains a DNN on adversarial examples by initializing its parameters with a DNN that is standardly pretrained on clean examples. RGKD and NR respectively entail a regularization term and an algorithm to preserve latent representations of clean examples during AFT. RGKD penalizes the distance between the representations of the standardly pretrained and AFT DNNs. NR switches input adversarial examples to nonadversarial ones when the representation changes significantly during AFT. By combining these components, ARREST achieves both high standard accuracy and robustness. Experimental results demonstrate that ARREST mitigates the tradeoff more effectively than previous AT-based methods do. Shin'ya Yamaguchi, Shoichiro Takeda, Sekitoshi Kanai, Naoki Makishima, Atsushi Ando, Ryo Masumura |
ICCV | 7 |
| 2023 | Distilling Knowledge of Bidirectional Language Model for Scene Text RecognitionabstractThis paper proposes a knowledge distillation method for an external bidirectional language model trained by masked language modeling to achieve high accuracy in scene text recognition. In Asian languages such as Japanese, it is necessary to perform text recognition in units of multiple words or sentences rather than individual words because words are not separated by spaces, and so high-level linguistic knowledge is needed to recognize text correctly. To enhance linguistic knowledge, several methods that use an external language model have been proposed, but these methods fail to consider future context well in performing text recognition because they revise the text candidates yielded by autoregressive text recognition models, which consider mainly past context. To overcome this deficiency, our key idea is to enhance a text recognition model by utilizing knowledge of an external bidirectional language model trained by masked language modeling, which reflects not only past but also future context. So as to actively consider future context in text recognition, our proposed method introduces a distillation loss term that makes the output probability of the text recognition model closer to that of the bidirectional language model. Experiments on Japanese scene text recognition demonstrate the effectiveness of the proposed method. Shota Orihashi, Yoshihiro Yamazaki, Mihiro Uchida, Akihiko Takashima, Ryo Masumura |
ICIP | 5 |
| 2023 | OnDA-DETR: Online Domain Adaptation for Detection Transformers with Self-Training FrameworkabstractThis paper presents a novel method for online domain adaptation (OnDA) for DEtection TRansformer (DETR)-based object detection models called OnDA-DETR. OnDA is a domain adaptation paradigm that adapts a model trained on the source domain data to perform well on the target domain in an online manner during testing, using only the unlabeled test data from the target domain. Due to challenging and realistic problem settings, OnDA has garnered significant attention. However, OnDA methods for DETR-based models, which have demonstrated excellent performance in object detection research fields, had not been developed. OnDA-DETR is the first OnDA method specifically designed for DETR-based models. OnDA-DETR incorporates a self-training framework that generates pseudo-labels for the unlabeled target domain data. To effectively incorporate the self-training framework into DETR-based models, we leverage recall-aware pseudo-labeling and quality-aware training in OnDA-DETR. Experimental results indicate that OnDA-DETR improves the performance of the source-trained model by about 3.0 % points through OnDA. Taiga Yamane, Naoki Makishima, Keita Suzuki, Atsushi Ando, Ryo Masumura |
ICIP | 6 |
| 2023 | Open-Set Recognition for Facial-Expression RecognitionabstractWe address distinguishing whether an input is a facial image by learning only a facial-expression recognition (FER) dataset. To avoid misclassification in FER, it is necessary to distinguish whether the input is a facial image. Unfortunately, collecting exhaustive non-face images is costly. Therefore, distinguishing whether the input is a facial image by learning only an FER dataset is important. A representative method for this task is learning reconstruction of only facial images and determining high-error samples between input images and reconstructed images as non-face images. However, reconstruction is difficult on facial images because such images contain detailed features. Our key idea to tackle the task without reconstruction is assuming that facial images will match several emotions, whereas non-face images will not match any emotion. Therefore, we propose a method for training a discriminator that determines whether the inputs and emotions match using counter-factual pairs in an FER dataset. A metric for the task is then obtained by taking into account each emotion in the posterior probability that inputs and emotions match, estimated by the discriminator. Experiments on the RAF-DB dataset vs. the Stanford Dogs dataset and AffectNet datasets showed the effectiveness of our method. Mihiro Uchida, Shota Orihashi, Akihiko Takashima, Yoshihiro Yamazaki, Ryo Masumura |
ICIP | 5 |
| 2023 | Audio-Visual Praise Estimation for Conversational Video based on Synchronization-Guided Multimodal Transformer
Nobukatsu Hojo, Saki Mizuno, Satoshi Kobashikawa, Ryo Masumura, Mana Ihori, Tomohiro Tanaka |
INTERSPEECH | 4 |
| 2023 | Transcribing Speech as Spoken and Written Dual Text Using an Autoregressive Model
Mana Ihori, Tomohiro Tanaka, Ryo Masumura, Saki Mizuno, Nobukatsu Hojo |
INTERSPEECH | 4 |
| 2023 | What are differences? Comparing DNN and Human by Their Performance and Characteristics in Speaker Age Estimation
Yuki Kitagishi, Naohiro Tawara, Atsunori Ogawa, Ryo Masumura, Taichi Asami |
INTERSPEECH | 4 |
| 2023 | Joint Autoregressive Modeling of End-to-End Multi-Talker Overlapped Speech Recognition and Utterance-level Timestamp Prediction
Naoki Makishima, Keita Suzuki, Atsushi Ando, Ryo Masumura |
INTERSPEECH | 5 |
| 2023 | End-to-End Joint Target and Non-Target Speakers ASR
Ryo Masumura, Naoki Makishima, Taiga Yamane, Yoshihiko Yamazaki, Saki Mizuno, Mana Ihori, Mihiro Uchida, Keita Suzuki, Hiroshi Sato 0002, Tomohiro Tanaka, Akihiko Takashima, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando |
INTERSPEECH | 1 |
| 2023 | Knowledge Distillation for Neural Transducer-based Target-Speaker ASR: Exploiting Parallel Mixture/Single-Talker Speech Data
Takafumi Moriya, Hiroshi Sato 0002, Tsubasa Ochiai, Marc Delcroix, Takanori Ashihara, Kohei Matsuura, Tomohiro Tanaka, Ryo Masumura, Atsunori Ogawa, Taichi Asami |
INTERSPEECH | 8 |
| 2023 | Downstream Task Agnostic Speech Enhancement with Self-Supervised Representation Loss
Hiroshi Sato 0002, Ryo Masumura, Tsubasa Ochiai, Marc Delcroix, Takafumi Moriya, Takanori Ashihara, Kentaro Shinayama, Saki Mizuno, Mana Ihori, Tomohiro Tanaka, Nobukatsu Hojo |
INTERSPEECH | 2 |
| 2023 | Multi-region CNN-Transformer for Micro-gesture Recognition in Face and Upper BodyabstractThis paper presents a novel task that recognizes from a video unintentional micro-gestures (UMGs), which are movements made by people unconsciously and unintentionally. Recognizing UMGs is crucial because they reveal a person’s underlying psychological state. Since a UMG is composed of subtle sequential movements, the recognition model must be able to capture accurate information in both the spatial and temporal directions. Therefore, we utilize a convolutional neural network (CNN) to capture information in the spatial direction and a Transformer to merge the features extracted by the CNN in the temporal direction. However, this model often misrecognizes UMGs because it is not possible to capture slight differences in movements, such as in the face and mouth regions. To address this issue, we propose a novel model for UMG recognition, the Multi-Region CNN-Transformer model, that inputs cropped videos from multiple upper body regions simultaneously. The key advance of our method is to capture subtle changes in regions such as the upper body, face, head, and mouth for recognizing UMGs. We demonstrate the effectiveness of the proposed method through experiments using our newly created UMG dataset for this task. Keita Suzuki, Ryo Masumura, Atsushi Ando, Naoki Makishima |
MMAsia | 3 |
| 2022 | Multi-Perspective Document RevisionabstractThis paper presents a novel multi-perspective document revision task. In conventional studies on document revision, tasks such as grammatical error correction, sentence reordering, and discourse relation classification have been performed individually; however, these tasks simultaneously should be revised to improve the readability and clarity of a whole document. Thus, our study defines multi-perspective document revision as a task that simultaneously revises multiple perspectives. To model the task, we design a novel Japanese multi-perspective document revision dataset that simultaneously handles seven perspectives to improve the readability and clarity of a document. Although a large amount of data that simultaneously handles multiple perspectives is needed to model multi-perspective document revision elaborately, it is difficult to prepare such a large amount of this data. Therefore, our study offers a multi-perspective document revision modeling method that can use a limited amount of matched data (i.e., data for the multi-perspective document revision task) and external partially-matched data (e.g., data for the grammatical error correction task). Experiments using our created dataset demonstrate the effectiveness of using multiple partially-matched datasets to model the multi-perspective document revision task. Mana Ihori, Tomohiro Tanaka, Ryo Masumura |
COLING | 4 |
| 2022 | Customer Satisfaction Estimation Using Unsupervised Representation Learning with Multi-Format Prediction LossabstractWe propose a new Customer Satisfaction Estimation (CSE) method that utilizes unsupervised representation learning. Though conventional methods have improved both the heuristic features and the estimation models, their performance is still insufficient as only small amounts of labeled training data can be expected. To mitigate this problem, the proposed method leverages a large amount of unlabeled data by unsupervised representation learning based on self-training. The key advance of the proposed method is to introduce a Multi-Format Prediction (MFP) loss to improve the performance of self-training for the inputs that contain both continuous and biased discrete features such as the number of occurrences of a particular word. MFP loss uses two loss functions based on regression and weighted binary classification to reconstruct both types of features with high accuracy. Experiments on real English contact center calls reveal the improved CSE performance attained by the proposed method. Atsushi Ando, Yumiko Murata, Ryo Masumura, Naoki Makishima, Takafumi Moriya, Takanori Ashihara, Hiroshi Sato 0002 |
ICASSP | 3 |
| 2022 | Hybrid RNN-T/Attention-Based Streaming ASR with Triggered Chunkwise Attention and Dual Internal Language Model IntegrationabstractIn this paper we propose improvements to our recently proposed hybrid RNN-T/Attention architecture that includes a shared encoder followed by recurrent neural network-transducer (RNN-T) and triggered attention-based decoders (TAD). The use of triggered attention enables the attention-based decoder (AD) to operate in a streaming manner. When a trigger point is detected by RNN-T, TAD uses the context from the start-of-speech up to that trigger point to compute the attention weights. Consequently, the computation costs and the memory consumptions are quadratically increased with the duration of the utterances because all input features must be stored and used to re-compute the attention weights. In this paper, we use a short context from a few frames prior to each trigger point for attention weight computation resulting in reduced computation and memory costs. We call the proposed framework triggered chunkwise AD (TCAD). We also investigate the effectiveness of internal language model (ILM) estimation approach using both ILMs of RNN-T and TCAD heads for improving RNN-T performance. We confirm in experiments with public and private datasets covering various scenarios that TCAD achieves superior recognition performance while reducing computation costs compared to TAD. Takafumi Moriya, Takanori Ashihara, Atsushi Ando, Hiroshi Sato 0002, Tomohiro Tanaka, Kohei Matsuura, Ryo Masumura, Marc Delcroix, Takahiro Shinozaki |
ICASSP | 7 |
| 2022 | Fully Shareable Scene Text Recognition Modeling for Horizontal and Vertical WritingabstractThis paper proposes a simple and efficient method of joint scene text recognition for both horizontal and vertical writing. Recently, end-to-end scene text recognition using the Transformer-based autoregressive encoder-decoder model offers high recognition accuracy. Research into this method has mainly focused on horizontally written text, but in several Asian countries, texts are also written vertically. To efficiently train a recognition model for jointly recognizing horizontal and vertical writing, several methods have been proposed that partially share model components between each writing direction. However, this approach lowers training efficiency because non-shareable components are trained only on just horizontal or vertical writing data. To increase training efficiency, our key idea is to consider writing direction in the continuous space obtained by a fully shareable model for horizontal and vertical writing. To this end, our proposed method gives the writing direction as an initial token to the autoregressive decoder while sharing all components for each writing direction. Furthermore, to incorporate common features between each writing into the model, the proposed method predicts character count before predicting the character string. Experiments on Japanese scene text recognition demonstrate the effectiveness of the proposed method. Shota Orihashi, Yoshihiro Yamazaki, Mihiro Uchida, Akihiko Takashima, Ryo Masumura |
ICIP | 5 |
| 2022 | Speaker consistency loss and step-wise optimization for semi-supervised joint training of TTS and ASR using unpaired text dataabstractIn this paper, we investigate the semi-supervised joint training of text to speech (TTS) and automatic speech recognition (ASR), where a small amount of paired data and a large amount of unpaired text data are available.Conventional studies form a cycle called the TTS-ASR pipeline, where the multispeaker TTS model synthesizes speech from text with a reference speech and the ASR model reconstructs the text from the synthesized speech, after which both models are trained with a cycle-consistency loss.However, the synthesized speech does not reflect the speaker characteristics of the reference speech and the synthesized speech becomes overly easy for the ASR model to recognize after training.This not only decreases the TTS model quality but also limits the ASR model improvement.To solve this problem, we propose improving the cycleconsistency-based training with a speaker consistency loss and step-wise optimization.The speaker consistency loss brings the speaker characteristics of the synthesized speech closer to that of the reference speech.In the step-wise optimization, we first freeze the parameter of the TTS model before both models are trained to avoid over-adaptation of the TTS model to the ASR model.Experimental results demonstrate the efficacy of the proposed method. Naoki Makishima, Atsushi Ando, Ryo Masumura |
INTERSPEECH | 4 |
| 2022 | End-to-End Joint Modeling of Conversation History-Dependent and Independent ASR Systems with Multi-History Training
Ryo Masumura, Yoshihiro Yamazaki, Saki Mizuno, Naoki Makishima, Mana Ihori, Mihiro Uchida, Hiroshi Sato 0002, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando |
INTERSPEECH | 1 |
| 2022 | Predicting VQVAE-based Character Acting Style from Quotation-Annotated Text for Audiobook Speech Synthesis
Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Yuki Saito 0001, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
INTERSPEECH | 6 |
| 2022 | Dialogue Acts Aided Important Utterance Detection Based on Multiparty and Multimodal Information
Fumio Nihei, Ryo Ishii, Yukiko I. Nakano, Kyosuke Nishida, Ryo Masumura, Atsushi Fukayama, Takao Nakamura |
INTERSPEECH | 5 |
| 2022 | Strategies to Improve Robustness of Target Speech Extraction to Enrollment VariationsabstractTarget speech extraction is a technique to extract the target speaker's voice from mixture signals using a pre-recorded enrollment utterance that characterize the voice characteristics of the target speaker.One major difficulty of target speech extraction lies in handling variability in "intra-speaker" characteristics, i.e., characteristics mismatch between target speech and an enrollment utterance.While most conventional approaches focus on improving average performance given a set of enrollment utterances, here we propose to guarantee the worst performance, which we believe is of great practical importance.In this work, we propose an evaluation metric called worstenrollment source-to-distortion ratio (SDR) to quantitatively measure the robustness towards enrollment variations.We also introduce a novel training scheme that aims at directly optimizing the worst-case performance by focusing on training with difficult enrollment cases where extraction does not perform well.In addition, we investigate the effectiveness of auxiliary speaker identification loss (SI-loss) as another way to improve robustness over enrollments.Experimental validation reveals the effectiveness of both worst-enrollment target training and SI-loss training to improve robustness against enrollment variations, by increasing speaker discriminability. Hiroshi Sato 0002, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Takafumi Moriya, Naoki Makishima, Mana Ihori, Tomohiro Tanaka, Ryo Masumura |
INTERSPEECH | 9 |
| 2022 | Interactive Co-Learning with Cross-Modal Transformer for Audio-Visual Emotion Recognition
Akihiko Takashima, Ryo Masumura, Atsushi Ando, Yoshihiro Yamazaki, Mihiro Uchida, Shota Orihashi |
INTERSPEECH | 2 |
| 2022 | Domain Adversarial Self-Supervised Speech Representation Learning for Improving Unknown Domain Downstream Tasks
Tomohiro Tanaka, Ryo Masumura, Hiroshi Sato 0002, Mana Ihori, Kohei Matsuura, Takanori Ashihara, Takafumi Moriya |
INTERSPEECH | 2 |
| 2022 | Multimodal Negotiation Corpus with Various Subjective Assessments for Social-Psychological Outcome Prediction from Non-Verbal CuesabstractThis study investigates social-psychological negotiation-outcome prediction (SPNOP), a novel task for estimating various subjective evaluation scores of negotiation, such as satisfaction and trust, from negotiation dialogue data. To investigate SPNOP, a corpus with various psychological measurements is beneficial because the interaction process of negotiation relates to many aspects of psychology. However, current negotiation corpora only include information related to objective outcomes or a single aspect of psychology. In addition, most use the “laboratory setting” that uses non-skilled negotiators and over simplified negotiation scenarios. There is a concern that such a gap with actual negotiation will intrinsically affect the behavior and psychology of negotiators in the corpus, which can degrade the performance of models trained from the corpus in real situations. Therefore, we created a negotiation corpus with three features; 1) was assessed with various psychological measurements, 2) used skilled negotiators, and 3) used scenarios of context-rich negotiation. We recorded video and audio of negotiations in Japanese to investigate SPNOP in the context of social signal processing. Experimental results indicate that social-psychological outcomes can be effectively estimated from multimodal information. Nobukatsu Hojo, Satoshi Kobashikawa, Saki Mizuno, Ryo Masumura |
LREC | 4 |
| 2022 | On the Use of Modality-Specific Large-Scale Pre-Trained Encoders for Multimodal Sentiment AnalysisabstractThis paper investigates the effectiveness and implementation of modality-specific large-scale pre-trained encoders for multimodal sentiment analysis (MSA). Although the effectiveness of pre-trained encoders in various fields has been reported, conventional MSA methods employ them for only linguistic modality, and their application has not been investigated. This paper compares the features yielded by large-scale pre-trained encoders with conventional heuristic features. One each of the largest pre-trained encoders publicly available for each modality are used; CLIP-ViT, WavLM, and BERT for visual, acoustic, and linguistic modalities, respectively. Experiments on two datasets reveal that methods with domain-specific pre-trained encoders attain better performance than those with conventional features in both unimodal and multimodal scenarios. We also find it better to use the outputs of the intermediate layers of the encoders than those of the output layer. The codes are available at https://github.com/ando-hub/MSA_Pretrain. Atsushi Ando, Ryo Masumura, Akihiko Takashima, Naoki Makishima, Keita Suzuki, Takafumi Moriya, Takanori Ashihara, Hiroshi Sato 0002 |
SLT | 2 |
| 2021 | Hierarchical Knowledge Distillation for Dialogue Sequence LabelingabstractThis paper presents a novel knowledge distillation method for dialogue sequence labeling. Dialogue sequence labeling is a supervised learning task that estimates labels for each utterance in the target dialogue document, and is useful for many applications such as dialogue act estimation. Accurate labeling is often realized by a hierarchically-structured large model consisting of utterance-level and dialogue-level networks that capture the contexts within an utterance and between utterances, respectively. However, due to its large model size, such a model cannot be deployed on resource-constrained devices. To overcome this difficulty, we focus on knowledge distillation which trains a small model by distilling the knowledge of a large and high performance teacher model. Our key idea is to distill the knowledge while keeping the complex contexts captured by the teacher model. To this end, the proposed method, hierarchical knowledge distillation, trains the small model by distilling not only the probability distribution of the label classification, but also the knowledge of utterance-level and dialogue-level contexts trained in the teacher model by training the model to mimic the teacher model's output in each level. Experiments on dialogue act estimation and call scene segmentation demonstrate the effectiveness of the proposed method. Shota Orihashi, Yoshihiro Yamazaki, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Ryo Masumura |
ASRU | 7 |
| 2021 | Speech Emotion Recognition Based on Listener Adaptive ModelsabstractThis paper presents a novel speech emotion recognition scheme that can deal with the individuality of emotion perception. Most conventional methods directly model the majority decision of multiple listener’s perceived emotions. However, emotion perception varies with the listener, which means the conventional methods can mismatch the recognition results to human perception. In order to mitigate this problem, we propose a Listener Adaptive (LA) model that reflects emotion recognition criteria of each listener. One-hot listener codes with several adaptation layers are employed in the LA model. The LA model yields the posterior probabilities of the listener-specific perceived emotions. Majority-voted emotion can be also estimated by averaging, in the LA model, the posterior probabilities for all listeners. Experiments on two emotional speech datasets demonstrate that the proposed approach offers improved listener-wise perceived emotion recognition performance in natural speech. Atsushi Ando, Ryo Masumura, Hiroshi Sato 0002, Takafumi Moriya, Takanori Ashihara, Yusuke Ijima, Tomoki Toda |
ICASSP | 2 |
| 2021 | MAPGN: Masked Pointer-Generator Network for Sequence-to-Sequence Pre-TrainingabstractThis paper presents a self-supervised learning method for pointer-generator networks to improve spoken-text normalization. Spoken-text normalization that converts spoken-style text into style normalized text is becoming an important technology for improving subsequent processing such as machine translation and summarization. The most successful spoken-text normalization method to date is sequence-to-sequence (seq2seq) mapping using pointer-generator networks that possess a copy mechanism from an input sequence. However, these models require a large amount of paired data of spoken-style text and style normalized text, and it is difficult to prepare such a volume of data. In order to construct spoken-text normalization model from the limited paired data, we focus on self-supervised learning which can utilize unpaired text data to improve seq2seq models. Unfortunately, conventional self-supervised learning methods do not assume that pointer-generator networks are utilized. Therefore, we propose a novel self-supervised learning method, MAsked Pointer-Generator Network (MAPGN). The proposed method can effectively pre-train the pointer-generator net-work by learning to fill masked tokens using the copy mechanism. Our experiments demonstrate that MAPGN is more effective for pointer-generator networks than the conventional self-supervised learning methods in two spoken-text normalization tasks. Mana Ihori, Naoki Makishima, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Ryo Masumura |
ICASSP | 6 |
| 2021 | Audio-Visual Speech Separation Using Cross-Modal Correspondence LossabstractWe present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual speech separation is a technique to estimate the individual speech signals from a mixture using the visual signals of the speakers. Conventional studies on audio-visual speech separation mainly train the separation model on the audio-only loss, which reflects the distance between the source signals and the separated signals. However, conventional losses do not reflect the characteristics of the speech signals, including the speaker’s characteristics and phonetic information, which leads to distortion or remaining noise. To address this problem, we propose the cross-modal correspondence (CMC) loss, which is based on the cooccurrence of the speech signal and the visual signal. Since the visual signal is not affected by background noise and contains speaker and phonetic information, using the CMC loss enables the audio-visual speech separation model to remove noise while preserving the speech characteristics. Experimental results demonstrate that the proposed method learns the cooccurrence on the basis of CMC loss, which improves separation performance. Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
ICASSP | 6 |
| 2021 | Hierarchical Transformer-Based Large-Context End-To-End ASR with Large-Context Knowledge DistillationabstractWe present a novel large-context end-to-end automatic speech recognition (E2E-ASR) model and its effective training method based on knowledge distillation. Common E2E-ASR models have mainly focused on utterance-level processing in which each utterance is independently transcribed. On the other hand, large-context E2E-ASR models, which take into account long-range sequential contexts beyond utterance boundaries, well handle a sequence of utterances such as discourses and conversations. However, the transformer architecture, which has recently achieved state-of-the-art ASR performance among utterance-level ASR systems, has not yet been introduced into the large-context ASR systems. We can expect that the transformer architecture can be leveraged for effectively capturing not only input speech contexts but also long-range sequential contexts beyond utterance boundaries. Therefore, this paper proposes a hierarchical transformer-based large-context E2E-ASR model that combines the transformer architecture with hierarchical encoder-decoder based large-context modeling. In addition, in order to enable the proposed model to use long-range sequential contexts, we also propose a large-context knowledge distillation that distills the knowledge from a pre-trained large-context language model in the training phase. We evaluate the effectiveness of the proposed model and proposed training method on Japanese discourse ASR tasks. Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
ICASSP | 1 |
| 2021 | Simpleflat: A Simple Whole-Network Pre-Training Approach for RNN Transducer-Based End-to-End Speech RecognitionabstractRecurrent neural network-transducer (RNN-T) is promising for building time-synchronous end-to-end automatic speech recognition (ASR) systems, in part because it does not need frame-wise alignment between input features and target labels in the training step. Although training without alignment is beneficial, it makes it difficult to discern the relation between input features and output token sequences. This, in effect, degrades RNN-T performance. Our solution is SimpleFlat (SF), a novel and simple whole-network pretraining approach for RNN-T. SF extracts frame-wise alignments on-the-fly from the training dataset, and does not require any external resources. We distribute equal numbers of target tokens to each frame following RNN-T encoder output lengths by repeating each token. The frame-wise tokens so created are shifted, and also used as the prediction network inputs. Therefore, SF can be implemented by cross entropy loss computation as in autoregressive model training. Experiments on Japanese and English ASR tasks demonstrate that SF can effectively improve various RNN-T architectures. Takafumi Moriya, Takanori Ashihara, Tomohiro Tanaka, Tsubasa Ochiai, Hiroshi Sato 0002, Atsushi Ando, Yusuke Ijima, Ryo Masumura, Yusuke Shinohara |
ICASSP | 8 |
| 2021 | Zero-Shot Joint Modeling of Multiple Spoken-Text-Style Conversion Tasks Using Switching TokensabstractIn this paper, we propose a novel spoken-text-style conversion method that can simultaneously execute multiple style conversion modules such as punctuation restoration and disfluency deletion without preparing matched datasets.In practice, transcriptions generated by automatic speech recognition systems are not highly readable because they often include many disfluencies and do not include punctuation marks.To improve their readability, multiple spoken-text-style conversion modules that individually model a single conversion task are cascaded because matched datasets that simultaneously handle multiple conversion tasks are often unavailable.However, the cascading is unstable against the order of tasks because of the chain of conversion errors.Besides, the computation cost of the cascading must be higher than the single conversion.To execute multiple conversion tasks simultaneously without preparing matched datasets, our key idea is to distinguish individual conversion tasks using the on-off switch.In our proposed zero-shot joint modeling, we switch the individual tasks using multiple switching tokens, enabling us to utilize a zero-shot learning approach to executing simultaneous conversions.Our experiments on joint modeling of disfluency deletion and punctuation restoration demonstrate the effectiveness of our method. Mana Ihori, Naoki Makishima, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Ryo Masumura |
Interspeech | 6 |
| 2021 | Enrollment-Less Training for Personalized Voice Activity DetectionabstractWe present a novel personalized voice activity detection (PVAD) learning method that does not require enrollment data during training.PVAD is a task to detect the speech segments of a specific target speaker at the frame level using enrollment speech of the target speaker.Since PVAD must learn speakers' speech variations to clarify the boundary between speakers, studies on PVAD used large-scale datasets that contain many utterances for each speaker.However, the datasets to train a PVAD model are often limited because substantial cost is needed to prepare such a dataset.In addition, we cannot utilize the datasets used to train the standard VAD because they often lack speaker labels.To solve these problems, our key idea is to use one utterance as both a kind of enrollment speech and an input to the PVAD during training, which enables PVAD training without enrollment speech.In our proposed method, called enrollment-less training, we augment one utterance so as to create variability between the input and the enrollment speech while keeping the speaker identity, which avoids the mismatch between training and inference.Our experimental results demonstrate the efficacy of the method. Naoki Makishima, Mana Ihori, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Ryo Masumura |
Interspeech | 6 |
| 2021 | Unified Autoregressive Modeling for Joint End-to-End Multi-Talker Overlapped Speech Recognition and Speaker Attribute EstimationabstractIn this paper, we present a novel modeling method for singlechannel multi-talker overlapped automatic speech recognition (ASR) systems.Fully neural network based end-to-end models have dramatically improved the performance of multi-taker overlapped ASR tasks.One promising approach for end-toend modeling is autoregressive modeling with serialized output training in which transcriptions of multiple speakers are recursively generated one after another.This enables us to naturally capture relationships between speakers.However, the conventional modeling method cannot explicitly take into account the speaker attributes of individual utterances such as gender and age information.In fact, the performance deteriorates when each speaker is the same gender or is close in age.To address this problem, we propose unified autoregressive modeling for joint end-to-end multi-talker overlapped ASR and speaker attribute estimation.Our key idea is to handle gender and age estimation tasks within the unified autoregressive modeling.In the proposed method, transformer-based autoregressive model recursively generates not only textual tokens but also attribute tokens of each speaker.This enables us to effectively utilize speaker attributes for improving multi-talker overlapped ASR.Experiments on Japanese multi-talker overlapped ASR tasks demonstrate the effectiveness of the proposed method. Ryo Masumura, Daiki Okamura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
Interspeech | 1 |
| 2021 | Streaming End-to-End Speech Recognition for Hybrid RNN-T/Attention Architecture
Takafumi Moriya, Tomohiro Tanaka, Takanori Ashihara, Tsubasa Ochiai, Hiroshi Sato 0002, Atsushi Ando, Ryo Masumura, Marc Delcroix, Taichi Asami |
Interspeech | 7 |
| 2021 | Cross-Modal Transformer-Based Neural Correction Models for Automatic Speech RecognitionabstractWe propose a cross-modal transformer-based neural correction models that refines the output of an automatic speech recognition (ASR) system so as to exclude ASR errors.Generally, neural correction models are composed of encoder-decoder networks, which can directly model sequence-to-sequence mapping problems.The most successful method is to use both input speech and its ASR output text as the input contexts for the encoder-decoder networks.However, the conventional method cannot take into account the relationships between these two different modal inputs because the input contexts are separately encoded for each modal.To effectively leverage the correlated information between the two different modal inputs, our proposed models encode two different contexts jointly on the basis of cross-modal self-attention using a transformer.We expect that cross-modal self-attention can effectively capture the relationships between two different modals for refining ASR hypotheses.We also introduce a shallow fusion technique to efficiently integrate the first-pass ASR model and our proposed neural correction model.Experiments on Japanese natural language ASR tasks demonstrated that our proposed models achieve better ASR performance than conventional neural correction models. Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Akihiko Takashima, Takafumi Moriya, Takanori Ashihara, Shota Orihashi, Naoki Makishima |
Interspeech | 2 |
| 2021 | End-to-End Rich Transcription-Style Automatic Speech Recognition with Semi-Supervised LearningabstractWe propose a semi-supervised learning method for building end-to-end rich transcription-style automatic speech recognition (RT-ASR) systems from small-scale rich transcriptionstyle and large-scale common transcription-style datasets.In spontaneous speech tasks, various speech phenomena such as fillers, word fragments, laughter and coughs, etc. are often included.While common transcriptions do not give special awareness to these phenomena, rich transcriptions explicitly convert them into special phenomenon tokens as well as textual tokens.In previous studies, the textual and phenomenon tokens were simultaneously estimated in an end-to-end manner.However, it is difficult to build accurate RT-ASR systems because large-scale rich transcription-style datasets are often unavailable.To solve this problem, our training method uses a limited rich transcription-style dataset and common transcriptionstyle dataset simultaneously.The Key process in our semisupervised learning is to convert the common transcription-style dataset into a pseudo-rich transcription-style dataset.To this end, we introduce style tokens which control phenomenon tokens are generated or not into transformer-based autoregressive modeling.We use this modeling for generating the pseudorich transcription-style datasets and for building RT-ASR system from the pseudo and original datasets.Our experiments on spontaneous ASR tasks showed the effectiveness of the proposed method. Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Akihiko Takashima, Shota Orihashi, Naoki Makishima |
Interspeech | 2 |
| 2021 | Utilizing Resource-Rich Language Datasets for End-to-End Scene Text Recognition in Resource-Poor LanguagesabstractThis paper presents a novel training method for end-to-end scene text recognition. End-to-end scene text recognition offers high recognition accuracy, especially when using the encoder-decoder model based on Transformer. To train a highly accurate end-to-end model, we need to prepare a large image-to-text paired dataset for the target language. However, it is difficult to collect this data, especially for resource-poor languages. To overcome this difficulty, our proposed method utilizes well-prepared large datasets in resource-rich languages such as English, to train the resource-poor encoder-decoder model. Our key idea is to build a model in which the encoder reflects knowledge of multiple languages while the decoder specializes in knowledge of just the resource-poor language. To this end, the proposed method pre-trains the encoder by using a multilingual dataset that combines the resource-poor language’s dataset and the resource-rich language’s dataset to learn language-invariant knowledge for scene text recognition. The proposed method also pre-trains the decoder by using the resource-poor language’s dataset to make the decoder better suited to the resource-poor language. Experiments on Japanese scene text recognition using a small, publicly available dataset demonstrate the effectiveness of the proposed method. Shota Orihashi, Yoshihiro Yamazaki, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Ryo Masumura |
MMAsia | 7 |
| 2021 | Large-Context Conversational Representation Learning: Self-Supervised Learning For Conversational DocumentsabstractThis paper presents a novel self-supervised learning method for handling conversational documents consisting of transcribed text of human-to-human conversations. One of the key technologies for understanding conversational documents is utterance-level sequential labeling, where labels are estimated from the documents in an utterance-by-utterance manner. The main issue with utterance-level sequential labeling is the difficulty of collecting labeled conversational documents, as manual annotations are very costly. To deal with this issue, we propose large-context conversational representation learning (LC-CRL), a self-supervised learning method specialized for conversational documents. A self-supervised learning task in LC-CRL involves the estimation of an utterance using all the surrounding utterances based on large-context language modeling. In this way, LC-CRL enables us to effectively utilize unlabeled conversational documents and thereby enhances the utterance-level sequential labeling. The results of experiments on scene segmentation tasks using contact center conversational datasets demonstrate the effectiveness of the proposed method. Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
SLT | 1 |
| 2021 | Neural candidate-aware language models for speech recognition
Tomohiro Tanaka, Ryo Masumura, Takanobu Oba |
Comput. Speech Lang. | 2 |
| 2020 | Large-Context Pointer-Generator Networks for Spoken-to-Written Style ConversionabstractThis paper introduces a spoken-to-written style conversion method that is suitable for handling a series of text such as discourses and conversations. Spoken-to-written style conversion can increase the readability of automatic speech recognition (ASR) outputs because ASR systems transcribe input speech into text in a literal manner; however, it generates several disfluencies and redundant expressions. The most successful method of text style conversion is sequence-to-sequence mapping using pointer-generator networks that possess a copy mechanism from an input sequence. However, pointer-generator networks cannot process a series of text serially because they are developed to handle isolated text. In fact, pointer-generator networks cannot consider relationships between current processing text and all preceding text. Therefore, this paper proposes large-context pointer-generator networks that combine pointer-generator networks with large-context encoder-decoder networks. In the proposed networks, all preceding written-style text can be considered to convert current spoken-style text into written-style text. In addition, the proposed networks introduce a large-context copy mechanism that can copy tokens from both current spoken-style text and preceding written-style text. Our experiments demonstrate the proposed networks yield better performance than conventional pointer-generator networks and large-context encoder-decoder networks. Mana Ihori, Akihiko Takashima, Ryo Masumura |
ICASSP | 3 |
| 2020 | Sequence-Level Consistency Training for Semi-Supervised End-to-End Automatic Speech RecognitionabstractThis paper presents a novel semi-supervised end-to-end automatic speech recognition (ASR) method that employs consistency training with the use of unlabeled data. In consistency training, unlabeled data can be utilized for constraining a model such that it becomes invariant to small deformation. In fact, considering consistency can make the model robust to a variety of input examples. While previous studies have applied consistency training to primitive classification problems, no studies have employed consistency training to tackle sequence-to-sequence generation problems including end-to- end ASR. One problem is that existing consistency training schemes cannot take sequence-level generation consistency into consideration. In this paper, we propose a sequence-level consistency training scheme specialized to handle sequence-to-sequence generation problems. Our key idea is to consider the consistency of the generation function by utilizing beam search decoding results. For semi- supervised learning, we adopt Transformer as the end-to-end ASR model, and SpecAugment as the deformation function in consistency training. Our experiments show that our semi-supervised learning proposal with sequence-level consistency training can efficiently improve ASR performance using unlabeled speech data. Ryo Masumura, Mana Ihori, Akihiko Takashima, Takafumi Moriya, Atsushi Ando, Yusuke Shinohara |
ICASSP | 1 |
| 2020 | Distilling Attention Weights for CTC-Based ASR SystemsabstractWe present a novel training approach for connectionist temporal classification (CTC) -based automatic speech recognition (ASR) systems. CTC models are promising for building both a conventional acoustic model and an end-to-end (E2E) ASR model. However, CTC models make it difficult to capture the correct timing of each output label because timing is not given explicitly in the training data. In this paper, we propose a new auxiliary task with frame-wise targets for CTC model enhancement. We utilize attention weights generated by an attention-based encoder-decoder model (S2S) for making the targets, called the attention matrix. The attention matrix is the sum of the products of the attention weights (spike timing information) and the corresponding target vectors (probability information), and used for S2S-to-CTC knowledge distillation loss computation. Therefore, the attention matrix makes the CTC models jointly train-able as regards spike timings and their posteriors. Experiments on Japanese ASR tasks demonstrate that our proposal is effective for CTC model training; it achieves a 10.2% (E2E) / 9.4% (acoustic model) relative reduction in the character/kana-syllable error rates compared to models trained using only CTC loss. Takafumi Moriya, Hiroshi Sato 0002, Tomohiro Tanaka, Takanori Ashihara, Ryo Masumura, Yusuke Shinohara |
ICASSP | 5 |
| 2020 | Memory Attentive Fusion: External Language Model Integration for Transformer-based Sequence-to-Sequence ModelabstractThis paper presents a novel fusion method for integrating an external language model (LM) into the Transformer based sequenceto-sequence (seq2seq) model.While paired data are basically required to train the seq2seq model, the external LM can be trained with only unpaired data.Thus, it is important to leverage memorized knowledge in the external LM for building the seq2seq model, since it is hard to prepare a large amount of paired data.However, the existing fusion methods assume that the LM is integrated with recurrent neural network-based seq2seq models instead of the Transformer.Therefore, this paper proposes a fusion method that can explicitly utilize network structures in the Transformer.The proposed method, called memory attentive fusion, leverages the Transformer-style attention mechanism that repeats source-target attention in a multi-hop manner for reading the memorized knowledge in the LM.Our experiments on two text-style conversion tasks demonstrate that the proposed method performs better than conventional fusion methods. Mana Ihori, Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi |
INLG | 2 |
| 2020 | A Transformer-Based Audio Captioning Model with Keyword EstimationabstractOne of the problems with automated audio captioning (AAC) is the indeterminacy in word selection corresponding to the audio event/scene.Since one acoustic event/scene can be described with several words, it results in a combinatorial explosion of possible captions and difficulty in training.To solve this problem, we propose a Transformer-based audio-captioning model with keyword estimation called TRACKE.It simultaneously solves the word-selection indeterminacy problem with the main task of AAC while executing the sub-task of acoustic event detection/acoustic scene classification (i.e., keyword estimation).TRACKE estimates keywords, which comprise a word set corresponding to audio events/scenes in the input audio, and generates the caption while referring to the estimated keywords to reduce word-selection indeterminacy.Experimental results on a public AAC dataset indicate that TRACKE achieved state-ofthe-art performance and successfully estimated both the caption and its keywords. Yuma Koizumi, Ryo Masumura, Kyosuke Nishida, Masahiro Yasuda, Shoichiro Saito |
INTERSPEECH | 2 |
| 2020 | Phoneme-to-Grapheme Conversion Based Large-Scale Pre-Training for End-to-End Automatic Speech Recognition
Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
INTERSPEECH | 1 |
| 2020 | Self-Distillation for Improving CTC-Transformer-Based ASR Systems
Takafumi Moriya, Tsubasa Ochiai, Shigeki Karita, Hiroshi Sato 0002, Tomohiro Tanaka, Takanori Ashihara, Ryo Masumura, Yusuke Shinohara, Marc Delcroix |
INTERSPEECH | 7 |
| 2020 | Unsupervised Domain Adaptation for Dialogue Sequence Labeling Based on Hierarchical Adversarial Training
Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Ryo Masumura |
INTERSPEECH | 4 |
| 2020 | Investigating Effective Additional Contextual Factors in DNN-Based Spontaneous Speech Synthesis
Yuki Yamashita, Tomoki Koriyama, Yuki Saito 0001, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
INTERSPEECH | 6 |
| 2020 | Parallel Corpus for Japanese Spoken-to-Written Style ConversionabstractWith the increase of automatic speech recognition (ASR) applications, spoken-to-written style conversion that transforms spoken-style text into written-style text is becoming an important technology to increase the readability of ASR transcriptions. To establish such conversion technology, a parallel corpus of spoken-style text and written-style text is beneficial because it can be utilized for building end-to-end neural sequence transformation models. Spoken-to-written style conversion involves multiple conversion problems including punctuation restoration, disfluency detection, and simplification. However, most existing corpora tend to be made for just one of these conversion problems. In addition, in Japanese, we have to consider not only general spoken-to-written style conversion problems but also Japanese-specific ones, such as language style unification (e.g., polite, frank, and direct styles) and omitted postpositional particle expressions restoration. Therefore, we created a new Japanese parallel corpus of spoken-style text and written-style text that can simultaneously handle general problems and Japanese-specific ones. To make this corpus, we prepared four types of spoken-style text and utilized a crowdsourcing service for manually converting them into written-style text. This paper describes the building setup of this corpus and reports the baseline results of spoken-to-written style conversion using the latest neural sequence transformation models. Mana Ihori, Akihiko Takashima, Ryo Masumura |
LREC | 3 |
| 2020 | Generating Responses that Reflect Meta Information in User-Generated Question Answer PairsabstractThis paper concerns the problem of realizing consistent personalities in neural conversational modeling by using user generated question-answer pairs as training data. Using the framework of role play-based question answering, we collected single-turn question-answer pairs for particular characters from online users. Meta information was also collected such as emotion and intimacy related to question-answer pairs. We verified the quality of the collected data and, by subjective evaluation, we also verified their usefulness in training neural conversational models for generating utterances reflecting the meta information, especially emotion. Takashi Kodama, Ryuichiro Higashinaka, Koh Mitsuda, Ryo Masumura, Yushi Aono, Ryuta Nakamura, Noritake Adachi, Hidetoshi Kawabata |
LREC | 4 |
| 2020 | DNN-based Speech Synthesis Using Abundant Tags of Spontaneous Speech CorpusabstractIn this paper, we investigate the effectiveness of using rich annotations in deep neural network (DNN)-based statistical speech synthesis. DNN-based frameworks typically use linguistic information as input features called context instead of directly using text. In such frameworks, we can synthesize not only reading-style speech but also speech with paralinguistic and nonlinguistic features by adding such information to the context. However, it is not clear what kind of information is crucial for reproducing paralinguistic and nonlinguistic features. Therefore, we investigate the effectiveness of rich tags in DNN-based speech synthesis according to the Corpus of Spontaneous Japanese (CSJ), which has a large amount of annotations on paralinguistic features such as prosody, disfluency, and morphological features. Experimental evaluation results shows that the reproducibility of paralinguistic features of synthetic speech was enhanced by adding such information as context. Yuki Yamashita, Tomoki Koriyama, Yuki Saito 0001, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
LREC | 6 |
| 2020 | Customer Satisfaction Estimation in Contact Center Calls Based on a Hierarchical Multi-Task ModelabstractThis article presents a novel customer satisfaction (CS) estimation method that outputs both turn-level and call-level estimations simultaneously. Our key idea is to directly apply turn-level estimation results to call-level estimation and optimize them jointly; previous works treat both as being independent. Our proposal applies long short-term memory recurrent neural networks (LSTM-RNNs) to turn-level and call-level CS estimation to capture long-range sequential context in contact center calls. In addition, both networks are hierarchically stacked so as to use turn-level estimation results for call-level estimation directly. In order to learn the relationship between the two tasks, we also introduce joint optimization training to the stacked model. Several analyses of turn-level and call-level CS are provided on acted and real calls to support the proposed method. Experiments show that the proposed framework outperforms the conventional methods in both turn-level and call-level estimations. Atsushi Ando, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono, Tomoki Toda |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Improving Speech-Based End-of-Turn Detection Via Cross-Modal Representation Learning with Punctuated Text DataabstractThis paper presents a novel training method for speech-based end-of-turn detection for which not only manually annotated speech data sets but also punctuated text data sets are utilized. The speech-based end-of-turn detection estimates whether a target speaker's utterance is ended or not using speech information. In previous studies, the speech-based end-of-turn detection models were trained using only speech data sets that contained manually annotated end-of-turn labels. However, since the amounts of annotated speech data sets are often limited, the end-of-turn detection models were unable to correctly handle a wide variety of speech patterns. In order to mitigate the data scarcity problem, our key idea is to leverage punctuated text data sets for building more effective speech-based end-of-turn detection. Therefore, the proposed method introduces cross-modal representation learning to construct a speech encoder and a text encoder that can map speech and text with the same lexical information into similar vector representations. This enables us to train speech-based end-of-turn detection models from the punctuated text data sets by tackling text-based sentence boundary detection. In experiments on contact center calls, we show that speech-based end-of-turn detection models using hierarchical recurrent neural networks can be improved through the use of punctuated text data sets. Ryo Masumura, Mana Ihori, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Takanobu Oba, Ryuichiro Higashinaka |
ASRU | 1 |
| 2019 | Generalized Large-Context Language Models Based on Forward-Backward Hierarchical Recurrent Encoder-Decoder ModelsabstractThis paper presents a generalized form of large-context language models (LCLMs) that can take linguistic contexts beyond utterance boundaries into consideration. In discourse-level and conversation-level automatic speech recognition (ASR) tasks, which have to handle a series of utterances, it is essential to capture long-range linguistic contexts beyond utterance boundaries. The LCLMs of previous studies mainly focused on utilizing past contexts, and none fully utilized future contexts because LMs typically process words in a time-ordered manner. Our key idea is to introduce the LCLMs into the situation where ASR results of the whole series of utterances are given by a first decoding pass. This situation makes it possible for the LCLMs to leverage future contexts. In this paper, we propose generalized LCLMs (GLCLMs) based on forward-backward hierarchical recurrent encoder-decoder models in which generative probabilities of individual utterances are computed by leveraging not only past contexts but also future contexts beyond utterance boundaries. In order to efficiently introduce GLCLMs to ASR, we also propose a global-context iterative rescoring method that repeatedly rescores the ASR hypotheses of an individual utterance by using surrounding ASR hypotheses. Experiments on discourse-level ASR tasks demonstrate the effectiveness of our GLCLM approach. Ryo Masumura, Mana Ihori, Tomohiro Tanaka, Itsumi Saito, Kyosuke Nishida, Takanobu Oba |
ASRU | 1 |
| 2019 | Large Context End-to-end Automatic Speech Recognition via Extension of Hierarchical Recurrent Encoder-decoder ModelsabstractThis paper describes a novel end-to-end automatic speech recognition (ASR) method that takes into consideration long-range sequential context information beyond utterance boundaries. In spontaneous ASR tasks such as those for discourses and conversations, the input speech often comprises a series of utterances. Accordingly, the relationships between the utterances should be leveraged for transcribing the individual utterances. While most previous end-to-end ASR methods only focus on utterance-level ASR that handles single utterances independently, the proposed method (which we call "large-context end-to-end ASR") can explicitly utilize relationships between a current target utterance and all preceding utterances. The method is modeled by combining an attention-based encoder-decoder model, which is one of the most representative end-to-end ASR models, with hierarchical recurrent encoder-decoder models, which are effective language models for capturing long-range sequential contexts beyond the utterance boundaries. Experiments on Japanese discourse speech tasks demonstrate the proposed method yields significant ASR performance improvements compared with the conventional utterance-level end-to-end ASR system. Ryo Masumura, Tomohiro Tanaka, Takafumi Moriya, Yusuke Shinohara, Takanobu Oba, Yushi Aono |
ICASSP | 1 |
| 2019 | Speech Emotion Recognition Based on Multi-Label Emotion Existence Model
Atsushi Ando, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 2 |
| 2019 | End-to-End Automatic Speech Recognition with a Reconstruction Criterion Using Speech-to-Text and Text-to-Speech Encoder-Decoders
Ryo Masumura, Hiroshi Sato 0002, Tomohiro Tanaka, Takafumi Moriya, Yusuke Ijima, Takanobu Oba |
INTERSPEECH | 1 |
| 2019 | Improving Conversation-Context Language Models with Multiple Spoken Language Understanding Models
Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Hosana Kamiyama, Takanobu Oba, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 1 |
| 2019 | Joint Maximization Decoder with Neural Converters for Fully Neural Network-Based Japanese Speech Recognition
Takafumi Moriya, Tomohiro Tanaka, Ryo Masumura, Yusuke Shinohara, Yoshikazu Yamaguchi, Yushi Aono |
INTERSPEECH | 4 |
| 2019 | A Joint End-to-End and DNN-HMM Hybrid Automatic Speech Recognition System with Transferring Sharable Knowledge
Tomohiro Tanaka, Ryo Masumura, Takafumi Moriya, Takanobu Oba, Yushi Aono |
INTERSPEECH | 2 |
| 2019 | Recurrent out-of-vocabulary word detection based on distribution of features
Taichi Asami, Ryo Masumura, Yushi Aono, Koichi Shinoda |
Comput. Speech Lang. | 2 |
| 2018 | Multi-task and Multi-lingual Joint Learning of Neural Lexical Utterance Classification based on Partially-shared ModelingabstractThis paper is an initial study on multi-task and multi-lingual joint learning for lexical utterance classification. A major problem in constructing lexical utterance classification modules for spoken dialogue systems is that individual data resources are often limited or unbalanced among tasks and/or languages. Various studies have examined joint learning using neural-network based shared modeling; however, previous joint learning studies focused on either cross-task or cross-lingual knowledge transfer. In order to simultaneously support both multi-task and multi-lingual joint learning, our idea is to explicitly divide state-of-the-art neural lexical utterance classification into language-specific components that can be shared between different tasks and task-specific components that can be shared between different languages. In addition, in order to effectively transfer knowledge between different task data sets and different language data sets, this paper proposes a partially-shared modeling method that possesses both shared components and components specific to individual data sets. We demonstrate the effectiveness of proposed method using Japanese and English data sets with three different lexical utterance classification tasks. Ryo Masumura, Tomohiro Tanaka, Ryuichiro Higashinaka, Hirokazu Masataki, Yushi Aono |
COLING | 1 |
| 2018 | Adversarial Training for Multi-task and Multi-lingual Joint Modeling of Utterance Intent ClassificationabstractThis paper proposes an adversarial training method for the multi-task and multi-lingual joint modeling needed for utterance intent classification.In joint modeling, common knowledge can be efficiently utilized among multiple tasks or multiple languages.This is achieved by introducing both languagespecific networks shared among different tasks and task-specific networks shared among different languages.However, the shared networks are often specialized in majority tasks or languages, so performance degradation must be expected for some minor data sets.In order to improve the invariance of shared networks, the proposed method introduces both language-specific task adversarial networks and task-specific language adversarial networks; both are leveraged for purging the task or language dependencies of the shared networks.The effectiveness of the adversarial training proposal is demonstrated using Japanese and English data sets for three different utterance intent classification tasks. Ryo Masumura, Yusuke Shinohara, Ryuichiro Higashinaka, Yushi Aono |
EMNLP | 1 |
| 2018 | Soft-Target Training with Ambiguous Emotional Utterances for DNN-Based Speech Emotion ClassificationabstractThis paper presents a novel emotion classification method for natural speech. One of the problems in the state-of-the-art method based on Deep Neural Network (DNN) is the paucity of the training data compared to model complexity. To solve this problem, this paper utilizes the ambiguous emotional utterances, utterances that have no dominant target emotion label. While previous work ignored ambiguous emotional utterances for training, the proposed method leverages all annotated labels via soft-target training. In addition, this paper modifies the soft-target training in order to effectively handle both clear and ambiguous emotional utterances. Experiments show that the proposed method yields performance improvements in terms of both weighted and unweighted accuracies. Atsushi Ando, Satoshi Kobashikawa, Hosana Kamiyama, Ryo Masumura, Yusuke Ijima, Yushi Aono |
ICASSP | 4 |
| 2018 | Neural Confnet Classification: Fully Neural Network Based Spoken Utterance Classification Using Word Confusion NetworksabstractThis paper describes neural ConfNet classification, a novel fully neural network based spoken utterance classification method that uses word confusion networks (ConfNets). Our motivation is to establish a spoken utterance classification method that can precisely understand natural language and robustly handle automatic speech recognition (ASR) errors. Remarkable progress has been made in neural networks for accurate modeling, however, most previous methods could not handle ASR errors since they were developed for reference transcriptions. Therefore, in our work we utilized ConfNets, which are compact and efficient graph representations of ASR hypotheses. Our idea is to regard the ConfNet as a sequence of bag-of-weighted-arcs and introduce a mechanism that converts the bag-of-weighted-arcs into a continuous representation called a modified weighted sum representation. This enables us to flexibly connect ConfNets to arbitrary model structures developed for reference transcriptions. We demonstrate the effectiveness of the neural ConfNet classification in dialogue act, extended named entity, and question type classification tasks. Ryo Masumura, Yusuke Ijima, Taichi Asami, Hirokazu Masataki, Ryuichiro Higashinaka |
ICASSP | 1 |
| 2018 | Automatic Question Detection from Acoustic and Phonetic Features Using Feature-wise Pre-training
Atsushi Ando, Reine Asakawa, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 3 |
| 2018 | Role Play Dialogue Aware Language Models Based on Conditional Hierarchical Recurrent Encoder-Decoder
Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Hirokazu Masataki, Yushi Aono |
INTERSPEECH | 1 |
| 2018 | Neural Error Corrective Language Models for Automatic Speech Recognition
Tomohiro Tanaka, Ryo Masumura, Hirokazu Masataki, Yushi Aono |
INTERSPEECH | 2 |
| 2018 | Neural Dialogue Context Online End-of-Turn DetectionabstractThis paper proposes a fully neural network based dialogue-context online end-of-turn detection method that can utilize longrange interactive information extracted from both target speaker's and interlocutor's utterances.In the proposed method, we combine multiple time-asynchronous long short-term memory recurrent neural networks, which can capture target speaker's and interlocutor's multiple sequential features, and their interactions.On the assumption of applying the proposed method to spoken dialogue systems, we introduce target speaker's acoustic sequential features and interlocutor's linguistic sequential features, each of which can be extracted in an online manner.Our evaluation confirms the effectiveness of taking dialogue context formed by the target speaker's utterances and interlocutor's utterances into consideration. Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Ryuichiro Higashinaka, Yushi Aono |
SIGDIAL Conference | 1 |
| 2017 | Domain adaptation of DNN acoustic models using knowledge distillationabstractConstructing deep neural network (DNN) acoustic models from limited training data is an important issue for the development of automatic speech recognition (ASR) applications that will be used in various application-specific acoustic environments. To this end, domain adaptation techniques that train a domain-matched model without overfitting by lever-aging pre-constructed source models are widely used. In this paper, we propose a novel domain adaptation method for DNN acoustic models based on the knowledge distillation framework. Knowledge distillation transfers the knowledge of a teacher model to a student model and offers better generalizability of the student model by controlling the shape of posterior probability distribution of the teacher model, which was originally proposed for model compression. We apply this framework to model adaptation. Our domain adaptation method avoids overfitting of the adapted model trained on limited data by transferring the knowledge of the source model to the adapted model by distillation. Experiments show that the proposed method can effectively avoid the overfitting of convolutional neural network based acoustic models and yield lower error rates than conventional adaptation methods. Taichi Asami, Ryo Masumura, Yoshikazu Yamaguchi, Hirokazu Masataki, Yushi Aono |
ICASSP | 2 |
| 2017 | Parallel phonetically aware DNNs and LSTM-RNNS for frame-by-frame discriminative modeling of spoken language identificationabstractParallel phonetically aware deep neural networks (PPA-DNNs) and long short-term memory recurrent neural networks (PPA-LSTM-RNNs) to enhance frame-by-frame discriminative modeling of spoken language identification are proposed. This idea is inspired by traditional systems based on parallel phoneme recognition followed by language modeling (PPRLM). The proposed methods utilize multiple senone bottleneck features individually extracted from language-dependent senone-based DNNs in a frame-by-frame manner. The multiple senone bottleneck features can yield phonetic awareness to frame-by-frame DNNs and LSTM-RNNs without losing compatibility to real time applications. In experiments, three senone-based DNNs are introduced in order to extract senone bottleneck features, and both single use and parallel use of them are examined. Furthermore, we also examine a combination of PPA-DNNs and PPA-LSTM-RNNs. The proposed method's effectiveness is investigated by comparison with a simple speech aware modeling and traditional systems based on PPRLM. Ryo Masumura, Taichi Asami, Hirokazu Masataki, Yushi Aono |
ICASSP | 1 |
| 2017 | Hierarchical LSTMs with Joint Learning for Estimating Customer Satisfaction from Contact Center Calls
Atsushi Ando, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 2 |
| 2017 | Prosody Aware Word-Level Encoder Based on BLSTM-RNNs for DNN-Based Speech Synthesis
Yusuke Ijima, Nobukatsu Hojo, Ryo Masumura, Taichi Asami |
INTERSPEECH | 3 |
| 2017 | Online End-of-Turn Detection from Speech Based on Stacked Time-Asynchronous Sequential Networks
Ryo Masumura, Taichi Asami, Hirokazu Masataki, Ryo Ishii, Ryuichiro Higashinaka |
INTERSPEECH | 1 |
| 2017 | Parallel Hierarchical Attention Networks with Shared Memory Reader for Multi-Stream Conversational Document Classification
Naoki Sawada, Ryo Masumura, Hiromitsu Nishizaki |
INTERSPEECH | 2 |
| 2016 | Recurrent Out-of-Vocabulary Word Detection Using Distribution of FeaturesabstractThe repeated use of out-of-vocabulary (OOV) words in a spo-\nken document seriously degrades a speech recognizer’s perfor-\nmance. This paper provides a novel method for accurately de-\ntecting such recurrent OOV words. Standard OOV word de-\ntection methods classify each word segment into in-vocabulary\n(IV) or OOV. This word-by-word classification tends to be af-\nfected by sudden vocal irregularities in spontaneous speech,\ntriggering false alarms. To avoid this sensitivity to the irreg-\nularities, our proposal focuses on consistency of the repeated\noccurrence of OOV words. The proposed method preliminar-\nily detects recurrent segments, segments that contain the same\nword, in a spoken document by open vocabulary spoken term\ndiscovery using a phoneme recognizer. If the recurrent seg-\nments are OOV words, features for OOV detection in those\nsegments should exhibit consistency. We capture this consis-\ntency by using the mean and variance (distribution) of features\n(DOF) derived from the recurrent segments, and use the DOF\nfor IV/OOV classification. Experiments illustrate that the pro-\nposed method’s use of the DOF significantly improves its per-\nformance in recurrent OOV word detection.\nIndex Terms: speech recognition, OOV word detection, recur-\nrent OOV words, distribution of features Taichi Asami, Ryo Masumura, Yushi Aono, Koichi Shinoda |
INTERSPEECH | 2 |
| 2016 | Language Identification Based on Generative Modeling of Posteriorgram Sequences Extracted from Frame-by-Frame DNNs and LSTM-RNNs
Ryo Masumura, Taichi Asami, Hirokazu Masataki, Yushi Aono, Sumitaka Sakauchi |
INTERSPEECH | 1 |
| 2015 | Hierarchical Latent Words Language Models for Robust Modeling to Out-Of Domain TasksabstractThis paper focuses on language modeling with adequate robustness to support different domain tasks.To this end, we propose a hierarchical latent word language model (h-LWLM).The proposed model can be regarded as a generalized form of the standard LWLMs.The key advance is introducing a multiple latent variable space with hierarchical structure.The structure can flexibly take account of linguistic phenomena not present in the training data.This paper details the definition as well as a training method based on layer-wise inference and a practical usage in natural language processing tasks with an approximation technique.Experiments on speech recognition show the effectiveness of h-LWLM in out-of domain tasks. Ryo Masumura, Taichi Asami, Takanobu Oba, Hirokazu Masataki, Sumitaka Sakauchi, Akinori Ito |
EMNLP | 1 |
| 2015 | Training data selection for acoustic modeling via submodular optimization of joint kullback-leibler divergence
Taichi Asami, Ryo Masumura, Hirokazu Masataki, Manabu Okamoto, Sumitaka Sakauchi |
INTERSPEECH | 2 |
| 2015 | Combinations of various language model technologies including data expansion and adaptation in spontaneous speech recognition
Ryo Masumura, Taichi Asami, Takanobu Oba, Hirokazu Masataki, Sumitaka Sakauchi, Akinori Ito |
INTERSPEECH | 1 |
| 2015 | Latent words recurrent neural network language models
Ryo Masumura, Taichi Asami, Takanobu Oba, Hirokazu Masataki, Sumitaka Sakauchi, Akinori Ito |
INTERSPEECH | 1 |
| 2015 | Discourse Relation Recognition by Comparing Various Units of Sentence Expression with Recursive Neural Network
Atsushi Otsuka, Toru Hirano, Chiaki Miyazaki, Ryo Masumura, Ryuichiro Higashinaka, Toshiro Makino, Yoshihiro Matsuo |
PACLIC | 4 |
| 2014 | Role play dialogue topic model for language model adaptation in multi-party conversation speech recognitionabstractThis paper introduces an unsupervised language model adaptation technique for multi-party conversation speech recognition. The use of topic models provides one of the most accurate frameworks for unsupervised language model adaptation since they can inject long-range topic information into language models. However, conventional topic models are not suitable for multi-party conversation because they assume that each speech set has each different topic. In a multi-party conversation, each speaker will share the same conversation topic and each speaker utterance will depend on both topic and speaker role. Accordingly, this paper proposes new concept of the “role play dialogue topic model” to utilize multiparty conversation attributes. The proposed topic model can share the topic distribution among each speaker and can also consider both topic and speaker role. The proposed topic model based adaptation realizes a new framework that sets multiple recognition hypotheses for each speaker and simultaneously adapts a language model for each speaker role. We use a call center dialogue data set in speech recognition experiments to show the effectiveness of the proposed method. Ryo Masumura, Takanobu Oba, Hirokazu Masataki, Osamu Yoshioka, Satoshi Takahashi |
ICASSP | 1 |
| 2014 | Read and spontaneous speech classification based on variance of GMM supervectors
Taichi Asami, Ryo Masumura, Hirokazu Masataki, Sumitaka Sakauchi |
INTERSPEECH | 2 |
| 2014 | Mixture of latent words language models for domain adaptation
Ryo Masumura, Taichi Asami, Takanobu Oba, Hirokazu Masataki, Sumitaka Sakauchi |
INTERSPEECH | 1 |
| 2013 | Use of latent words language models in ASR: A sampling-based implementationabstractThis paper applies the latent words language model (LWLM) to automatic speech recognition (ASR). LWLMs are trained taking into account related words, i.e., grouping of similar words in terms of meaning and syntactic role. This means, for example, if a technical word and a general word play a similar syntactic role, they are given a similar probability. This is expected that the LWLM performs robustly over multiple domains. Furthermore, we can expect that the interpolation of the LWLM and a standard n-gram LM will be effective since each of the LMs have different learning criterion. In addition, this paper also describes an approximation method of the LWLM for ASR, in which words are randomly sampled on the LWLM and then a standard word n-gram language model is trained. This enables us one-pass decoding. Our experimental results show that the LWLM performs comparable to the hierarchical Pitman-Yor language model (HPYLM) in a target domain task, and more robustly performs in out-domain tasks. Moreover, an interpolation model with the HPYLM provides a lower word error rate in all the tasks. Ryo Masumura, Hirokazu Masataki, Takanobu Oba, Osamu Yoshioka, Satoshi Takahashi |
ICASSP | 1 |
| 2013 | Viterbi decoding for latent words language models using gibbs sampling
Ryo Masumura, Hirokazu Masataki, Takanobu Oba, Osamu Yoshioka, Satoshi Takahashi |
INTERSPEECH | 1 |
| 2011 | Training a Language Model Using Webdata for Large Vocabulary Japanese Spontaneous Speech Recognition
Ryo Masumura, Seongjun Hahm, Akinori Ito |
INTERSPEECH | 1 |
| 2011 | Language Model Expansion Using Webdata for Spoken Document RetrievalabstractIn recent years, there has been increasing demand for ad hoc retrieval of spoken documents. We can use existing text retrieval methods by transcribing spoken documents into text data using a Large Vocabulary Continuous Speech Recognizer (LVCSR). However, retrieval performance is severely deteriorated by recognition errors and out-of-vocabulary (OOV) words. To solve these problems, we previously proposed an expansion method that compensates the transcription by using text data downloaded from the Web. In this paper, we introduce two improvements to the existing document expansion framework. First, we use a large-scale sample database of webdata as the source of relevant documents, thus avoiding the bias introduced by choosing keywords in the existing methods. Next, we use a document retrieval method based on a statistical language model (SLM), which is a popular framework in information retrieval, and also propose a new smoothing method considering recognition errors and missing keywords. Retrieval experiments show that the proposed methods yield a good results. Index Terms: Spoken document retrieval, statistical language models, World Wide Web Ryo Masumura, Seongjun Hahm, Akinori Ito |
INTERSPEECH | 1 |