EDBT 2026 Demo / reviewers in the wild / expert
Atsushi Ando
dblp:173/6654
· DBLP profile ↗
43ranked-venue papers
10as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 9 first-author · 29 since 2021Artificial intelligence and machine learning · 28 · 6 first-author · 17 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Microphone array geometry-independent multi-talker distant ASR: NTT system for DASR task of the CHiME-8 challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando, Takatomo Kano, Hiroshi Sato 0002, Rintaro Ikeshita, Takafumi Moriya, Shota Horiguchi, Kohei Matsuura, Atsunori Ogawa, Alexis Plaquet, Takanori Ashihara, Tsubasa Ochiai, Masato Mimura, Marc Delcroix, Tomohiro Nakatani, Taichi Asami, Shoko Araki |
Comput. Speech Lang. | 3 |
| 2025 | Predictive ASR and Turn-taking Prediction at Once: Towards More Responsive Spoken Dialog SystemabstractSpoken dialog systems usually wait for users to finish speaking before generating responses, resulting in response delays. A possible solution for reducing the response delay is to predict future words and/or turn-ends while the user is speaking. To realize this, we propose a method to jointly perform predictive automatic speech recognition and turn-taking prediction. Our model receives partial utterances as input and performs speech recognition, future word prediction, and turntaking prediction via autoregressive decoding. It enables turntaking prediction based on prosodic and linguistic cues of observed partial utterances and predicted future linguistic cues. We also incorporate dialogue contexts to improve the performance. Experiments on the Switchboard corpus showed that our multi-task model outperforms a single-task model in turn-taking prediction. We found that conditioning turn-taking prediction on predicted words improved performance when words were correctly predicted. Ryo Fukuda, Takatomo Kano, Naohiro Tawara, Marc Delcroix, Atsunori Ogawa, Yuya Chiba, Atsushi Ando |
ASRU | 7 |
| 2025 | Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?abstractNeural speaker diarization is widely used for overlapaware speaker diarization, but it requires large multi-speaker datasets for training. To meet this data requirement, large datasets are often constructed by combining multiple corpora, including those originally designed for multi-speaker automatic speech recognition (ASR). However, ASR datasets often feature loosely defined segment boundaries that do not align with the stricter conventions of diarization benchmarks. In this work, we show that such boundary looseness significantly impacts the diarization error rate, reducing evaluation reliability. We also reveal that models trained on data with varying boundary precision tend to learn dataset-specific looseness, leading to poor generalization across out-of-domain datasets. Training with standardized tight boundaries via forced alignment improves not only diarization performance, especially in streaming scenarios, but also ASR performance when combined with simple post-processing. Shota Horiguchi, Naohiro Tawara, Takanori Ashihara, Atsushi Ando, Marc Delcroix |
ASRU | 4 |
| 2025 | Speech Emotion Recognition Based on Large-Scale Automatic Speech RecognizerabstractThis paper proposes a novel speech emotion recognition (SER) method that fully leverages the architecture of Whisper, a large-scale automatic speech recognition (ASR) model. The conventional SER models using a pre-trained speech encoder may fail to capture linguistic content since their decoders are too simple. Our proposed method addresses this shortcoming by adopting the decoder of Whisper, which has been discarded in conventional SER, to leverage its language modeling capability. The proposed method introduces special tokens corresponding to the target emotions and then fine-tunes the entire Whisper model. Furthermore, we also propose a new training scheme suitable for Whisper, named serialized multi-task learning (SerialMTL), to consider various speech information as context for the objective SER task. In SerialMTL, the model initially predicts subtask tokens, such as transcription and gender tokens, and then estimates the emotion token. An advantage of the proposed method is the simplicity of the model structure, even when adding any new subtasks. Experimental results show that our model, based on the entire Whisper, achieves better SER performance than the conventional model and further improves with SerialMTL training via ASR and gender recognition subtasks. Ryo Fukuda, Takatomo Kano, Atsushi Ando, Atsunori Ogawa |
ICASSP | 3 |
| 2025 | Guided Speaker EmbeddingabstractThis paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped multi-speaker audio processing are typically two-staged: i) segment-level processing and ii) inter-segment speaker matching. Speaker embeddings are often used for the latter purpose. Typical speaker embedding extraction approaches only use single-speaker intervals to avoid corrupting the embeddings with speech from interference speakers. However, this often makes speaker embeddings impossible to extract because sufficiently long non-overlapping intervals are not always available. In this paper, we propose using speaker activities as clues to extract the embedding of the speaker-of-interest directly from overlapping speech. Specifically, we concatenate the activity of target and non-target speakers to acoustic features before being fed to the model. We also condition the attention weights used for pooling so that the attention weights of the intervals in which the target speaker is inactive are zero. The effectiveness of the proposed method is demonstrated in speaker verification and speaker diarization. Shota Horiguchi, Takafumi Moriya, Atsushi Ando, Takanori Ashihara, Hiroshi Sato 0002, Naohiro Tawara, Marc Delcroix |
ICASSP | 3 |
| 2025 | Mamba-based Segmentation Model for Speaker DiarizationabstractMamba is a newly proposed architecture that behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization, as attention-based models have unsuitable memory requirements for long-form audio, and traditional RNN capabilities are too limited. In this paper, we propose to assess the potential of Mamba for diarization by comparing the state-of-the-art neural segmentation of the pyannote pipeline with our proposed Mamba-based variant. Mamba’s stronger processing capabilities allow usage of longer local windows, which significantly improve diarization quality by making the speaker embedding extraction more reliable. We find Mamba to be a superior alternative to both traditional RNN and the tested attention-based model. Our proposed Mamba-based system achieves state-of-the-art performance on three widely used diarization datasets. Alexis Plaquet, Naohiro Tawara, Marc Delcroix, Shota Horiguchi, Atsushi Ando, Shoko Araki |
ICASSP | 5 |
| 2025 | Multi-channel Speaker Counting for EEND-VC-based Speaker Diarization on Multi-domain ConversationabstractThis paper proposes a speaker counting scheme using multichannel microphones for end-to-end neural diarization with a vector clustering (EEND-VC) speaker diarization pipeline. The EEND-VC-based system estimates the number of speakers by clustering speaker embeddings from small chunks. However, conventional speaker counting struggles in short sessions with limited available embeddings. We address this issue by leveraging the most possible embeddings from multichannel signals to increase the number of embeddings. One challenge in using embeddings across channels is the biases caused by channel differences. To mitigate this issue, we extend the EEND-VC pipeline with two modifications: (1) applying speech enhancement before extracting speaker embedding to capture the speaker characteristics even from short chunks and (2) grouping microphones based on inter-channel correlation to perform speaker counting within each group and then aggregating these channel-wise results. The proposed scheme was integrated into our CHiME-8 diarization pipeline, achieving superior speaker counting accuracy compared to the CHiME-8 baseline, with 54.2% and 61.4% improvements in the development and evaluation sets, respectively. Naohiro Tawara, Atsushi Ando, Shota Horiguchi, Marc Delcroix |
ICASSP | 2 |
| 2025 | Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
Shota Horiguchi, Takanori Ashihara, Marc Delcroix, Atsushi Ando, Naohiro Tawara |
INTERSPEECH | 4 |
| 2025 | Pretraining Multi-Speaker Identification for Neural Speaker Diarization
Shota Horiguchi, Atsushi Ando, Naohiro Tawara, Marc Delcroix |
INTERSPEECH | 2 |
| 2024 | NTT Speaker Diarization System for Chime-7: Multi-Domain, Multi-Microphone end-to-end and Vector Clustering DiarizationabstractThis paper details our speaker diarization system designed for multi-domain, multi-microphone casual conversations. The proposed diarization pipeline uses weighted prediction error (WPE)based dereverberation as a front end, and separately applies end-to-end neural diarization with vector clustering (EEND-VC) to each channel. It integrates the diarization result obtained from each channel using diarization output voting error reduction plus overlap (DOVER-Lap). To harness the knowledge from the target domain and the results integrated across all channels, we apply self-supervised adaptation for each session by retraining the EEND-VC with pseudo-labels derived from DOVER-Lap. We incorporated our proposed system into NTT’s submission for a distant automatic speech recognition task in the CHiME-7 challenge. Our system obtained third place in the diarization performance by improving the development and evaluation sets by 65 % and 62 % compared to the organizer-provided, VC-based baseline diarization system. Naohiro Tawara, Marc Delcroix, Atsushi Ando, Atsunori Ogawa |
ICASSP | 3 |
| 2024 | Factor-Conditioned Speaking-Style Captioning
Atsushi Ando, Takafumi Moriya, Shota Horiguchi, Ryo Masumura |
INTERSPEECH | 1 |
| 2024 | SOMSRED: Sequential Output Modeling for Joint Multi-talker Overlapped Speech Recognition and Speaker Diarization
Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Atsushi Ando, Ryo Masumura |
INTERSPEECH | 6 |
| 2024 | Unified Multi-Talker ASR with and without Target-speaker Enrollment
Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Mana Ihori, Naotaka Kawata, Shota Orihashi, Kazutoshi Shinoda, Taiga Yamane, Saki Mizuno, Keita Suzuki, Nobukatsu Hojo, Takafumi Moriya, Atsushi Ando |
INTERSPEECH | 14 |
| 2024 | SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
Hiroshi Sato 0002, Takafumi Moriya, Masato Mimura, Shota Horiguchi, Tsubasa Ochiai, Takanori Ashihara, Atsushi Ando, Kentaro Shinayama, Marc Delcroix |
INTERSPEECH | 7 |
| 2024 | Recursive Attentive Pooling For Extracting Speaker Embeddings From Multi-Speaker RecordingsabstractThis paper proposes a method for extracting speaker embedding for each speaker from a variable-length recording containing multiple speakers. Speaker embeddings are crucial not only for speaker recognition but also for various multi-speaker speech applications such as speaker diarization and target-speaker speech processing. Despite the challenges of obtaining a single speaker’s speech without pre-registration in multi-speaker scenarios, most studies on speaker embedding extraction focus on extracting embeddings only from single-speaker recordings. Some methods have been proposed for extracting speaker embeddings directly from multi-speaker recordings, but they typically require preparing a model for each possible number of speakers or involve complicated training procedures. The proposed method computes the embeddings of multiple speakers by focusing on different parts of the frame-wise embeddings extracted from the input multi-speaker audio. This is achieved by recursively computing attention weights for pooling the frame-wise embeddings. Additionally, we propose using the calculated attention weights to estimate the number of speakers in the recording, which allows the same model to be applied to various numbers of speakers. Experimental evaluations demonstrate the effectiveness of the proposed method in speaker verification and diarization tasks. Shota Horiguchi, Atsushi Ando, Takafumi Moriya, Takanori Ashihara, Hiroshi Sato 0002, Naohiro Tawara, Marc Delcroix |
SLT | 2 |
| 2023 | Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuracy-Robustness TradeoffabstractThis paper addresses the tradeoff between standard accuracy on clean examples and robustness against adversarial examples in deep neural networks (DNNs). Although adversarial training (AT) improves robustness, it degrades the standard accuracy, thus yielding the tradeoff. To mitigate this tradeoff, we propose a novel AT method called ARREST, which comprises three components: (i) adversarial finetuning (AFT), (ii) representation-guided knowledge distillation (RGKD), and (iii) noisy replay (NR). AFT trains a DNN on adversarial examples by initializing its parameters with a DNN that is standardly pretrained on clean examples. RGKD and NR respectively entail a regularization term and an algorithm to preserve latent representations of clean examples during AFT. RGKD penalizes the distance between the representations of the standardly pretrained and AFT DNNs. NR switches input adversarial examples to nonadversarial ones when the representation changes significantly during AFT. By combining these components, ARREST achieves both high standard accuracy and robustness. Experimental results demonstrate that ARREST mitigates the tradeoff more effectively than previous AT-based methods do. Shin'ya Yamaguchi, Shoichiro Takeda, Sekitoshi Kanai, Naoki Makishima, Atsushi Ando, Ryo Masumura |
ICCV | 6 |
| 2023 | OnDA-DETR: Online Domain Adaptation for Detection Transformers with Self-Training FrameworkabstractThis paper presents a novel method for online domain adaptation (OnDA) for DEtection TRansformer (DETR)-based object detection models called OnDA-DETR. OnDA is a domain adaptation paradigm that adapts a model trained on the source domain data to perform well on the target domain in an online manner during testing, using only the unlabeled test data from the target domain. Due to challenging and realistic problem settings, OnDA has garnered significant attention. However, OnDA methods for DETR-based models, which have demonstrated excellent performance in object detection research fields, had not been developed. OnDA-DETR is the first OnDA method specifically designed for DETR-based models. OnDA-DETR incorporates a self-training framework that generates pseudo-labels for the unlabeled target domain data. To effectively incorporate the self-training framework into DETR-based models, we leverage recall-aware pseudo-labeling and quality-aware training in OnDA-DETR. Experimental results indicate that OnDA-DETR improves the performance of the source-trained model by about 3.0 % points through OnDA. Taiga Yamane, Naoki Makishima, Keita Suzuki, Atsushi Ando, Ryo Masumura |
ICIP | 5 |
| 2023 | Joint Autoregressive Modeling of End-to-End Multi-Talker Overlapped Speech Recognition and Utterance-level Timestamp Prediction
Naoki Makishima, Keita Suzuki, Atsushi Ando, Ryo Masumura |
INTERSPEECH | 4 |
| 2023 | End-to-End Joint Target and Non-Target Speakers ASR
Ryo Masumura, Naoki Makishima, Taiga Yamane, Yoshihiko Yamazaki, Saki Mizuno, Mana Ihori, Mihiro Uchida, Keita Suzuki, Hiroshi Sato 0002, Tomohiro Tanaka, Akihiko Takashima, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando |
INTERSPEECH | 15 |
| 2023 | Multi-region CNN-Transformer for Micro-gesture Recognition in Face and Upper BodyabstractThis paper presents a novel task that recognizes from a video unintentional micro-gestures (UMGs), which are movements made by people unconsciously and unintentionally. Recognizing UMGs is crucial because they reveal a person’s underlying psychological state. Since a UMG is composed of subtle sequential movements, the recognition model must be able to capture accurate information in both the spatial and temporal directions. Therefore, we utilize a convolutional neural network (CNN) to capture information in the spatial direction and a Transformer to merge the features extracted by the CNN in the temporal direction. However, this model often misrecognizes UMGs because it is not possible to capture slight differences in movements, such as in the face and mouth regions. To address this issue, we propose a novel model for UMG recognition, the Multi-Region CNN-Transformer model, that inputs cropped videos from multiple upper body regions simultaneously. The key advance of our method is to capture subtle changes in regions such as the upper body, face, head, and mouth for recognizing UMGs. We demonstrate the effectiveness of the proposed method through experiments using our newly created UMG dataset for this task. Keita Suzuki, Ryo Masumura, Atsushi Ando, Naoki Makishima |
MMAsia | 4 |
| 2022 | Customer Satisfaction Estimation Using Unsupervised Representation Learning with Multi-Format Prediction LossabstractWe propose a new Customer Satisfaction Estimation (CSE) method that utilizes unsupervised representation learning. Though conventional methods have improved both the heuristic features and the estimation models, their performance is still insufficient as only small amounts of labeled training data can be expected. To mitigate this problem, the proposed method leverages a large amount of unlabeled data by unsupervised representation learning based on self-training. The key advance of the proposed method is to introduce a Multi-Format Prediction (MFP) loss to improve the performance of self-training for the inputs that contain both continuous and biased discrete features such as the number of occurrences of a particular word. MFP loss uses two loss functions based on regression and weighted binary classification to reconstruct both types of features with high accuracy. Experiments on real English contact center calls reveal the improved CSE performance attained by the proposed method. Atsushi Ando, Yumiko Murata, Ryo Masumura, Naoki Makishima, Takafumi Moriya, Takanori Ashihara, Hiroshi Sato 0002 |
ICASSP | 1 |
| 2022 | Hybrid RNN-T/Attention-Based Streaming ASR with Triggered Chunkwise Attention and Dual Internal Language Model IntegrationabstractIn this paper we propose improvements to our recently proposed hybrid RNN-T/Attention architecture that includes a shared encoder followed by recurrent neural network-transducer (RNN-T) and triggered attention-based decoders (TAD). The use of triggered attention enables the attention-based decoder (AD) to operate in a streaming manner. When a trigger point is detected by RNN-T, TAD uses the context from the start-of-speech up to that trigger point to compute the attention weights. Consequently, the computation costs and the memory consumptions are quadratically increased with the duration of the utterances because all input features must be stored and used to re-compute the attention weights. In this paper, we use a short context from a few frames prior to each trigger point for attention weight computation resulting in reduced computation and memory costs. We call the proposed framework triggered chunkwise AD (TCAD). We also investigate the effectiveness of internal language model (ILM) estimation approach using both ILMs of RNN-T and TCAD heads for improving RNN-T performance. We confirm in experiments with public and private datasets covering various scenarios that TCAD achieves superior recognition performance while reducing computation costs compared to TAD. Takafumi Moriya, Takanori Ashihara, Atsushi Ando, Hiroshi Sato 0002, Tomohiro Tanaka, Kohei Matsuura, Ryo Masumura, Marc Delcroix, Takahiro Shinozaki |
ICASSP | 3 |
| 2022 | Speaker consistency loss and step-wise optimization for semi-supervised joint training of TTS and ASR using unpaired text dataabstractIn this paper, we investigate the semi-supervised joint training of text to speech (TTS) and automatic speech recognition (ASR), where a small amount of paired data and a large amount of unpaired text data are available.Conventional studies form a cycle called the TTS-ASR pipeline, where the multispeaker TTS model synthesizes speech from text with a reference speech and the ASR model reconstructs the text from the synthesized speech, after which both models are trained with a cycle-consistency loss.However, the synthesized speech does not reflect the speaker characteristics of the reference speech and the synthesized speech becomes overly easy for the ASR model to recognize after training.This not only decreases the TTS model quality but also limits the ASR model improvement.To solve this problem, we propose improving the cycleconsistency-based training with a speaker consistency loss and step-wise optimization.The speaker consistency loss brings the speaker characteristics of the synthesized speech closer to that of the reference speech.In the step-wise optimization, we first freeze the parameter of the TTS model before both models are trained to avoid over-adaptation of the TTS model to the ASR model.Experimental results demonstrate the efficacy of the proposed method. Naoki Makishima, Atsushi Ando, Ryo Masumura |
INTERSPEECH | 3 |
| 2022 | End-to-End Joint Modeling of Conversation History-Dependent and Independent ASR Systems with Multi-History Training
Ryo Masumura, Yoshihiro Yamazaki, Saki Mizuno, Naoki Makishima, Mana Ihori, Mihiro Uchida, Hiroshi Sato 0002, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando |
INTERSPEECH | 14 |
| 2022 | Interactive Co-Learning with Cross-Modal Transformer for Audio-Visual Emotion Recognition
Akihiko Takashima, Ryo Masumura, Atsushi Ando, Yoshihiro Yamazaki, Mihiro Uchida, Shota Orihashi |
INTERSPEECH | 3 |
| 2022 | On the Use of Modality-Specific Large-Scale Pre-Trained Encoders for Multimodal Sentiment AnalysisabstractThis paper investigates the effectiveness and implementation of modality-specific large-scale pre-trained encoders for multimodal sentiment analysis (MSA). Although the effectiveness of pre-trained encoders in various fields has been reported, conventional MSA methods employ them for only linguistic modality, and their application has not been investigated. This paper compares the features yielded by large-scale pre-trained encoders with conventional heuristic features. One each of the largest pre-trained encoders publicly available for each modality are used; CLIP-ViT, WavLM, and BERT for visual, acoustic, and linguistic modalities, respectively. Experiments on two datasets reveal that methods with domain-specific pre-trained encoders attain better performance than those with conventional features in both unimodal and multimodal scenarios. We also find it better to use the outputs of the intermediate layers of the encoders than those of the output layer. The codes are available at https://github.com/ando-hub/MSA_Pretrain. Atsushi Ando, Ryo Masumura, Akihiko Takashima, Naoki Makishima, Keita Suzuki, Takafumi Moriya, Takanori Ashihara, Hiroshi Sato 0002 |
SLT | 1 |
| 2021 | Speech Emotion Recognition Based on Listener Adaptive ModelsabstractThis paper presents a novel speech emotion recognition scheme that can deal with the individuality of emotion perception. Most conventional methods directly model the majority decision of multiple listener’s perceived emotions. However, emotion perception varies with the listener, which means the conventional methods can mismatch the recognition results to human perception. In order to mitigate this problem, we propose a Listener Adaptive (LA) model that reflects emotion recognition criteria of each listener. One-hot listener codes with several adaptation layers are employed in the LA model. The LA model yields the posterior probabilities of the listener-specific perceived emotions. Majority-voted emotion can be also estimated by averaging, in the LA model, the posterior probabilities for all listeners. Experiments on two emotional speech datasets demonstrate that the proposed approach offers improved listener-wise perceived emotion recognition performance in natural speech. Atsushi Ando, Ryo Masumura, Hiroshi Sato 0002, Takafumi Moriya, Takanori Ashihara, Yusuke Ijima, Tomoki Toda |
ICASSP | 1 |
| 2021 | Simpleflat: A Simple Whole-Network Pre-Training Approach for RNN Transducer-Based End-to-End Speech RecognitionabstractRecurrent neural network-transducer (RNN-T) is promising for building time-synchronous end-to-end automatic speech recognition (ASR) systems, in part because it does not need frame-wise alignment between input features and target labels in the training step. Although training without alignment is beneficial, it makes it difficult to discern the relation between input features and output token sequences. This, in effect, degrades RNN-T performance. Our solution is SimpleFlat (SF), a novel and simple whole-network pretraining approach for RNN-T. SF extracts frame-wise alignments on-the-fly from the training dataset, and does not require any external resources. We distribute equal numbers of target tokens to each frame following RNN-T encoder output lengths by repeating each token. The frame-wise tokens so created are shifted, and also used as the prediction network inputs. Therefore, SF can be implemented by cross entropy loss computation as in autoregressive model training. Experiments on Japanese and English ASR tasks demonstrate that SF can effectively improve various RNN-T architectures. Takafumi Moriya, Takanori Ashihara, Tomohiro Tanaka, Tsubasa Ochiai, Hiroshi Sato 0002, Atsushi Ando, Yusuke Ijima, Ryo Masumura, Yusuke Shinohara |
ICASSP | 6 |
| 2021 | Phoneme Duration Modeling Using Speech Rhythm-Based Speaker Embeddings for Multi-Speaker Speech Synthesis
Kenichi Fujita, Atsushi Ando, Yusuke Ijima |
Interspeech | 2 |
| 2021 | Streaming End-to-End Speech Recognition for Hybrid RNN-T/Attention Architecture
Takafumi Moriya, Tomohiro Tanaka, Takanori Ashihara, Tsubasa Ochiai, Hiroshi Sato 0002, Atsushi Ando, Ryo Masumura, Marc Delcroix, Taichi Asami |
Interspeech | 6 |
| 2020 | Sequence-Level Consistency Training for Semi-Supervised End-to-End Automatic Speech RecognitionabstractThis paper presents a novel semi-supervised end-to-end automatic speech recognition (ASR) method that employs consistency training with the use of unlabeled data. In consistency training, unlabeled data can be utilized for constraining a model such that it becomes invariant to small deformation. In fact, considering consistency can make the model robust to a variety of input examples. While previous studies have applied consistency training to primitive classification problems, no studies have employed consistency training to tackle sequence-to-sequence generation problems including end-to- end ASR. One problem is that existing consistency training schemes cannot take sequence-level generation consistency into consideration. In this paper, we propose a sequence-level consistency training scheme specialized to handle sequence-to-sequence generation problems. Our key idea is to consider the consistency of the generation function by utilizing beam search decoding results. For semi- supervised learning, we adopt Transformer as the end-to-end ASR model, and SpecAugment as the deformation function in consistency training. Our experiments show that our semi-supervised learning proposal with sequence-level consistency training can efficiently improve ASR performance using unlabeled speech data. Ryo Masumura, Mana Ihori, Akihiko Takashima, Takafumi Moriya, Atsushi Ando, Yusuke Shinohara |
ICASSP | 5 |
| 2020 | Customer Satisfaction Estimation in Contact Center Calls Based on a Hierarchical Multi-Task ModelabstractThis article presents a novel customer satisfaction (CS) estimation method that outputs both turn-level and call-level estimations simultaneously. Our key idea is to directly apply turn-level estimation results to call-level estimation and optimize them jointly; previous works treat both as being independent. Our proposal applies long short-term memory recurrent neural networks (LSTM-RNNs) to turn-level and call-level CS estimation to capture long-range sequential context in contact center calls. In addition, both networks are hierarchically stacked so as to use turn-level estimation results for call-level estimation directly. In order to learn the relationship between the two tasks, we also introduce joint optimization training to the stacked model. Several analyses of turn-level and call-level CS are provided on acted and real calls to support the proposed method. Experiments show that the proposed framework outperforms the conventional methods in both turn-level and call-level estimations. Atsushi Ando, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono, Tomoki Toda |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Improving Speech-Based End-of-Turn Detection Via Cross-Modal Representation Learning with Punctuated Text DataabstractThis paper presents a novel training method for speech-based end-of-turn detection for which not only manually annotated speech data sets but also punctuated text data sets are utilized. The speech-based end-of-turn detection estimates whether a target speaker's utterance is ended or not using speech information. In previous studies, the speech-based end-of-turn detection models were trained using only speech data sets that contained manually annotated end-of-turn labels. However, since the amounts of annotated speech data sets are often limited, the end-of-turn detection models were unable to correctly handle a wide variety of speech patterns. In order to mitigate the data scarcity problem, our key idea is to leverage punctuated text data sets for building more effective speech-based end-of-turn detection. Therefore, the proposed method introduces cross-modal representation learning to construct a speech encoder and a text encoder that can map speech and text with the same lexical information into similar vector representations. This enables us to train speech-based end-of-turn detection models from the punctuated text data sets by tackling text-based sentence boundary detection. In experiments on contact center calls, we show that speech-based end-of-turn detection models using hierarchical recurrent neural networks can be improved through the use of punctuated text data sets. Ryo Masumura, Mana Ihori, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Takanobu Oba, Ryuichiro Higashinaka |
ASRU | 4 |
| 2019 | Does the Lombard Effect Improve Emotional Communication in Noise? - Analysis of Emotional Speech Acted in NoiseabstractSpeakers usually adjust their way of talking in noisy environments involuntarily for effective communication. This adaptation is known as the Lombard effect. Although speech accompanying the Lombard effect can improve the intelligibility of a speaker's voice, the changes in acoustic features (e.g. fundamental frequency, speech intensity, and spectral tilt) caused by the Lombard effect may also affect the listener's judgment of emotional content. To the best of our knowledge, there is no published study on the influence of the Lombard effect in emotional speech. Therefore, we recorded parallel emotional speech waveforms uttered by 12 speakers under both quiet and noisy conditions in a professional recording studio in order to explore how the Lombard effect interacts with emotional speech. By analyzing confusion matrices and acoustic features, we aim to answer the following questions: 1) Can speakers express their emotions correctly even under adverse conditions? 2) Can listeners recognize the emotion contained in speech signals even under noise? 3) How does emotional speech uttered in noise differ from emotional speech uttered in quiet conditions in terms of acoustic characteristic? Yi Zhao 0006, Atsushi Ando, Shinji Takaki, Junichi Yamagishi, Satoshi Kobashikawa |
INTERSPEECH | 2 |
| 2019 | Speech Emotion Recognition Based on Multi-Label Emotion Existence Model
Atsushi Ando, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 1 |
| 2019 | Improving Conversation-Context Language Models with Multiple Spoken Language Understanding Models
Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Hosana Kamiyama, Takanobu Oba, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 3 |
| 2018 | Soft-Target Training with Ambiguous Emotional Utterances for DNN-Based Speech Emotion ClassificationabstractThis paper presents a novel emotion classification method for natural speech. One of the problems in the state-of-the-art method based on Deep Neural Network (DNN) is the paucity of the training data compared to model complexity. To solve this problem, this paper utilizes the ambiguous emotional utterances, utterances that have no dominant target emotion label. While previous work ignored ambiguous emotional utterances for training, the proposed method leverages all annotated labels via soft-target training. In addition, this paper modifies the soft-target training in order to effectively handle both clear and ambiguous emotional utterances. Experiments show that the proposed method yields performance improvements in terms of both weighted and unweighted accuracies. Atsushi Ando, Satoshi Kobashikawa, Hosana Kamiyama, Ryo Masumura, Yusuke Ijima, Yushi Aono |
ICASSP | 1 |
| 2018 | Automatic Question Detection from Acoustic and Phonetic Features Using Feature-wise Pre-training
Atsushi Ando, Reine Asakawa, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 1 |
| 2018 | Role Play Dialogue Aware Language Models Based on Conditional Hierarchical Recurrent Encoder-Decoder
Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Hirokazu Masataki, Yushi Aono |
INTERSPEECH | 3 |
| 2018 | Neural Dialogue Context Online End-of-Turn DetectionabstractThis paper proposes a fully neural network based dialogue-context online end-of-turn detection method that can utilize longrange interactive information extracted from both target speaker's and interlocutor's utterances.In the proposed method, we combine multiple time-asynchronous long short-term memory recurrent neural networks, which can capture target speaker's and interlocutor's multiple sequential features, and their interactions.On the assumption of applying the proposed method to spoken dialogue systems, we introduce target speaker's acoustic sequential features and interlocutor's linguistic sequential features, each of which can be extracted in an online manner.Our evaluation confirms the effectiveness of taking dialogue context formed by the target speaker's utterances and interlocutor's utterances into consideration. Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Ryuichiro Higashinaka, Yushi Aono |
SIGDIAL Conference | 3 |
| 2017 | Hierarchical LSTMs with Joint Learning for Estimating Customer Satisfaction from Contact Center Calls
Atsushi Ando, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 1 |
| 2017 | Interaction and Transition Model for Speech Emotion Recognition in Dialogue
Ruo Zhang, Atsushi Ando, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 2 |
| 2015 | Agreement and disagreement utterance detection in conversational speech by extracting and integrating local features
Atsushi Ando, Taichi Asami, Manabu Okamoto, Hirokazu Masataki, Sumitaka Sakauchi |
INTERSPEECH | 1 |